Rendering on-device

Realtime avatars
that run on iPhone.

Every other realtime avatar is a video call with a server. This one is a face on the glass — drawn where it’s shown, reacting while you’re still talking.

Shipping today25 fps on the Neural Enginenothing streamed

ON DEVICE
25.0 FPS · ANE
0 FRAMES STREAMED
The idle state

Watch the mouth when nobody is speaking.

Mid-sentence, every avatar looks fine. The tell is the pause — a mouth left parted, flapping between words, on a face that never quite stops moving. Buyers running a bake-off score the idle state first, because it is the most honest signal of render quality there is.

It goes quiet in the pause

Mouth closure in silence measures 0.80, against 0.227 for the commercial lipsync tool we trained against. Idle jitter 2.7–4 reversals per second against their 8.45. Their face flutters while you talk. Ours settles.

Still, without looking frozen

It blinks, it attends, it nods when you land a point — driven by your voice and decided on the phone. A face that stops talking should look like it is listening, not like it has hung.

No buffering wheel on a human face

There is no video to stall. On a subway, on hotel wifi, on one bar of signal, the face keeps drawing at full rate. Only the words ever have to travel.
Why on-device

No server draws this face.

25
frames per second, on the Neural Engine
~32MB
of avatar models, delivered over the air
0
bytes of audio, video, transcript or frames reach us
0ms
of network between audio and mouth
simultaneous conversations — every one on its own phone
keeps working if our servers do not
Two architectures side by side. Left, a cloud avatar: the phone sends your voice to a datacenter, which animates the face, renders and encodes it, and streams H.264 back for every frame — four network hops before the mouth moves, and the face stops when the signal does. Right, Yoob: a model is delivered once, then animation, rendering and speech all run on the phone's Neural Engine at 25 fps. Nothing about the face leaves the device, and it keeps going when the signal drops.
The fine print

Where this doesn’t win.

iOS onlyAndroid in development, Web on the roadmap. Need them now — Spatius and bitHuman ship both.
Characters take daysConcierge bake, setup plus an annual licence. No self-serve yet. Stock characters ship same day.
A mouth, not a sceneComposited onto a living host, not a full body in an arbitrary environment.
Your LLM still needs a linkRendering and speech are local. The brain is your provider, your key, your bill.
We lose under ~55k min/moSpatius and bitHuman are cheaper down there. The slider above says so itself.
Pricing

Free for people. Priced per app for products.

Companion

Personal

An hour a day, free, permanently — on-device face and voice, no ads, no card. Premium voice is $19/mo and stays on-device.

SDK

Commercial

$99 a month plus $0.005 a conversation minute, on-device TTS included. Rendered frames are never counted and never charged.

Questions

The ones that decide it.

You’ll start metering after your Series A, right?

A build you’ve already shipped can’t be re-priced — the renderer is compiled into your binary and never asks us for permission to draw. Beyond that it’s contractual: no per-minute, per-session or per-MAU charge on any tier, and shipped apps keep the terms they shipped under.

Other vendors render on the client too. So what?

bitHuman does, and they charge $0.01 a minute where we charge $0.005 with on-device voice included. Others stream compact pose data and rasterise it in the browser — still a server generating motion for every frame of every session. We charge for conversation time, never for frames, and there is no concurrency tier because there is no capacity of ours to reserve. Full comparison.

Is it really zero network?

No. Rendering and speech never touch it — no audio, video, frames or transcripts. Your LLM does. Model delivery is one authenticated fetch on first launch. Ask for a packet capture and check.

Give it a face that’s actually there.