Rendering on-device

Realtime avatars
that run on iPhone.

Every other realtime avatar is a video call with a server. This one is a face on the glass — drawn where it’s shown, reacting while you’re still talking.

Shipping today25 fps on the Neural Enginenothing streamed

ON DEVICE
25.0 FPS · ANE
0 FRAMES STREAMED
The idle state

Watch the mouth when nobody is speaking.

Mid-sentence, every avatar looks fine. The tell is the pause — a mouth left parted, flapping between words, on a face that never quite stops moving. Buyers running a bake-off score the idle state first, because it is the most honest signal of render quality there is.

It goes quiet in the pause

Mouth closure in silence measures 0.80, against 0.227 for the commercial lipsync tool we trained against. Idle jitter 2.7–4 reversals per second against their 8.45. Their face flutters while you talk. Ours settles.

Still, without looking frozen

It blinks, it attends, it nods when you land a point — driven by your voice and decided on the phone. A face that stops talking should look like it is listening, not like it has hung.

No buffering wheel on a human face

There is no video to stall. On a subway, on hotel wifi, on one bar of signal, the face keeps drawing at full rate. Only the words ever have to travel.
Why on-device

No server draws this face.

25
frames per second, on the Neural Engine
~32MB
of avatar models, delivered over the air
0
bytes of video or frames reach us; audio only with Yoob voice
0ms
of network between audio and mouth
simultaneous conversations — every one on its own phone
10min
a live session keeps rendering for up to 10 minutes of a Yoob outage; new sessions need our API
Two architectures side by side. Left, a cloud avatar: the phone sends your voice to a datacenter, which animates the face, renders and encodes it, and streams H.264 back for every frame — four network hops before the mouth moves, and the face stops when the signal does. Right, Yoob: a model is delivered once, then animation, rendering and speech all run on the phone's Neural Engine at 25 fps. Nothing about the face leaves the device, and a live session keeps rendering for up to 10 minutes when the signal drops.
The fine print

Where this doesn’t win.

No Android yetiOS and desktop web (npm @yoob/avatar) ship today; Android is in development and mobile browsers aren’t supported. Spatius and bitHuman ship Android.
Characters take daysConcierge bake, setup plus an annual licence. No self-serve yet. Stock characters ship same day.
A mouth, not a sceneComposited onto a living host, not a full body in an arbitrary environment.
The conversation still needs a linkRendering is local. The brain and voice are your provider on your key, or Yoob voice through our relay.
Not the cheapest per minuteSpatius and bitHuman list lower per-minute rates than our $0.03. The cost calculator says so itself.
Pricing

Free for people. Pay as you go for products.

Companion

Personal

An hour a day, free, permanently — on-device face and voice, no ads, no card. Premium voice is $19/mo and stays on-device.

SDK

Commercial

$0.03 a conversation minute, or $0.08 with Yoob voice, from prepaid credit. Test keys and trial credit are free. Rendered frames are never counted and never charged.

Questions

The ones that decide it.

Will the price change once we depend on you?

We already meter, and the rate is public: $0.03 a conversation minute, or $0.08 with Yoob voice. There is no per-session, per-MAU or per-frame charge, and a monthly spend cap keeps the bill where you set it.

Other vendors render on the client too. So what?

bitHuman does, and they charge $0.01 a minute self-hosted. We charge $0.03, or $0.08 with a realtime voice included. Others stream compact pose data and rasterise it in the browser — still a server generating motion for every frame of every session. We charge for conversation time, never for frames, and there is no concurrency tier because there is no capacity of ours to reserve. Full comparison.

Is it really zero network?

No. Rendering happens on the device, so no video or frames cross the network. With your own voice provider, audio and transcripts go from the device to that provider and never reach Yoob servers. With Yoob voice, audio goes through our relay at voice.yoob.com to OpenAI. Character files download once, with a short-lived grant, and heartbeats carry only a session token. If Yoob is down, a running session keeps rendering for up to 10 minutes; new sessions need our API. Everything the SDK sends.

Give it a face that’s actually there.