UGC & Creator Ads

What is the number-reading voice cloning technique and how does it work?

Last updated August 1, 2026

The number-reading voice cloning technique is the calibration step in Google Omni's avatar setup: you read only double-digit numbers to the camera — no sentences, no paragraphs — and the model infers your intonation, pauses, and syllable handling from that alone. In testing, the resulting voice replica came out stronger than the facial replica.

To run it, start Omni's avatar setup, and when the calibration session begins, read the double-digit numbers it shows you aloud to the camera. That number-reading session, plus the face calibration, is the entire input — the model builds both your visual avatar and your voice clone from it.

Why numbers work as the sample: double-digit numbers force a wide spread of phonemes, stress patterns, and micro-pauses without any semantic content, so the model gets a dense prosodic signal in a short session. Omni's voice model demonstrably infers intonation, pause behavior, and syllable handling from this alone — testers confirmed it never asked for a sentence or a read-aloud paragraph. As one tester put it after the session: the output was a decent replica of the face but an incredible replica of the voice — across 30+ test outputs, voice replication consistently outperformed facial replication.

Set expectations correctly: this is an Omni-specific calibration workflow, not an industry-standard voice cloning method. Most voice cloning systems train from any speech sample — recorded sentences, existing audio — and don't prescribe number reading. Omni's approach is notable precisely because the input is so minimal and non-semantic, and because Google Omni is currently the only AI video model offering native avatar generation, so the voice clone plugs directly into generated video rather than living as a separate audio tool.

One practical note for using the cloned voice in avatar clips: keep a single speaker in frame and plan speaking segments around the model's consistency window — research on Omni found 6–7 seconds is the ceiling where lip sync performs reliably before degradation sets in.

Watch some of these to see what works for you:

See the number-reading voice cloning session and avatar results in action

It was just double digit numbers that I was saying to the camera — it didn't ask me to say a sentence, read out a paragraph.

— invideo's creative team, from hands-on testing of Google Omni's avatar calibration

Share

More on UGC & Creator Ads