Technology
A dubbing engine built for performance, not just speech.
Zero-shot voice cloning, an explicit emotion bank, accent control, isochrony scoring, live dubbing and consent-tracked provenance in one pipeline, with every knob exposed.
An explicit, controllable emotion bank.
Thirteen named emotions with a continuous intensity dial (mild to extreme), settable per line or inline mid-sentence. Sustained emotions can be layered, and source-audio emotion can be auto-detected and preserved through translation instead of flattened. Every setting is reproducible: dial it in once, regenerate the same performance every time.
Inline syntax: "[happy] So good to see you! [sad:strong] I just wish it were better news." Intensity levels are mild · moderate · strong · extreme, or any value 0–2.
Accent control with adjustable intensity.
Blend a speaker’s delivery toward a target accent with a 0–100% intensity slider while their voice identity is preserved. The blend operates on the speech-delivery conditioning only, never the timbre, so the speaker still sounds like themselves.
Accent centroids ship per language after passing a speaker-similarity guardrail (≥ 0.60 at half strength).
Zero-shot cloning, hardened for real references.
No per-voice training. References are assembled automatically from a speaker’s best segments, denoised adaptively, and quality-scored (duration, SNR, clipping) with warnings surfaced before you dub, so even a speaker with seconds of usable audio gets the best possible clone.
Transcript of the rythmoband demo dialog:
- ANNA: Dub take three, rolling. (German: Synchro, Take drei. Läuft.)
- MARCUS: Watch the sync bar. (German: Achte auf die Sync-Linie.)
- ANNA: Catch up, you're late! (German: Aufholen, du bist spät!)
- MARCUS: Now… hold it. (German: Jetzt… halten.)
- ANNA: Cut. Print it. (German: Schnitt. Kopieren.)
Timing that survives the edit.
Slot-aware synthesis picks a speaking rate so every line fills its original screen time, a decode-time rate controller steers token budgets, and DubScore grades the finished dub segment by segment. Weak lines are flagged and regenerable individually, without reprocessing the timeline.
Max drift 12 ms across the full runtime · 214 segments
Live dubbing with an honest latency meter.
A streaming ASR → translation → cloned-TTS cascade driven by one latency knob (conversational, balanced, broadcast). Session voices are pre-warmed and cached on the GPU, chunks synthesize over a priority lane, and the console shows the measured end-to-end lag next to the target, not an aspirational number.
Netflix (dialog-gated)
Target −27 LKFS ± 0.5 · ceiling −2 dBTP
Consent, watermarking and provenance built in.
Every clone requires a consent reference: dual-signed marketplace agreements, per-project actor consent, or an attested release. Output audio carries an inaudible watermark, C2PA manifests are signed, and a consent-vault report collates the full evidence chain for a studio’s compliance team.
Consent record
verifiedconsent_ref: esign_5f2a · signed 2026-05-14
Acoustic watermark
verifiedscheme: audioseal · inaudible · present ✓
C2PA manifest
verifiedsignature valid · asserts project · model · consent_ref
sealed into every export · verification API returns the full chain
Voice Artist Program
Artists are paid per use, and keep their rights.
Verified artists in our curated marketplace earn a royalty every time their voice is used in a production, with a per-use ledger they can audit. Consent is dual-signed; commercial rights never transfer.
See it on your own content.
Bring a scene; leave with a dubbed, QC-scored, provenance-sealed cut.