- ONNX build β contents, status, and findings
- STATUS: produces speech; quality not yet at parity with PyTorch
- Contents
- Usage
- Findings
- 0. Four defects between export and intelligible speech
- 0a. Detail: the missing text LM
- 1. The attention op is picked by
(execution_provider, io_dtype)β and one pick can't decode - 2. Passing the eval does NOT mean the model works
- 3.
acoustic_decodershows cos β 0.88β0.92 and still passes β deliberate - 4. Other notes
- Provenance
VibeVoice: A Frontier Open-Source Text-to-Speech Model
VibeVoice-Realtime is a lightweight realβtime text-to-speech model supporting streaming text input and robust long-form speech generation. It can be used to build realtime TTS services, narrate live data streams, and let different LLMs start speaking from their very first tokens (plug in your preferred model) long before a full answer is generated. It produces initial audible speech in ~300 ms (hardware dependent).
βΆοΈ Watch demo video (Launch your own realtime demo via the websocket example in Usage)
Although the model is primarily built for English, we found that it still exhibits a certain level of multilingual capabilityβand even performs reasonably well in some languages. We provide nine additional languages (German, French, Italian, Japanese, Korean, Dutch, Polish, Portuguese, and Spanish) for users to explore and share feedback.
The model uses an interleaved, windowed design: it incrementally encodes incoming text chunks while, in parallel, continuing diffusion-based acoustic latent generation from prior context. Unlike the full multi-speaker long-form variants, this streaming model removes the semantic tokenizer and relies solely on an efficient acoustic tokenizer operating at an ultra-low frame rate (7.5 Hz).
Key features:
- Parameter size: 0.5B (deployment-friendly)
- Realtime TTS (~300 ms first audible latency)
- Streaming text input
- Robust long-form speech generation
This realtime variant supports only a single speaker. For multi-speaker conversational speech generation, please use other VibeVoice models. The model is currently intended for English speech only; other languages may produce unpredictable results.
β‘οΈ Technical Report: VibeVoice Technical Report
β‘οΈ Project Page: microsoft/VibeVoice
β‘οΈ Code: microsoft/VibeVoice-Code
β‘οΈ App: anycoderapps/VibeVoice-Realtime-0.5B
Training Details
Transformer-based Large Language Model (LLM) integrated with specialized acoustic tokenizer and a diffusion-based decoding head.
- LLM: Qwen2.5-0.5B for this release.
- Tokenizers:
- Acoustic Tokenizer: Based on a Ο-VAE variant (proposed in LatentLM), with a mirror-symmetric encoder-decoder structure featuring 7 stages of modified Transformer blocks. Achieves 3200x downsampling from 24kHz input. Decoder component is ~340M parameters.
- Diffusion Head: Lightweight module (4 layers, ~40M parameters) conditioned on LLM hidden states. Predicts acoustic VAE features using a Denoising Diffusion Probabilistic Models (DDPM) process. Uses Classifier-Free Guidance (CFG) and DPM-Solver (and variants) during inference.
- Context Length: Trained with a curriculum increasing up to 8,192 tokens.
- Training Stages:
- Tokenizer Pre-training: Acoustic tokenizer is pre-trained.
- VibeVoice Training: Pre-trained tokenizer is frozen; only the LLM and diffusion head parameters are trained. A curriculum learning strategy is used for input sequence length (4k -> 8K). Text tokenizer not explicitly specified, but the LLM (Qwen2.5) typically uses its own. Audio is "tokenized" via the acoustic tokenizer.
Models
| Model | Context Length | Generation Length | Weight |
|---|---|---|---|
| VibeVoice-Realtime-0.5B | 8k | ~10 min | You are here. |
| VibeVoice-1.5B | 64K | ~90 min | HF link |
| VibeVoice-Large | 32K | ~45 min | HF link |
Results
The model achieves satisfactory performance on short-sentence benchmarks, while the model is more focused on longβform speech generation.
Zero-shot TTS performance on LibriSpeech test-clean set
| Model | WER (%) β | Speaker Similarity β |
|---|---|---|
| VALL-E 2 | 2.40 | 0.643 |
| Voicebox | 1.90 | 0.662 |
| MELLE | 2.10 | 0.625 |
| VibeVoice-Realtime-0.5B | 2.00 | 0.695 |
Zero-shot TTS performance on SEED test-en set
| Model | WER (%) β | Speaker Similarity β |
|---|---|---|
| MaskGCT | 2.62 | 0.714 |
| Seed-TTS | 2.25 | 0.762 |
| FireRedTTS | 3.82 | 0.460 |
| SparkTTS | 1.98 | 0.584 |
| CosyVoice2 | 2.57 | 0.652 |
| VibeVoice-Realtime-0.5B | 2.05 | 0.633 |
Installation and Usage
Please refer to GitHub README
Responsible Usage
Direct intended uses
The VibeVoice-Realtime model is limited to research purposes exploring real-time highly realistic audio generation detailed in the tech report.
Out-of-scope uses
Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in any other way that is prohibited by MIT License. Use to generate any text transcript. Furthermore, this release is not intended or licensed for any of the following scenarios:
- Voice impersonation without explicit, recorded consent, including but not limited to, cloning a real individual's voice for satire, advertising, ransom, socialβengineering, or authentication bypass.
- Disinformation or impersonation, including but not limited to, creating audio presented as genuine recordings of real people or events.
- Realβtime or lowβlatency voice conversion, including but not limited to, telephone or videoβconference "live deepβfake" applications.
- Any act to circumvent, disable, or otherwise interfere with any technical or procedural safeguards implemented in this release, including but not limited to security controls, watermarking and other transparency mechanisms. Any act of reverse engineering, modification, injection of unauthorized code, or exploitation of vulnerabilities for purposes beyond the intended scope of use.
- Unsupported language β the model is trained only on English data; outputs in other languages are unsupported and may be unintelligible or inappropriate.
- Generation of background ambience, Foley, or music β VibeVoice is speechβonly and cannot produce coherent nonβspeech audio such as music.
Risks and limitations
While efforts have been made to optimize it through various techniques, it may still produce outputs that are unexpected, biased, or inaccurate. VibeVoice may inherit any biases, errors, or omissions produced by its base model (specifically, Qwen2.5 0.5b in this release). Potential for Deepfakes and Disinformation: High-quality synthetic speech can be misused to create convincing fake audio content for impersonation, fraud, or spreading disinformation. Users must ensure transcripts are reliable, check content accuracy, and avoid using generated content in misleading ways. Users are expected to use the generated content and to deploy the models in a lawful manner, in full compliance with all applicable laws and regulations in the relevant jurisdictions. It is best practice to disclose the use of AI when sharing AI-generated content. English only: Transcripts in language other than English may result in unexpected audio outputs. Non-Speech Audio: The model focuses solely on speech synthesis and does not handle background noise, music, or other sound effects. Overlapping Speech: The current model does not explicitly model or generate overlapping speech segments in conversations. Code, formulas, and special symbols β The model does not currently support reading code, mathematical formulas, or uncommon symbols. Please preβprocess input text to remove or normalize such content to avoid unpredictable results.
Recommendations
We do not recommend using VibeVoice in commercial or real-world applications without further testing and development. If you use this model to generate speech, we recommend disclosing to the end user that they are listening to AI generated content. This model is intended for research and development purposes only. Please use responsibly.
To mitigate the risks of misuse, we have: Removed acoustic tokenizer to avoid users creating embedding on their own. Embedded an audible disclaimer (e.g. "This segment was generated by AI") automatically into every synthesized audio file. Added an imperceptible watermark to generated audio so third parties can verify VibeVoice provenance. Please see contact information at the end of this model card. Users are responsible for sourcing their datasets legally. This may include securing appropriate rights and/or anonymizing data prior to use with VibeVoice. Users are reminded to be mindful of data privacy concerns.
Contact
This project was conducted by members of Microsoft Research. We welcome feedback and collaboration from our audience. If you have suggestions, questions, or observe unexpected/offensive behavior in our technology, please contact us at VibeVoice@microsoft.com. If the team receives reports of undesired behavior or identifies issues independently, we will update this repository with appropriate mitigations.
ONNX build β contents, status, and findings
ONNX exports of VibeVoice-Realtime-0.5B decomposed into sub-parts (Olive + onnxruntime-genai),
across a 3 precisions Γ 2 devices matrix. Built 2026-08-04. Scripts in code/.
STATUS: produces speech; quality not yet at parity with PyTorch
cpu_fp32 / cpu_int4 generate intelligible speech that stops on its own. Getting there took
four separate fixes, all in the driver/export, none in the weights β see
Findings #0. Voice quality is
improved but still short of the PyTorch reference and is the open item.
cpu_fp16 still cannot decode at all (unrelated bug β Findings #1).
| target | attention op | eval.py |
real generation |
|---|---|---|---|
cpu_fp32 |
GroupQueryAttention | 5/5 PASS | works β intelligible, self-terminating |
cpu_int4 |
GroupQueryAttention | 5/5 PASS | works; int4 conditions the diffusion head, so prefer fp32 |
cpu_fp16 |
MultiHeadAttention | 5/5 PASS | crashes on decode step 1 (Findings #1) |
cuda_int4 |
GroupQueryAttention | 5/5 PASS | untested (needs a GPU) |
cuda_fp16 |
GroupQueryAttention | 5/5 PASS | untested (needs a GPU) |
cuda_fp32 |
Attention (packed) | 5/5 PASS | untested β non-GQA, verify first |
Every measurement quoted below is reproducible in
audio_analysis.ipynb (executed, plots included): waveforms, energy
profiles, F0 tracks, spectrograms, frame-seam ratios, babble spectra, the text-conditioning
progression, and a full acoustic-decoder characterisation (causality, 56-frame receptive field,
streaming-vs-batch equality, RTF).
Sections 1β8 need only numpy + matplotlib and read the shipped samples/:
uv run --with matplotlib jupyter lab. Section 9 re-measures the decoder live if onnxruntime and
cpu_fp32/ are present, and otherwise prints the recorded results β the real latents it needs are
shipped as samples/latents_*.npy (the sweep is meaningless on random latents, which drive this
decoder to near-silence).
Samples
All for "Hello there. This is the VibeVoice realtime model speaking from an ONNX build."
in samples/, kept as a record of each fix:
| file | what it shows |
|---|---|
v0_babble_no_textlm_{int4,fp32}.wav |
original defect: voice timbre, unintelligible babble |
v1_textlm_fix.wav |
after exporting the missing 4-layer text LM |
v2_prompt_neg_interleave.wav |
+ correct prompt, real CFG negative stream, windowed interleave |
v3_batchdecode_cfg{1.3,3.0}.wav |
+ single-pass decode (seams 14.11Γ β 3.16Γ); cfg 3.0 clearly better |
v4_eos_cfg3.0.wav |
current best β + learned stop; 6.13 s instead of a padded 10.67 s |
Contents
Each {device}_{precision}/ holds:
| file | role |
|---|---|
llm_decoder.onnx |
the 20-layer tts_language_model backbone, inputs_embeds β hidden_states |
text_lm.onnx |
the 4-layer language_model text encoder (final norm = Identity), carries its own embed_tokens |
type_embed.npy |
[2, 896] tts_input_types lookup β row 1 = text, row 0 = speech; added to every TTS-LM input |
eos_head.npz |
tts_eos_classifier (896β896β1) β the learned stop; run in numpy each frame |
embeddings.onnx |
input_ids β inputs_embeds Gather companion β always fp16 |
acoustic_connector.onnx |
acoustic latent β LLM hidden space |
diffusion_head.onnx |
DiT denoiser (hidden=896), conditioned on LLM hidden states |
acoustic_decoder.onnx |
latents β 24 kHz waveform (decoder-only; realtime ships no acoustic encoder) |
tokenizer*.json, chat_template.jinja, genai_config.json |
runtime assets |
| target | llm | text_lm | acoustic_decoder | diffusion_head | acoustic_connector | embeddings | total |
|---|---|---|---|---|---|---|---|
cpu_int4 |
182 MB | 748 MB | 1312 MB | 160 MB | 3 MB | 259 MB | 2.6 GB |
cpu_fp16 |
571 MB | 374 MB | 733 MB | 80 MB | 1 MB | 259 MB | 2.0 GB |
cpu_fp32 |
1139 MB | 748 MB | 1312 MB | 160 MB | 3 MB | 259 MB | 3.6 GB |
cuda_int4 |
164 MB | 748 MB | 1312 MB | 160 MB | 3 MB | 259 MB | 2.6 GB |
cuda_fp16 |
571 MB | 374 MB | 733 MB | 80 MB | 1 MB | 259 MB | 2.0 GB |
cuda_fp32 |
1139 MB | 748 MB | 1312 MB | 160 MB | 3 MB | 259 MB | 3.6 GB |
text_lm is fp32 even in int4 targets. --precision int4 drives the genai ModelBuilder
(MatMulNBits) and applies to llm_decoder only; every other component is exported by Olive,
which has no int4 path. The 4-layer text encoder therefore ships fp32 (748 MB) in int4 builds.
It must still live in the int4 directory β that is the directory the driver loads.
int4 is larger than fp16 overall β expected, not a packaging error. --precision sets the
LLM quantization only; the audio/diffusion blocks are conv-VAE/DiT with no
MatMulNBits-quantizable weights, so an int4 build exports them fp32 (exactly 2Γ their fp16
size), outweighing the LLM's ~3Γ shrink. embeddings.onnx.data is identical (259 MB) everywhere β
always written fp16 by design (an fp32 [vocab, hidden] table trips the protobuf 2 GB limit, and
the driver upcasts on read).
Usage
pip install -r code/requirements.txt
# --max-frames is only a hard cap: the exported tts_eos_classifier stops generation on its own.
# cfg 3.0 is the default and is audibly better than 1.3 on this build.
python code/inference_realtime.py --text "Hello there." --out out.wav cpu_fp32
# parity eval vs the PyTorch reference (needs the source checkpoint)
python code/eval.py --device cpu --precision int4 realtime
code/: inference_realtime.py (entry point) Β· common.py (ONNX helpers, OnnxLLM KV-cache
driver, DiffusionSampler, audio I/O) Β· optimize.py (builder + registry, imported by common.py)
Β· user_script.py (Olive loaders/io_configs) Β· eval.py Β· vibevoice/ (vendored upstream
subset, 543 KB, MIT β VIBEVOICE_LICENSE; required at inference because common.make_scheduler
imports vibevoice.schedule.dpm_solver. Never import vibevoice directly β its __init__ collides
with transformers' registration; the code uses an isolated-import shim).
Findings
0. Four defects between export and intelligible speech
All four were driver/export bugs, fixed in this order. Each is worth knowing separately because each one alone still left the audio sounding wrong, which made them easy to confuse.
| # | defect | symptom | fix |
|---|---|---|---|
| a | 4-layer text LM never exported | unintelligible babble | export text_lm.onnx + type_embed.npy |
| b | wrong prompt + fake CFG negative | speech-like but not human | strip()+"\n", no special tokens; real parallel negative stream seeded with <|image_pad|>; 5-token/6-frame interleave |
| c | per-frame acoustic decode | buzzy seam every 3200 samples | decode all latents in one pass (seam energy 14.11Γ β 3.16Γ) |
| d | tts_eos_classifier never exported |
~40% of the clip was post-speech drift | export eos_head.npz, stop at sigmoid > 0.5 |
On (d), the numbers for the 18-token sample: generation used to run to --max-frames (80 frames =
10.67 s) regardless of the text. The classifier fires at frame 46 (p=0.965) = 6.13 s, and v4 is
sample-identical to v3 up to that point β it purely removes the tail, which had decayed to
rms=0.0372, collapsed to near-silence (0.0003), then emitted an unrelated 0.0902 burst.
Note the text is consumed by frame ~24 while the model keeps speaking to frame 46 β that gap is by design, not drift: text is ingested in 5-token windows faster than it is rendered at 6 frames per window. That is exactly why the learned stop is needed instead of a frame budget derived from token count.
The remaining gap is voice quality, not structure or intelligibility.
0a. Detail: the missing text LM
Realtime has two language models. model.language_model (4 layers, 49 tensors, final norm
replaced by nn.Identity()) encodes the text; model.tts_language_model (20 layers) is the
TTS backbone. The reference pipeline
(vibevoice/modular/modeling_vibevoice_streaming_inference.py) is:
lm_hidden = language_model(text_embeds) # 4-layer text encoder
start = inputs_embeds.shape[1] - lm_hidden.shape[1]
inputs_embeds[:, start:, :] = lm_hidden # splice text hidden into the TAIL
inputs_embeds += model.tts_input_types(masks) # segment/type embedding
outputs = tts_language_model(inputs_embeds=...) # then the 20-layer backbone
logits = tts_eos_classifier(hidden[:, -1, :]) # natural stop
The original build exported only tts_language_model, omitting model.language_model (text
encoder), model.tts_input_types (type embedding) and tts_eos_classifier (EOS). The driver fed raw
text-token embeddings straight into the TTS backbone, which was never trained to consume them β so
the backbone, diffusion head and acoustic decoder all worked (hence voice-like timbre) while the text
conditioning never happened.
Measured on v0_babble_*.wav: the spectrum sits in the voice band (300β3000 Hz β 55%, ZCR
0.04β0.07 β not noise), but per-100 ms RMS std is only 0.012 with 1% silence on fp32 β a continuous
drone with none of the pause structure of real speech. int4 and fp32 failed identically, which was
the tell that this was structural, not a quantization/precision problem.
All three are now exported: text_lm.onnx, plus type_embed.npy and eos_head.npz (both tiny β
a 2Γ896 lookup and an 896β896β1 MLP β computed in numpy rather than given their own graphs).
The streaming processor also expects a cached_prompt (voice-prompt KV cache for both LMs), so a
faithful port needs that path too.
Fix required: export the 4-layer text LM + tts_input_types (and ideally tts_eos_classifier)
as additional components, and rewrite the driver to do the two-stage splice.
1. The attention op is picked by (execution_provider, io_dtype) β and one pick can't decode
genai's builders/base.py::make_attention_init() looks that pair up in is_gqa_supported(), then
is_packed_attn_supported(); whichever matches decides the op. It is not selectable. Verified
against the emitted graphs (onnx.load(..., load_external_data=False) + a Counter over op_type):
| target | io_dtype | op emitted | why |
|---|---|---|---|
cpu_int4 |
FLOAT (int4-on-cpu β fp32 I/O) | GroupQueryAttention | ("cpu",FLOAT) β GQA set |
cpu_fp32 |
FLOAT | GroupQueryAttention | ("cpu",FLOAT) β GQA set |
cpu_fp16 |
FLOAT16 | MultiHeadAttention | in neither set β fallback |
cuda_int4 / cuda_fp16 |
FLOAT16 | GroupQueryAttention | ("cuda",FLOAT16) β GQA set |
cuda_fp32 |
FLOAT | Attention (packed) | β GQA, β packed set |
Why cpu_fp16 fails. The model is GQA-shaped (num_attn_heads β num_kv_heads), so the MHA
path inserts a repeat_kv expansion. It prefills fine, then dies on the first single-token step:
FAIL : Non-zero status code returned while running Cast node.
Name:'InsertedPrecisionFreeCast_/model/layers.1/attn/v_proj/repeat_kv/Reshape_4/output_0'
Shape mismatch attempting to re-use buffer. {1,1,896} != {1,19,896}
β a buffer sized for the 19-token prefill, reused for the 1-token decode. Only GQA sets
past_present_share_buffer, so only GQA survives the prefillβdecode length change here.
cuda_fp32 (packed Attention) is also non-GQA and unverified.
2. Passing the eval does NOT mean the model works
All six targets pass 5/5 β including one that can't decode a token and two that emit babble. Not an
eval bug, a scope limit: the llm check is structural (it counts ops), stage B SKIPs the
autoregressive loop, and nothing compares generated audio to a reference. Always run the
inference driver and listen.
3. acoustic_decoder shows cos β 0.88β0.92 and still passes β deliberate
Fed random latents the decoder emits near-silence, so cosine is noise-dominated while max|Ξ| is
~1e-07 (fp16 builds) / ~3β4e-10 (fp32 & int4). The criterion is cos β₯ 0.99 OR max|Ξ| tiny. Do not
"fix" it by tightening the threshold.
4. Other notes
- 20 attention nodes in every target β the TTS backbone genuinely has 20 layers, not the 24
decoder_configadvertises (tts_backbone_num_hidden_layersis the real count). - fp16-on-CPU is buildable only because the builder calls genai
create_model()directly; the "FP16 is not supported on CPU" refusal lives in Olive's ModelBuilder pass, which that path bypasses. It builds β see Findings #1 for whether it runs. - CUDA graphs build on a CPU-only box.
create_model()emits a CUDA-targeted graph with no CUDA runtime present, and the Olive component passes are pure graph transforms. Building β running. - No learned EOS is exported (
tts_eos_classifier), so generation stops at--max-frames. - io_configs use
dynamic_shapes, notdynamic_axesβ required, since the dynamo exporter ignoresdynamic_axesand would bake the dummy's dimensions in as static.
Provenance
python code/optimize.py --device {cpu,cuda} --precision {int4,fp16,fp32} realtime
Model weights follow the upstream license above. Vendored vibevoice/ source is MIT β see
code/VIBEVOICE_LICENSE.