initial release
This commit is contained in:
@@ -0,0 +1,900 @@
|
||||
# Architecture
|
||||
|
||||
Technical reference for omnivoice.cpp, the GGML port of OmniVoice
|
||||
(k2-fsa/OmniVoice). This document covers the model, the conversion to
|
||||
GGUF, the inference pipeline, the GGML graph conventions, and the CLI
|
||||
tools.
|
||||
|
||||
## Upstream model
|
||||
|
||||
OmniVoice (Xiaomi / k2-fsa, Apache 2.0) is a multilingual zero-shot
|
||||
text-to-speech system covering 646 languages. It targets three modes :
|
||||
|
||||
voice cloning reference audio plus reference transcript drive the
|
||||
target speaker identity
|
||||
voice design six attribute categories (gender, age, pitch, style,
|
||||
volume, emotion) drive a synthesised speaker
|
||||
auto voice no reference, the model picks a coherent speaker per
|
||||
utterance
|
||||
|
||||
The system is a non autoregressive, mask-predict (MaskGIT) generative
|
||||
model running on top of a Qwen3 backbone with custom audio
|
||||
input/output and a separate audio tokenizer. The audio tokenizer is
|
||||
the Higgs Audio v2 codec (`bosonai/higgs-audio-v2-tokenizer`,
|
||||
Apache 2.0), which combines a HuBERT semantic stream, a DAC acoustic
|
||||
stream, and an 8-codebook residual vector quantiser at 25 frames per
|
||||
second over 24 kHz mono audio.
|
||||
|
||||
Single public checkpoint : `k2-fsa/OmniVoice` (3.1 GB).
|
||||
|
||||
Backbone Qwen3 0.6B (28 layers, hidden 1024, GQA 16/8)
|
||||
Audio codebooks 8 residual, 1024 entries each plus 1 mask token
|
||||
Audio framerate 25 Hz
|
||||
Hop length 960 samples
|
||||
Sample rate 24 kHz mono
|
||||
Semantic SR 16 kHz (HuBERT input)
|
||||
MaskGIT steps 32 default, configurable
|
||||
|
||||
## Build
|
||||
|
||||
```
|
||||
git clone --recurse-submodules https://github.com/ServeurpersoCom/omnivoice.cpp.git
|
||||
cd omnivoice.cpp
|
||||
./buildcuda.sh # NVIDIA GPU
|
||||
./buildvulkan.sh # AMD/Intel GPU (Vulkan)
|
||||
./buildcpu.sh # CPU only
|
||||
./buildall.sh # all backends, runtime DL loading
|
||||
```
|
||||
|
||||
The GGML submodule lives at `https://github.com/ServeurpersoCom/ggml.git`
|
||||
and provides two custom ops required by the codec :
|
||||
`GGML_OP_SNAKE` and `GGML_OP_COL2IM_1D`. Both have CPU, CUDA, Metal,
|
||||
and Vulkan kernels.
|
||||
|
||||
## Model conversion
|
||||
|
||||
```
|
||||
./checkpoints.sh # hf download k2-fsa/OmniVoice -> checkpoints/OmniVoice/
|
||||
./convert.py # 2 GGUFs in BF16 -> models/
|
||||
./quantize.sh # base LM Q8_0 (tokenizer stays at native dtype)
|
||||
```
|
||||
|
||||
Outputs :
|
||||
|
||||
```
|
||||
models/omnivoice-base-BF16.gguf 1.2 GB LLM + audio_emb + audio_heads + tokenizer
|
||||
models/omnivoice-base-Q8_0.gguf 626 MB quantized base, 1.9x smaller
|
||||
models/omnivoice-tokenizer-F32.gguf 702 MB HuBERT + DAC + RVQ + fc/fc2 (native F32)
|
||||
```
|
||||
|
||||
The audio tokenizer GGUF preserves the source dtype 1:1. The reference
|
||||
checkpoint stores the codec at F32, so the GGUF stays F32 to avoid
|
||||
truncation noise across the 8-stage RVQ residual chain. Late codebooks
|
||||
fall below 50 percent codebook match against the reference if any
|
||||
intermediate weight is rounded to BF16.
|
||||
|
||||
Quantisation policy : Q8_0 only on the base LM. The 612 M parameter
|
||||
backbone is small enough that lower quants degrade quality without
|
||||
meaningful size gains.
|
||||
|
||||
## GGUF layout
|
||||
|
||||
`omnivoice-base-{quant}.gguf` (arch `omnivoice-lm`) :
|
||||
|
||||
```
|
||||
metadata
|
||||
general.architecture omnivoice-lm
|
||||
block_count 28
|
||||
embedding_length 1024
|
||||
feed_forward_length 3072
|
||||
head_count 16
|
||||
head_count_kv 8 (GQA 2:1)
|
||||
key_length 128
|
||||
vocab_size 151676
|
||||
context_length 40960
|
||||
layer_norm_rms_eps 1e-6
|
||||
rope_freq_base 1e6
|
||||
omnivoice.tie_word_embeddings true
|
||||
omnivoice.num_audio_codebook 8
|
||||
omnivoice.audio_vocab_size 1025
|
||||
omnivoice.audio_mask_id 1024
|
||||
omnivoice.audio_codebook_weights [8, 8, 6, 6, 4, 4, 2, 2]
|
||||
omnivoice.special.denoise 151669
|
||||
omnivoice.special.lang_start 151670
|
||||
omnivoice.special.lang_end 151671
|
||||
omnivoice.special.instruct_start 151672
|
||||
omnivoice.special.instruct_end 151673
|
||||
omnivoice.special.text_start 151674
|
||||
omnivoice.special.text_end 151675
|
||||
tokenizer (Qwen2 BPE, 151676 vocab, 151387 merges, 33 added_tokens)
|
||||
|
||||
tensors (312)
|
||||
llm.embed_tokens.weight (151676, 1024)
|
||||
llm.norm.weight (1024,)
|
||||
llm.layers.0..27.{q,k,v,o}_proj.weight GQA, no bias
|
||||
llm.layers.0..27.self_attn.{q_norm, k_norm}.weight per-head RMSNorm (128,)
|
||||
llm.layers.0..27.{input,post_attention}_layernorm.weight RMSNorm
|
||||
llm.layers.0..27.mlp.{gate,up,down}_proj.weight SwiGLU, no bias
|
||||
audio_embeddings.weight (8200, 1024) 8 codebooks * 1025 vocab
|
||||
audio_heads.weight (8200, 1024) audio output, no bias
|
||||
```
|
||||
|
||||
`omnivoice-tokenizer-{quant}.gguf` (arch `omnivoice-tokenizer`) :
|
||||
|
||||
```
|
||||
metadata
|
||||
omnivoice.sample_rate 24000
|
||||
omnivoice.semantic_sample_rate 16000
|
||||
omnivoice.downsample_factor 320
|
||||
omnivoice.codebook_size 1024
|
||||
omnivoice.codebook_dim 64
|
||||
omnivoice.acoustic.encoder_hidden_size 64
|
||||
omnivoice.acoustic.decoder_hidden_size 1024
|
||||
omnivoice.acoustic.hidden_size 256
|
||||
omnivoice.acoustic.n_codebooks 9 (only 8 used)
|
||||
omnivoice.acoustic.hop_length 960
|
||||
omnivoice.acoustic.upsampling_ratios [8, 5, 4, 2, 3]
|
||||
omnivoice.acoustic.downsampling_ratios [8, 5, 4, 2, 3]
|
||||
omnivoice.semantic.hidden_size 768 (HuBERT base)
|
||||
omnivoice.semantic.intermediate_size 3072
|
||||
omnivoice.semantic.num_attention_heads 12
|
||||
omnivoice.semantic.num_hidden_layers 12
|
||||
omnivoice.semantic.num_feat_extract_layers 7
|
||||
omnivoice.semantic.conv_dim [512]*7
|
||||
omnivoice.semantic.conv_kernel [10, 3, 3, 3, 3, 2, 2]
|
||||
omnivoice.semantic.conv_stride [5, 2, 2, 2, 2, 2, 2]
|
||||
omnivoice.semantic.num_conv_pos_embeddings 128
|
||||
omnivoice.semantic.num_conv_pos_embedding_groups 16
|
||||
omnivoice.semantic.layer_norm_eps 1e-5
|
||||
|
||||
tensors (486)
|
||||
acoustic_encoder.* DAC encoder, 5 blocks, downsamples 8 5 4 2 3
|
||||
acoustic_decoder.* DAC decoder, 5 blocks, upsamples 8 5 4 2 3
|
||||
encoder_semantic.* semantic conv blocks
|
||||
semantic_model.* HuBERT base, weight_norm folded
|
||||
quantizer.quantizers.0..7.{codebook.embed, project_in.{w,b}, project_out.{w,b}}
|
||||
fc.{weight, bias} 1024 -> 1024 (after concat acoustic + semantic)
|
||||
fc2.{weight, bias} 1024 -> 256 (before DAC decoder)
|
||||
```
|
||||
|
||||
Single weight_norm fold at convert time :
|
||||
`semantic_model.encoder.pos_conv_embed.conv.weight`, formula
|
||||
`weight = v * g / ||v||_{dim=(0,1)}` matching
|
||||
`torch._weight_norm(v, g, dim=2)`. Validated bit-perfect, max abs diff
|
||||
3.9e-7 against the PyTorch reference.
|
||||
|
||||
## Component architecture
|
||||
|
||||
### Qwen3 0.6B backbone with custom IO
|
||||
|
||||
Standard Qwen3 modulo two changes :
|
||||
|
||||
input embed hybrid text plus audio, weighted sum across 8
|
||||
codebooks gated by `audio_mask`
|
||||
output head custom `audio_heads` Linear (8200, 1024), no text
|
||||
`lm_head`
|
||||
|
||||
```
|
||||
input_ids [B, 8, S] int (text on row 0, audio codes on rows 1..7)
|
||||
audio_mask [B, S] bool
|
||||
|
||||
text_emb = embed_tokens(input_ids[:, 0, :]) (B, S, 1024)
|
||||
shifted = input_ids * audio_mask + offsets[None, :, None]
|
||||
offsets = arange(8) * 1025
|
||||
audio_emb = audio_embeddings(shifted).sum(dim=1) (B, S, 1024)
|
||||
inputs = where(audio_mask, audio_emb, text_emb) (B, S, 1024)
|
||||
|
||||
x = qwen3_forward(inputs, attention_mask, position_ids) (B, S, 1024)
|
||||
|
||||
logits_flat = x @ audio_heads.weight.T (B, S, 8200)
|
||||
logits = reshape (B, 8, S, 1025)
|
||||
```
|
||||
|
||||
Qwen3 specifics already in llama.cpp :
|
||||
|
||||
28 layers, hidden 1024, intermediate 3072
|
||||
16 query heads + 8 KV heads (GQA 2:1), head_dim 128
|
||||
per-head RMSNorm on Q and K (q_norm, k_norm shape (128,)) before RoPE
|
||||
no bias on Q/K/V/O/MLP
|
||||
RoPE theta = 1e6
|
||||
SwiGLU MLP
|
||||
tie_word_embeddings = true (`lm_head` absent, output goes through audio_heads)
|
||||
|
||||
### MaskGIT decoder
|
||||
|
||||
Iterative non autoregressive decoder, no KV cache. Each step is a full
|
||||
prefill of the LLM on the current input.
|
||||
|
||||
Prompt (per item, broadcast across 8 codebooks) :
|
||||
|
||||
```
|
||||
[<|denoise|>]?
|
||||
<|lang_start|> {iso_code or "None"} <|lang_end|>
|
||||
<|instruct_start|> {style or "None"} <|instruct_end|>
|
||||
<|text_start|> {ref_text + " " + text} <|text_end|>
|
||||
{ref_audio_codes}?
|
||||
{MASK x num_target_tokens}
|
||||
```
|
||||
|
||||
Unconditional prompt for CFG = the trailing `num_target_tokens` mask
|
||||
tokens only. Batched (cond + uncond) doubles the batch dim.
|
||||
|
||||
```
|
||||
for step in 0..num_step-1 :
|
||||
forward(input_ids, audio_mask, attention_mask) (2B, 8, S, 1025)
|
||||
log_probs = log_softmax(c + cfg_scale * (c - u))
|
||||
log_probs[..., MASK_ID] = -inf
|
||||
if class_temp > 0 :
|
||||
keep_top_k_ratio(log_probs, 0.1)
|
||||
gumbel_sample(temp = class_temp)
|
||||
pred = argmax(log_probs)
|
||||
score = log_probs.max - layer_idx * layer_penalty (5.0)
|
||||
if pos_temp > 0 :
|
||||
score += gumbel * pos_temp
|
||||
score[already_unmasked] = -inf
|
||||
topk_idx = topk(score.flatten(), schedule[step])
|
||||
tokens[topk_idx] = pred[topk_idx]
|
||||
update batch_input_ids cond and uncond
|
||||
```
|
||||
|
||||
Schedule of newly unmasked positions per step is computed from
|
||||
`_get_time_steps(t_start=0, t_end=1, num_step, t_shift=0.1)` then
|
||||
`ceil(N_total * (t[step+1] - t[step]))`. 32 steps default.
|
||||
|
||||
KV cache is not usable across MaskGIT steps. The attention is fully
|
||||
bidirectional, so the prefix hidden states depend on the current
|
||||
target state through every layer. As tokens get progressively
|
||||
unmasked the K and V tensors of the prefix at every layer above the
|
||||
embeddings drift, which forbids the standard prefix-cache trick that
|
||||
works for causal LMs. Each step is therefore a full prefill of the
|
||||
LLM at cost `2 * forward_full(B, S)` (the 2 accounts for the cond +
|
||||
uncond CFG rows).
|
||||
|
||||
Determinism. With `class_temperature = 0` and `position_temperature = 0`
|
||||
the decoder is bit deterministic. Higher temperatures rely on a
|
||||
seedable Philox4x32-10 PRNG. The pipeline threads the Philox counter
|
||||
across MaskGIT calls so that chunked inference matches the global RNG
|
||||
drift of the PyTorch reference.
|
||||
|
||||
#### Inner-loop optimisations
|
||||
|
||||
The num_step iterations of one chunk run on a fixed shape (`S`, `K`, `B'`) so
|
||||
the per-step overhead can be cut without touching the math.
|
||||
|
||||
`pipeline_tts_llm_forward_batched` accepts a `T_audio` parameter that
|
||||
narrows the GPU output to the audio window only. The MaskGIT decoder
|
||||
reads cond logits at `[S - T, S)` on row 0 and uncond logits at
|
||||
`[0, T)` on row 1, so the full `[B', V, K, S]` tensor is wasteful.
|
||||
With `T_audio > 0` the function builds two `ggml_view_4d` over those
|
||||
ranges, makes them contiguous via `ggml_cont`, and only those two
|
||||
sub-tensors are flagged as graph outputs. The GPU to CPU copy shrinks
|
||||
from `B' * V * K * S` floats to `2 * V * K * T_audio` floats, around
|
||||
5.6x less on the typical voice cloning shape (S ~ 1880, T ~ 350).
|
||||
When `T_audio == 0` the function falls back to the full output, used
|
||||
by the debug dump path that needs every position.
|
||||
|
||||
`MaskgitBatchedCtx` holds the inputs that stay constant across the
|
||||
steps : the F32 audio mask and its complement, the RoPE position
|
||||
vector, and the F16 attention bias. The bias is the heaviest piece,
|
||||
`B' * S * S` F16 conversions per step (about 7 M ops on the typical
|
||||
shape). `pipeline_tts_llm_batched_ctx_init` precomputes the bias once
|
||||
per chunk, with a single `ggml_fp32_to_fp16` call for each of the two
|
||||
distinct values (1.0 and 0.0) hoisted out of the conversion loop. The
|
||||
context also keeps the original int32 pointers so the debug loop path
|
||||
can hand them down to the single forward unchanged.
|
||||
|
||||
Both optimisations preserve the math exactly, the only side effect is
|
||||
a slight reordering of the GPU FP32 reductions when the scheduler
|
||||
fuses the new output nodes differently, which moves the audio cosine
|
||||
similarity by a few times 1e-6. Token-level results stay 100 percent
|
||||
exact against the PyTorch reference.
|
||||
|
||||
### Audio tokenizer pipeline
|
||||
|
||||
Encode (voice cloning reference path) :
|
||||
|
||||
```
|
||||
ref_audio @ 24 kHz (1, 1, T_samples)
|
||||
-> resample 16 kHz (kaiser polyphase)
|
||||
-> pad 160 each side
|
||||
-> HuBERT.feature_extractor (320x downsample)
|
||||
-> HuBERT.feature_projection (LayerNorm + Linear)
|
||||
-> + pos_conv_embed (folded)
|
||||
-> 12 transformer layers
|
||||
-> mean over 13 hidden states (1, 768, T_sem)
|
||||
-> downsample by 2 (semantic_downsample_factor)
|
||||
-> SemanticEncoder (conv blocks) (1, 768, T_frames)
|
||||
-> e_acoustic (DAC encoder, 5 down-blocks) (1, 256, T_frames)
|
||||
-> concat dim=1 (1, 1024, T_frames)
|
||||
-> fc Linear (1024 -> 1024) (1, 1024, T_frames)
|
||||
-> RVQ encode (8 codebooks residual) (1, 8, T_frames) int @ 25 fps
|
||||
```
|
||||
|
||||
Decode (TTS path) :
|
||||
|
||||
```
|
||||
codes [B, 8, T] int
|
||||
-> RVQ decode :
|
||||
for k in 0..7 :
|
||||
e_k = codebook[k].embed[codes[k, :]] (B, T, 64)
|
||||
p_k = e_k @ project_out[k].W.T + bias[k] (B, T, 1024)
|
||||
out += p_k
|
||||
-> transpose (B, 1024, T)
|
||||
-> fc2 Linear (1024 -> 256) (B, 256, T)
|
||||
-> acoustic_decoder DAC :
|
||||
conv1 (256 -> 1024, k=7, pad=3)
|
||||
for block in 0..4, ratios [8, 5, 4, 2, 3] :
|
||||
snake1 (alpha)
|
||||
conv_t1 (IC -> OC, k=2*r, stride=r,
|
||||
padding=ceil(r/2), output_padding=r%2)
|
||||
for res_unit in 0..2, dilations [1, 3, 9] :
|
||||
snake1 (alpha)
|
||||
conv1 (OC, OC, k=7, dil=d, pad=3*d)
|
||||
snake2 (alpha)
|
||||
conv2 (OC, OC, k=1)
|
||||
residual add
|
||||
snake1 (alpha) final
|
||||
conv2 (32 -> 1, k=7, pad=3)
|
||||
-> audio (B, 1, 960*T)
|
||||
```
|
||||
|
||||
960x upsample = 8 * 5 * 4 * 2 * 3. T_in @ 25 fps -> T_out @ 24 kHz exact.
|
||||
|
||||
### DAC decoder block channels
|
||||
|
||||
```
|
||||
block 0 : IC=1024 OC=512 stride=8 K=16 pad=4 output_pad=0
|
||||
block 1 : IC=512 OC=256 stride=5 K=10 pad=3 output_pad=1
|
||||
block 2 : IC=256 OC=128 stride=4 K=8 pad=2 output_pad=0
|
||||
block 3 : IC=128 OC=64 stride=2 K=4 pad=1 output_pad=0
|
||||
block 4 : IC=64 OC=32 stride=3 K=6 pad=2 output_pad=1
|
||||
final : 32 -> 1
|
||||
```
|
||||
|
||||
PyTorch ConvTranspose1d formula :
|
||||
`T_out = (T_in - 1)*stride - 2*padding + dilation*(kernel - 1) + output_padding + 1`
|
||||
|
||||
With our parameters (d=1, k=2*s, p=ceil(s/2), op=s%2) the formula
|
||||
collapses to `T_out = stride * T_in` exactly for all five blocks.
|
||||
|
||||
### Snake activation
|
||||
|
||||
DAC reference formula (Hugging Face `Snake1d.forward`) :
|
||||
`y = x + (alpha + 1e-9).reciprocal() * sin(alpha * x)^2`
|
||||
|
||||
`ggml_snake(x, a, inv_b)` computes `y = x + sin^2(a * x) * inv_b`.
|
||||
Mapping :
|
||||
|
||||
a = alpha (loaded direct, BF16 to F32)
|
||||
inv_b = 1/(alpha + 1e-9) (precomputed CPU side at load, F32)
|
||||
|
||||
Both stored as F32 `[1, C]` tensors. `alpha` shape in checkpoint :
|
||||
`(1, C, 1)`, ggml ne = (1, C, 1). C lives on ne[1]. Loader reads C from
|
||||
`mt->ne[1]`.
|
||||
|
||||
### ConvTranspose1d via GEMM + col2im_1d
|
||||
|
||||
PyTorch `nn.ConvTranspose1d(IC, OC, kernel=K, stride=s, padding=p)` with
|
||||
weight shape `(IC, OC, K)`. GGML decomposition :
|
||||
|
||||
```
|
||||
1. Permute weight (IC, OC, K) PyTorch -> (IC, K*OC) ggml at load time.
|
||||
Layout : dst[(oc*K + k) * IC + ic] = src[ic*OC*K + oc*K + k]
|
||||
This makes k vary faster than oc inside the K*OC axis, matching
|
||||
what ggml_compute_forward_col2im_1d_impl expects :
|
||||
col_data[(oc * K + k) + t_in * K_OC]
|
||||
|
||||
2. Build runtime graph :
|
||||
xt = ggml_cont(ctx, ggml_transpose(ctx, x)) # [IC, T_in]
|
||||
col = ggml_mul_mat(ctx, w, xt) # [K*OC, T_in]
|
||||
y = ggml_col2im_1d(ctx, col, stride, OC, padding) # [T_no_op, OC]
|
||||
if (output_pad > 0)
|
||||
y = ggml_pad(ctx, y, output_pad, 0, 0, 0) # right-pad zeros
|
||||
if (bias)
|
||||
y = ggml_add(ctx, y, bias_2d)
|
||||
```
|
||||
|
||||
Validated math : `T_no_op = (T_in - 1)*stride + K - 2*pad`. Adding
|
||||
`output_pad` right-pad gives the PyTorch output size exactly.
|
||||
|
||||
### RVQ codec
|
||||
|
||||
Per-codebook tensors (k = 0..7) :
|
||||
|
||||
```
|
||||
codebook.embed (1024, 64) PyTorch -> ggml ne=(64, 1024)
|
||||
project_in.weight (64, 1024) PyTorch -> ggml ne=(1024, 64) encode-only
|
||||
project_in.bias (64,) encode-only
|
||||
project_out.weight (1024, 64) PyTorch -> ggml ne=(64, 1024)
|
||||
project_out.bias (1024,)
|
||||
```
|
||||
|
||||
Decode graph (per codebook k, accumulated) :
|
||||
|
||||
```
|
||||
codes_k = ggml_view_1d(codes, T, k * stride) # [T] i32
|
||||
e_k = ggml_get_rows(embed[k], codes_k) # [64, T]
|
||||
p_k = ggml_mul_mat(project_out_w[k], e_k) # [1024, T]
|
||||
p_k = ggml_add(p_k, project_out_b[k])
|
||||
acc += p_k
|
||||
```
|
||||
|
||||
Encode (residual loop) :
|
||||
|
||||
```
|
||||
residual = embeddings_in
|
||||
for k in 0..7 :
|
||||
e_k = project_in[k](residual)
|
||||
codes_k = argmin_i ||e_k - codebook[k].embed[i]||^2
|
||||
quantized = project_out[k](codebook[k].embed[codes_k])
|
||||
residual -= quantized
|
||||
```
|
||||
|
||||
### HuBERT semantic encoder
|
||||
|
||||
12 transformer layers Pre-LN, GELU FFN, MHA 12 heads * 64 dim, biases
|
||||
on all QKVO. Pre-conv feature extractor : 7 Conv1D layers, kernels
|
||||
`[10, 3, 3, 3, 3, 2, 2]`, strides `[5, 2, 2, 2, 2, 2, 2]`, GroupNorm on
|
||||
the first only, GELU between. Feature projection LayerNorm + Linear
|
||||
(512 -> 768). Positional embedding via grouped Conv1D (128 kernel,
|
||||
16 groups), `weight_norm` folded at convert time. Final LayerNorm.
|
||||
|
||||
Output computation :
|
||||
|
||||
```
|
||||
mean(stack(all_13_hidden_states, dim=1), dim=1) # (B, T_sem, 768)
|
||||
```
|
||||
|
||||
This is unusual : the encoder averages across the initial input plus
|
||||
the 12 transformer layer outputs, not just the last hidden state.
|
||||
|
||||
## Long-form TTS pipeline
|
||||
|
||||
`pipeline_tts_synthesize_long` orchestrates inputs longer than the
|
||||
chunking threshold. It mirrors `_generate_chunked` in
|
||||
`omnivoice/models/omnivoice.py`.
|
||||
|
||||
```
|
||||
1. Estimate total target tokens via duration_estimate_tokens.
|
||||
2. If T_total <= chunk_threshold_sec * frame_rate, run a single shot
|
||||
pipeline_tts_synthesize and skip chunking.
|
||||
3. Otherwise split text on punctuation with chunk_text_punctuation,
|
||||
targeting chunk_duration_sec seconds per chunk.
|
||||
4. Generate chunks sequentially :
|
||||
- chunk 0 with no reference (auto voice / voice design path) or
|
||||
with the external reference (cloning path)
|
||||
- in the auto voice case, the audio tokens of chunk 0 become the
|
||||
voice prompt for chunks 1..N, locking in the speaker identity
|
||||
5. Cross-fade decoded chunks with cross_fade_chunks(rate, 0.3 s).
|
||||
6. Apply post-processing on the merged waveform.
|
||||
```
|
||||
|
||||
A shared Philox counter `ctr_lo` is threaded across MaskGIT calls so
|
||||
PRNG state advances continuously between chunks, matching the global
|
||||
`torch.cuda.manual_seed` behaviour on the reference side.
|
||||
|
||||
### Text chunking
|
||||
|
||||
`chunk_text_punctuation(text, chunk_len, min_chunk_len)` splits text on
|
||||
sentence-ending punctuation (skipping abbreviation periods), then
|
||||
merges sentences into chunks of at most `chunk_len` UTF-8 codepoints.
|
||||
Undersized chunks (< `min_chunk_len`) are merged into a neighbour.
|
||||
The function operates on UTF-8 strings and treats length as codepoints,
|
||||
matching Python `len(str)` semantics. Per chunk character budget :
|
||||
|
||||
```
|
||||
n_chars = utf8_codepoint_count(full_text)
|
||||
avg_tokens_per_char = T_total / n_chars
|
||||
chunk_len = (int)(chunk_duration_sec * frame_rate / avg_tokens_per_char)
|
||||
```
|
||||
|
||||
`add_punctuation(text)` appends a terminal `.` (Latin) or its CJK
|
||||
equivalent when missing. Used on the reference transcript when
|
||||
`preprocess_prompt` is on.
|
||||
|
||||
### Audio post-processing
|
||||
|
||||
`audio-postproc.h` is a strict math port of `omnivoice/utils/audio.py`
|
||||
plus the relevant `pydub.silence` routines. All public functions take
|
||||
and return float32 mono PCM in [-1, 1] at the pipeline rate (24 kHz).
|
||||
Silence detection runs on int16 samples to match pydub bit-for-bit.
|
||||
|
||||
```
|
||||
remove_silence(buf, min_silence_ms, keep_silence_ms,
|
||||
seek_step_ms, threshold_dbfs)
|
||||
|
||||
Splits buf on contiguous silent regions where every
|
||||
seek_step_ms-long frame stays below threshold_dbfs (RMS, S16,
|
||||
default -50 dBFS), keeps keep_silence_ms of leading and trailing
|
||||
silence around each retained segment, and concatenates the result.
|
||||
|
||||
cross_fade_chunks(chunks, rate, fade_seconds)
|
||||
|
||||
Concatenates audio chunks with a linear cross-fade of fade_seconds
|
||||
at each junction.
|
||||
|
||||
peak_normalize_half(buf)
|
||||
|
||||
Scales buf so peak |x| equals 0.5. Used in pure auto voice when no
|
||||
reference RMS is available.
|
||||
|
||||
fade_and_pad(buf, rate, fade_seconds, pad_seconds)
|
||||
|
||||
Applies a linear fade in / fade out and pads silence at the start
|
||||
and end. Default fade 0.05 s, pad 0.05 s. Mirrors final reference
|
||||
post-step.
|
||||
```
|
||||
|
||||
`ref_rms` is plumbed end to end and decides the volume branch :
|
||||
|
||||
ref_rms < 0 pure auto voice, peak_normalize_half on the
|
||||
cross-faded waveform
|
||||
ref_rms < 0.1 quiet reference, rescale by ref_rms / 0.1
|
||||
otherwise no-op
|
||||
|
||||
When a reference WAV is provided, the CLI computes its RMS on the F32
|
||||
samples after optional silence trimming and passes it down. The same
|
||||
quantity is used on the PyTorch side.
|
||||
|
||||
### Voice modes
|
||||
|
||||
```
|
||||
auto voice no ref-wav. Chunk 0 generates with no reference,
|
||||
subsequent chunks reuse chunk 0 audio tokens as the
|
||||
voice prompt. peak_normalize_half on output.
|
||||
|
||||
voice design no ref-wav, --instruct provides one or more attribute
|
||||
markers (gender, age, pitch, style, volume, emotion)
|
||||
resolved by voice_design.h to the EN/ZH instruct
|
||||
string the reference uses. Chunking behaves like
|
||||
auto voice.
|
||||
|
||||
voice cloning --ref-wav and --ref-text provided. The reference is
|
||||
resampled to 16 kHz, run through the audio tokenizer
|
||||
encoder, and the resulting RVQ codes are reused as
|
||||
the voice prompt for every chunk. The reference RMS
|
||||
sets the target loudness.
|
||||
```
|
||||
|
||||
## Public API
|
||||
|
||||
Two layers, picked by use case.
|
||||
|
||||
### Top-level public ABI : src/omnivoice.h
|
||||
|
||||
Single-header, plain C99, linkage `extern "C"`. The opaque `ov_context`
|
||||
handle aggregates the GGML backend pair, the LM pipeline, the audio
|
||||
tokenizer codec, the BPE tokenizer and the voice-design vocabulary.
|
||||
One init, one free, one synthesize call covers the full TTS path.
|
||||
Same names, same struct layout, same calling convention from C, C++,
|
||||
Python ctypes, Rust bindgen, Go cgo or any other binding generator.
|
||||
|
||||
```c
|
||||
#include "omnivoice.h"
|
||||
|
||||
struct ov_init_params iparams;
|
||||
ov_init_default_params(&iparams);
|
||||
iparams.model_path = "models/omnivoice-base-Q8_0.gguf";
|
||||
iparams.codec_path = "models/omnivoice-tokenizer-F32.gguf";
|
||||
|
||||
struct ov_context * ov = ov_init(&iparams);
|
||||
|
||||
struct ov_tts_params params;
|
||||
ov_tts_default_params(¶ms);
|
||||
params.text = "Hello world.";
|
||||
params.lang = "English";
|
||||
|
||||
struct ov_audio audio = { 0 };
|
||||
enum ov_status rc = ov_synthesize(ov, ¶ms, &audio);
|
||||
if (rc == OV_STATUS_OK) {
|
||||
/* audio.samples is a malloc'd buffer of audio.n_samples floats
|
||||
at audio.sample_rate Hz, audio.channels = 1 (mono) */
|
||||
}
|
||||
ov_audio_free(&audio);
|
||||
ov_free(ov);
|
||||
```
|
||||
|
||||
Status codes :
|
||||
|
||||
```
|
||||
OV_STATUS_OK 0
|
||||
OV_STATUS_INVALID_PARAMS -1 (mutually exclusive ref inputs etc.)
|
||||
OV_STATUS_INSTRUCT_INVALID -2 (instruct rejected by VoiceDesign)
|
||||
OV_STATUS_GENERATE_FAILED -3 (any internal generate / decode fail)
|
||||
OV_STATUS_OOM -4 (output samples allocation failed)
|
||||
OV_STATUS_CANCELLED -5 (cancel callback returned true)
|
||||
```
|
||||
|
||||
`ov_tts_params` exposes `cancel` and `cancel_user_data`. The pipeline
|
||||
polls between chunks of long-form output, so cancel granularity is
|
||||
roughly `chunk_duration_sec` (15 s by default).
|
||||
|
||||
The MaskGIT sampler config is flattened directly into `ov_tts_params`
|
||||
as seven `mg_*` fields ; `ov_tts_default_params` initialises them to
|
||||
the reference defaults (`num_step=32, guidance_scale=2.0, t_shift=0.1,
|
||||
layer_penalty_factor=5.0, position_temperature=5.0,
|
||||
class_temperature=0.0, seed=42`).
|
||||
|
||||
`ov_version()` returns a static string of the form
|
||||
`"MAJOR.MINOR.PATCH (git-hash, date)"`. The macros `OV_VERSION_MAJOR`,
|
||||
`OV_VERSION_MINOR`, `OV_VERSION_PATCH` are also available at
|
||||
compile time for feature-detection.
|
||||
|
||||
### ABI guarantee
|
||||
|
||||
`tests/abi-c.c` is built on every build with
|
||||
`-std=c99 -Wall -Werror -pedantic`. It includes the public header,
|
||||
calls every entry through its early-return path, and is wired into
|
||||
the default build target. Any regression that breaks plain C
|
||||
consumability fails the main build, not an opt-in step.
|
||||
|
||||
The static library `libomnivoice-core.a` is the default build
|
||||
artefact. For binding consumers, configure with
|
||||
`-DOMNIVOICE_SHARED=ON` to add a `libomnivoice.so` (or `.dll` /
|
||||
`.dylib`) shared target that exports only the `ov_*` symbols ;
|
||||
every internal `pipeline_*` and `backend_*` stays hidden behind
|
||||
`-fvisibility=hidden`. Install rules follow `GNUInstallDirs`.
|
||||
|
||||
### Low-level API : src/pipeline-tts.h, src/pipeline-codec.h
|
||||
|
||||
Direct access to the LM forward (`pipeline_tts_llm_forward`,
|
||||
`pipeline_tts_llm_forward_batched`), the MaskGIT-only path
|
||||
(`pipeline_tts_generate`), the codec encode / decode
|
||||
(`pipeline_codec_encode`, `pipeline_codec_decode`), the instruct
|
||||
resolver (`pipeline_tts_resolve_instruct`) and the manual init / free
|
||||
(`pipeline_tts_load`, `pipeline_codec_load`, `backend_init`).
|
||||
|
||||
Used by `--llm-test` and `--maskgit-test` in `omnivoice-tts`, by
|
||||
`omnivoice-codec` for the standalone codec roundtrip, and by the
|
||||
Python cossim harness through dump files.
|
||||
|
||||
This layer is intentionally not part of the public ABI (C++ types in
|
||||
the signatures, no visibility export). It exists for the in-tree
|
||||
debug paths and stays available as long as the bundled CLI tools
|
||||
need it. The handle layer above is the recommended entry for
|
||||
everything else.
|
||||
|
||||
## CLI tools
|
||||
|
||||
### omnivoice-tts
|
||||
|
||||
End-to-end synthesis : text on stdin, WAV file on disk. Verbatim
|
||||
`--help` (the binary also prints an `omnivoice.cpp <hash> (<date>)`
|
||||
banner line first) :
|
||||
|
||||
```
|
||||
Usage: omnivoice-tts --model <gguf> --codec <gguf> [options] -o <out.wav> < text.txt
|
||||
|
||||
Required:
|
||||
--model <gguf> LLM GGUF (F32 / BF16 / Q8_0)
|
||||
--codec <gguf> Codec GGUF (omnivoice-tokenizer-*.gguf)
|
||||
-o <path> Output WAV (24 kHz mono). '-' streams to stdout (pipe friendly).
|
||||
|
||||
Input:
|
||||
stdin Target text to synthesise. With -o '-', stdin is read
|
||||
incrementally and synthesis starts as soon as the first
|
||||
sentence boundary is reached. With -o file.wav, stdin is
|
||||
read fully then synthesised in one shot.
|
||||
--srt <path> Dub an SRT: synth each cue into its time slot, write one
|
||||
timeline WAV ready to mux. Pairs with --ref-wav / --ref-rvq
|
||||
for a cloned voice. Per cue duration comes from the SRT.
|
||||
|
||||
Optional:
|
||||
--format <fmt> WAV output format: wav16, wav24, wav32 (default: wav16)
|
||||
--lang <str> Language label (default 'None')
|
||||
--instruct <str> Style instruction (default 'None')
|
||||
--duration <sec> Output duration in seconds (default: estimate from text)
|
||||
--no-denoise Omit the <|denoise|> prefix
|
||||
--ref-wav <path> Reference WAV for voice cloning
|
||||
--ref-text <path> Transcript file for the reference (required with --ref-wav / --ref-rvq)
|
||||
--ref-rvq <path> Pre-encoded reference codes from omnivoice-codec (replaces --ref-wav)
|
||||
--seed <int> Sampling seed (default: -1 for random)
|
||||
--steps <int> MaskGIT decode steps (default: 32, fewer is faster)
|
||||
--no-preprocess-prompt Skip ref-wav silence trim and ref-text terminal punctuation
|
||||
--chunk-duration <sec> Long-form chunk duration (default: 15.0, <= 0 disables chunking)
|
||||
--chunk-threshold <sec> Activate chunking above this estimated duration (default: 30.0)
|
||||
--stream-by-line Flush synthesis at each newline, one WAV header per line (-o '-')
|
||||
|
||||
Debug:
|
||||
--no-fa Disable flash attention
|
||||
--clamp-fp16 Clamp hidden states to FP16 range
|
||||
--dump <dir> Dump intermediate tensors (f32) to <dir>
|
||||
--llm-test <input.bin> Full LLM forward, dump audio_logits
|
||||
--maskgit-test Greedy MaskGIT decoder, dump audio_tokens [K, T]
|
||||
(no codec decode, reads target text from stdin)
|
||||
```
|
||||
|
||||
### omnivoice-codec
|
||||
|
||||
Audio tokenizer round-trip : WAV to RVQ codes, RVQ codes to WAV.
|
||||
Verbatim `--help` :
|
||||
|
||||
```
|
||||
Usage: omnivoice-codec --model <gguf> -i <input>
|
||||
|
||||
Required:
|
||||
--model <gguf> Codec GGUF (omnivoice-tokenizer-*.gguf)
|
||||
-i <path> Input. WAV -> encode, .rvq -> decode
|
||||
|
||||
Optional:
|
||||
--format <fmt> WAV output format: wav16, wav24, wav32 (default: wav16)
|
||||
|
||||
Output is auto-named next to input : clip.wav -> clip.rvq, clip.rvq -> clip.wav.
|
||||
Encode applies the TTS reference preprocessing (RMS auto-gain, silence trim,
|
||||
hop truncation); the resulting .rvq feeds omnivoice-tts --ref-rvq directly.
|
||||
```
|
||||
|
||||
The `.rvq` file is a small binary container with shape `[8, T]` int32
|
||||
codes plus a header carrying the sample rate and frame rate.
|
||||
|
||||
## Module map
|
||||
|
||||
```
|
||||
src/
|
||||
backend.h GGML backend init, scheduler factory, env override
|
||||
weight-ctx.h Generic weight context for GGUF loaders
|
||||
gguf-weights.h mmap GGUF, gf_load_tensor, gf_get_*
|
||||
audio-io.h WAV read, mono write (S16 / S24 / F32)
|
||||
audio-resample.h Kaiser polyphase 24 kHz <-> 16 kHz
|
||||
audio-postproc.h remove_silence, peak_normalize_half, fade_and_pad,
|
||||
cross_fade_chunks. Strict pydub / utils.audio port.
|
||||
wav.h WAV header reader (PCM16/24/F32, mono/stereo)
|
||||
philox.h Philox4x32-10 counter-based PRNG, PyTorch CUDA aligned
|
||||
debug.h Tensor dumper for cossim tests
|
||||
|
||||
bpe.h Qwen2 / GPT-2 byte-level BPE tokenizer, GGUF loader
|
||||
lang-map.h Language name to ISO 639-3 ID resolution
|
||||
(auto-generated from omnivoice/utils/lang_map.py)
|
||||
voice-design.h Speaker attribute validation and EN / ZH instruct
|
||||
resolution (mirrors voice_design.py)
|
||||
text-chunker.h chunk_text_punctuation, add_punctuation, END_PUNCTUATION
|
||||
duration-estimator.h RuleDurationEstimator port (per-script weights,
|
||||
Unicode category fallback)
|
||||
|
||||
rvq-codec.h Residual VQ encode + decode (8 codebooks)
|
||||
dac-decoder.h DAC acoustic decoder (5 blocks, ratios 8 5 4 2 3)
|
||||
dac-encoder.h DAC acoustic encoder (mirror of decoder)
|
||||
semantic-enc.h SemanticEncoder convs (768 -> 768)
|
||||
hubert-enc.h HuBERT base (feature extractor + pos_conv +
|
||||
12 transformer layers + final LN)
|
||||
|
||||
qwen3-enc.h Qwen3 transformer building blocks
|
||||
omnivoice-llm.h OmniVoice TTS LLM weights and graph helpers
|
||||
prompt-tts.h Prompt builder (denoise + lang + instruct + text +
|
||||
ref + mask) and CFG batch stacking
|
||||
maskgit-tts.h Iterative non autoregressive decoder, configurable
|
||||
step count (32 default), CFG, layer penalty, gumbel
|
||||
sampling, deterministic in greedy mode
|
||||
|
||||
pipeline-codec.{h,cpp} Audio tokenizer end-to-end (encode and decode)
|
||||
pipeline-tts.{h,cpp} Full TTS orchestration, single shot and chunked,
|
||||
plus low-level entries kept available for the
|
||||
debug paths and the cossim test harness
|
||||
omnivoice.{h,cpp} Public ABI : opaque ov_context handle, plain C99
|
||||
header in extern "C", consumable from C, C++,
|
||||
Python ctypes, Rust bindgen, Go cgo
|
||||
|
||||
tools/
|
||||
omnivoice-tts.cpp CLI : text to WAV (auto / design / clone)
|
||||
omnivoice-codec.cpp CLI : codes <-> WAV
|
||||
quantize.cpp GGUF requantizer
|
||||
version.cmake Embeds the git short hash into the binary
|
||||
|
||||
tests/
|
||||
debug-tts-cossim.py Byte-level comparison of every pipeline stage
|
||||
against the PyTorch reference, voice design path
|
||||
debug-clone-cossim.py Same, voice cloning path
|
||||
cross-decode.py Cross check : decode C++ tokens through PyTorch
|
||||
codec and vice versa
|
||||
prompt.txt Long-form English TTS sample
|
||||
ref-audio.wav Voice cloning reference clip
|
||||
ref-text.txt Transcript matching ref-audio.wav
|
||||
abi-c.c Plain C99 smoke test for the public ABI ; built
|
||||
with -Wall -Werror -pedantic on every build,
|
||||
locks in C consumability and symbol linkage
|
||||
```
|
||||
|
||||
## GGML conventions
|
||||
|
||||
### Tensor shape and layout
|
||||
|
||||
PyTorch shape `(out, in)` for a Linear weight stores as ggml
|
||||
`ne[0]=in, ne[1]=out`. The GGUF tensor-shape array is reversed, so
|
||||
reading `reversed(t.shape)` from gguf-py yields the PyTorch shape
|
||||
directly.
|
||||
|
||||
For PyTorch `Conv1d` weight `(OC, IC, K)`, ggml ne is `(K, IC, OC)`.
|
||||
The kernel axis is innermost (contiguous in memory).
|
||||
|
||||
For PyTorch `ConvTranspose1d` weight `(IC, OC, K)`, the convert-time
|
||||
permutation to ggml `(IC, K*OC)` rearranges
|
||||
`(oc*K + k) * IC + ic` so that `ggml_col2im_1d` receives the correct
|
||||
column matrix.
|
||||
|
||||
`ggml_mul_mat(A, B)` : with A.ne[0] = K (must match B.ne[0]),
|
||||
A.ne[1] = M, B.ne[1] = N, output has ne = (N, M). In PyTorch terms,
|
||||
A is `(M, K)`, B is `(N, K)`, output is `(M, N)`, which equals
|
||||
`A @ B^T`.
|
||||
|
||||
### Custom GGML ops
|
||||
|
||||
Provided by the `ServeurpersoCom/ggml` fork :
|
||||
|
||||
`ggml_snake(ctx, x, a, inv_b)` : `y = x + sin^2(a * x) * inv_b`.
|
||||
F32 / F16 / BF16 input/output. CPU + CUDA + Metal + Vulkan.
|
||||
|
||||
`ggml_col2im_1d(ctx, a, s0, oc, p0)` : scatter-add `[K*OC, T_in]`
|
||||
columns into `[T_out, OC]` signal where
|
||||
`T_out = (T_in - 1)*s0 + K - 2*p0`. Layout requires k to vary faster
|
||||
than oc inside the K*OC axis. F32 / F16 / BF16. All backends. The
|
||||
fork also folds the padding crop into this op via the `p0` parameter,
|
||||
removing a follow-up `ggml_view` for the typical ConvTranspose1d use
|
||||
case.
|
||||
|
||||
### Backend lifecycle
|
||||
|
||||
`backend_init("MOD")` then `backend_sched_new(bp, max_nodes)`. Backend
|
||||
handles are shared across modules in the same binary, refcounted. The
|
||||
GPU backend is the default, the CPU backend is kept as a scheduler
|
||||
fallback.
|
||||
|
||||
## Validation
|
||||
|
||||
The reference comparison harness is `tests/debug-tts-cossim.py` and
|
||||
`tests/debug-clone-cossim.py`. They run the same input through the
|
||||
PyTorch reference (with TF32 disabled, eager attention) and through the
|
||||
C++ binary, dump each pipeline stage to disk, and report cosine
|
||||
similarity per stage. Latest run, chunked path, English long-form
|
||||
prompt :
|
||||
|
||||
```
|
||||
TTS chunked: Logits cos=1.000000 max 3.5e-04
|
||||
Step0 pred_tokens 99.93% (2 FP flips)
|
||||
Tokens 1.000000 exact 100.00%
|
||||
Audio 0.999991
|
||||
|
||||
Clone chunked: Lf hidden cos=1.000000 max 1.9e-03
|
||||
Logits cos=1.000000
|
||||
Step0 pred_tokens 100.00%
|
||||
Tokens 1.000000 exact 100.00%
|
||||
Audio 0.999989
|
||||
```
|
||||
|
||||
The few Step0 token flips are argmax ties at the FP epsilon (~2e-5
|
||||
between top1 and top2 at those positions), inherent to the mixed cuBLAS
|
||||
vs GGML kernel arithmetic. They resorb over the 32 MaskGIT steps so
|
||||
the final tokens match bit for bit and decoded audio cosine is
|
||||
> 0.9999.
|
||||
|
||||
## Glossary
|
||||
|
||||
RVQ Residual Vector Quantisation. Stack of codebooks where
|
||||
each one quantises the residual from the previous
|
||||
codebook reconstruction.
|
||||
|
||||
DAC Descript Audio Codec. Convolutional encoder/decoder over
|
||||
residual VQ codes.
|
||||
|
||||
HuBERT Hidden-Unit BERT. Transformer encoder pretrained with
|
||||
masked acoustic unit prediction. Used here to extract
|
||||
semantic embeddings from raw audio.
|
||||
|
||||
Snake Periodic activation introduced in BigVGAN,
|
||||
`y = x + (1/alpha) * sin^2(alpha * x)`. Replaces
|
||||
LeakyReLU in the DAC encoder/decoder.
|
||||
|
||||
CFG Classifier-Free Guidance. The model is run twice
|
||||
(conditional and unconditional) and the outputs combined
|
||||
as `c + scale * (c - u)` to amplify the conditional
|
||||
signal.
|
||||
|
||||
MaskGIT Masked Generative Image Transformer (Chang et al.,
|
||||
arXiv:2202.04200). Iterative non autoregressive decoder
|
||||
where masked tokens are progressively unmasked over a
|
||||
fixed number of steps, prioritising high-confidence
|
||||
positions per step. Originally introduced for image
|
||||
generation, adapted here to audio codes.
|
||||
|
||||
Philox Counter-based PRNG used by PyTorch CUDA. Thread safe and
|
||||
skip-ahead friendly, well suited to deterministic
|
||||
chunked inference.
|
||||
Reference in New Issue
Block a user