36 KiB
Architecture
Technical reference for omnivoice.cpp, the GGML port of OmniVoice (k2-fsa/OmniVoice). This document covers the model, the conversion to GGUF, the inference pipeline, the GGML graph conventions, and the CLI tools.
Upstream model
OmniVoice (Xiaomi / k2-fsa, Apache 2.0) is a multilingual zero-shot text-to-speech system covering 646 languages. It targets three modes :
voice cloning reference audio plus reference transcript drive the target speaker identity voice design six attribute categories (gender, age, pitch, style, volume, emotion) drive a synthesised speaker auto voice no reference, the model picks a coherent speaker per utterance
The system is a non autoregressive, mask-predict (MaskGIT) generative
model running on top of a Qwen3 backbone with custom audio
input/output and a separate audio tokenizer. The audio tokenizer is
the Higgs Audio v2 codec (bosonai/higgs-audio-v2-tokenizer,
Apache 2.0), which combines a HuBERT semantic stream, a DAC acoustic
stream, and an 8-codebook residual vector quantiser at 25 frames per
second over 24 kHz mono audio.
Single public checkpoint : k2-fsa/OmniVoice (3.1 GB).
Backbone Qwen3 0.6B (28 layers, hidden 1024, GQA 16/8) Audio codebooks 8 residual, 1024 entries each plus 1 mask token Audio framerate 25 Hz Hop length 960 samples Sample rate 24 kHz mono Semantic SR 16 kHz (HuBERT input) MaskGIT steps 32 default, configurable
Build
git clone --recurse-submodules https://github.com/ServeurpersoCom/omnivoice.cpp.git
cd omnivoice.cpp
./buildcuda.sh # NVIDIA GPU
./buildvulkan.sh # AMD/Intel GPU (Vulkan)
./buildcpu.sh # CPU only
./buildall.sh # all backends, runtime DL loading
The GGML submodule lives at https://github.com/ServeurpersoCom/ggml.git
and provides two custom ops required by the codec :
GGML_OP_SNAKE and GGML_OP_COL2IM_1D. Both have CPU, CUDA, Metal,
and Vulkan kernels.
Model conversion
./checkpoints.sh # hf download k2-fsa/OmniVoice -> checkpoints/OmniVoice/
./convert.py # 2 GGUFs in BF16 -> models/
./quantize.sh # base LM Q8_0 (tokenizer stays at native dtype)
Outputs :
models/omnivoice-base-BF16.gguf 1.2 GB LLM + audio_emb + audio_heads + tokenizer
models/omnivoice-base-Q8_0.gguf 626 MB quantized base, 1.9x smaller
models/omnivoice-tokenizer-F32.gguf 702 MB HuBERT + DAC + RVQ + fc/fc2 (native F32)
The audio tokenizer GGUF preserves the source dtype 1:1. The reference checkpoint stores the codec at F32, so the GGUF stays F32 to avoid truncation noise across the 8-stage RVQ residual chain. Late codebooks fall below 50 percent codebook match against the reference if any intermediate weight is rounded to BF16.
Quantisation policy : Q8_0 only on the base LM. The 612 M parameter backbone is small enough that lower quants degrade quality without meaningful size gains.
GGUF layout
omnivoice-base-{quant}.gguf (arch omnivoice-lm) :
metadata
general.architecture omnivoice-lm
block_count 28
embedding_length 1024
feed_forward_length 3072
head_count 16
head_count_kv 8 (GQA 2:1)
key_length 128
vocab_size 151676
context_length 40960
layer_norm_rms_eps 1e-6
rope_freq_base 1e6
omnivoice.tie_word_embeddings true
omnivoice.num_audio_codebook 8
omnivoice.audio_vocab_size 1025
omnivoice.audio_mask_id 1024
omnivoice.audio_codebook_weights [8, 8, 6, 6, 4, 4, 2, 2]
omnivoice.special.denoise 151669
omnivoice.special.lang_start 151670
omnivoice.special.lang_end 151671
omnivoice.special.instruct_start 151672
omnivoice.special.instruct_end 151673
omnivoice.special.text_start 151674
omnivoice.special.text_end 151675
tokenizer (Qwen2 BPE, 151676 vocab, 151387 merges, 33 added_tokens)
tensors (312)
llm.embed_tokens.weight (151676, 1024)
llm.norm.weight (1024,)
llm.layers.0..27.{q,k,v,o}_proj.weight GQA, no bias
llm.layers.0..27.self_attn.{q_norm, k_norm}.weight per-head RMSNorm (128,)
llm.layers.0..27.{input,post_attention}_layernorm.weight RMSNorm
llm.layers.0..27.mlp.{gate,up,down}_proj.weight SwiGLU, no bias
audio_embeddings.weight (8200, 1024) 8 codebooks * 1025 vocab
audio_heads.weight (8200, 1024) audio output, no bias
omnivoice-tokenizer-{quant}.gguf (arch omnivoice-tokenizer) :
metadata
omnivoice.sample_rate 24000
omnivoice.semantic_sample_rate 16000
omnivoice.downsample_factor 320
omnivoice.codebook_size 1024
omnivoice.codebook_dim 64
omnivoice.acoustic.encoder_hidden_size 64
omnivoice.acoustic.decoder_hidden_size 1024
omnivoice.acoustic.hidden_size 256
omnivoice.acoustic.n_codebooks 9 (only 8 used)
omnivoice.acoustic.hop_length 960
omnivoice.acoustic.upsampling_ratios [8, 5, 4, 2, 3]
omnivoice.acoustic.downsampling_ratios [8, 5, 4, 2, 3]
omnivoice.semantic.hidden_size 768 (HuBERT base)
omnivoice.semantic.intermediate_size 3072
omnivoice.semantic.num_attention_heads 12
omnivoice.semantic.num_hidden_layers 12
omnivoice.semantic.num_feat_extract_layers 7
omnivoice.semantic.conv_dim [512]*7
omnivoice.semantic.conv_kernel [10, 3, 3, 3, 3, 2, 2]
omnivoice.semantic.conv_stride [5, 2, 2, 2, 2, 2, 2]
omnivoice.semantic.num_conv_pos_embeddings 128
omnivoice.semantic.num_conv_pos_embedding_groups 16
omnivoice.semantic.layer_norm_eps 1e-5
tensors (486)
acoustic_encoder.* DAC encoder, 5 blocks, downsamples 8 5 4 2 3
acoustic_decoder.* DAC decoder, 5 blocks, upsamples 8 5 4 2 3
encoder_semantic.* semantic conv blocks
semantic_model.* HuBERT base, weight_norm folded
quantizer.quantizers.0..7.{codebook.embed, project_in.{w,b}, project_out.{w,b}}
fc.{weight, bias} 1024 -> 1024 (after concat acoustic + semantic)
fc2.{weight, bias} 1024 -> 256 (before DAC decoder)
Single weight_norm fold at convert time :
semantic_model.encoder.pos_conv_embed.conv.weight, formula
weight = v * g / ||v||_{dim=(0,1)} matching
torch._weight_norm(v, g, dim=2). Validated bit-perfect, max abs diff
3.9e-7 against the PyTorch reference.
Component architecture
Qwen3 0.6B backbone with custom IO
Standard Qwen3 modulo two changes :
input embed hybrid text plus audio, weighted sum across 8
codebooks gated by audio_mask
output head custom audio_heads Linear (8200, 1024), no text
lm_head
input_ids [B, 8, S] int (text on row 0, audio codes on rows 1..7)
audio_mask [B, S] bool
text_emb = embed_tokens(input_ids[:, 0, :]) (B, S, 1024)
shifted = input_ids * audio_mask + offsets[None, :, None]
offsets = arange(8) * 1025
audio_emb = audio_embeddings(shifted).sum(dim=1) (B, S, 1024)
inputs = where(audio_mask, audio_emb, text_emb) (B, S, 1024)
x = qwen3_forward(inputs, attention_mask, position_ids) (B, S, 1024)
logits_flat = x @ audio_heads.weight.T (B, S, 8200)
logits = reshape (B, 8, S, 1025)
Qwen3 specifics already in llama.cpp :
28 layers, hidden 1024, intermediate 3072
16 query heads + 8 KV heads (GQA 2:1), head_dim 128
per-head RMSNorm on Q and K (q_norm, k_norm shape (128,)) before RoPE
no bias on Q/K/V/O/MLP
RoPE theta = 1e6
SwiGLU MLP
tie_word_embeddings = true (lm_head absent, output goes through audio_heads)
MaskGIT decoder
Iterative non autoregressive decoder, no KV cache. Each step is a full prefill of the LLM on the current input.
Prompt (per item, broadcast across 8 codebooks) :
[<|denoise|>]?
<|lang_start|> {iso_code or "None"} <|lang_end|>
<|instruct_start|> {style or "None"} <|instruct_end|>
<|text_start|> {ref_text + " " + text} <|text_end|>
{ref_audio_codes}?
{MASK x num_target_tokens}
Unconditional prompt for CFG = the trailing num_target_tokens mask
tokens only. Batched (cond + uncond) doubles the batch dim.
for step in 0..num_step-1 :
forward(input_ids, audio_mask, attention_mask) (2B, 8, S, 1025)
log_probs = log_softmax(c + cfg_scale * (c - u))
log_probs[..., MASK_ID] = -inf
if class_temp > 0 :
keep_top_k_ratio(log_probs, 0.1)
gumbel_sample(temp = class_temp)
pred = argmax(log_probs)
score = log_probs.max - layer_idx * layer_penalty (5.0)
if pos_temp > 0 :
score += gumbel * pos_temp
score[already_unmasked] = -inf
topk_idx = topk(score.flatten(), schedule[step])
tokens[topk_idx] = pred[topk_idx]
update batch_input_ids cond and uncond
Schedule of newly unmasked positions per step is computed from
_get_time_steps(t_start=0, t_end=1, num_step, t_shift=0.1) then
ceil(N_total * (t[step+1] - t[step])). 32 steps default.
KV cache is not usable across MaskGIT steps. The attention is fully
bidirectional, so the prefix hidden states depend on the current
target state through every layer. As tokens get progressively
unmasked the K and V tensors of the prefix at every layer above the
embeddings drift, which forbids the standard prefix-cache trick that
works for causal LMs. Each step is therefore a full prefill of the
LLM at cost 2 * forward_full(B, S) (the 2 accounts for the cond +
uncond CFG rows).
Determinism. With class_temperature = 0 and position_temperature = 0
the decoder is bit deterministic. Higher temperatures rely on a
seedable Philox4x32-10 PRNG. The pipeline threads the Philox counter
across MaskGIT calls so that chunked inference matches the global RNG
drift of the PyTorch reference.
Inner-loop optimisations
The num_step iterations of one chunk run on a fixed shape (S, K, B') so
the per-step overhead can be cut without touching the math.
pipeline_tts_llm_forward_batched accepts a T_audio parameter that
narrows the GPU output to the audio window only. The MaskGIT decoder
reads cond logits at [S - T, S) on row 0 and uncond logits at
[0, T) on row 1, so the full [B', V, K, S] tensor is wasteful.
With T_audio > 0 the function builds two ggml_view_4d over those
ranges, makes them contiguous via ggml_cont, and only those two
sub-tensors are flagged as graph outputs. The GPU to CPU copy shrinks
from B' * V * K * S floats to 2 * V * K * T_audio floats, around
5.6x less on the typical voice cloning shape (S ~ 1880, T ~ 350).
When T_audio == 0 the function falls back to the full output, used
by the debug dump path that needs every position.
MaskgitBatchedCtx holds the inputs that stay constant across the
steps : the F32 audio mask and its complement, the RoPE position
vector, and the F16 attention bias. The bias is the heaviest piece,
B' * S * S F16 conversions per step (about 7 M ops on the typical
shape). pipeline_tts_llm_batched_ctx_init precomputes the bias once
per chunk, with a single ggml_fp32_to_fp16 call for each of the two
distinct values (1.0 and 0.0) hoisted out of the conversion loop. The
context also keeps the original int32 pointers so the debug loop path
can hand them down to the single forward unchanged.
Both optimisations preserve the math exactly, the only side effect is a slight reordering of the GPU FP32 reductions when the scheduler fuses the new output nodes differently, which moves the audio cosine similarity by a few times 1e-6. Token-level results stay 100 percent exact against the PyTorch reference.
Audio tokenizer pipeline
Encode (voice cloning reference path) :
ref_audio @ 24 kHz (1, 1, T_samples)
-> resample 16 kHz (kaiser polyphase)
-> pad 160 each side
-> HuBERT.feature_extractor (320x downsample)
-> HuBERT.feature_projection (LayerNorm + Linear)
-> + pos_conv_embed (folded)
-> 12 transformer layers
-> mean over 13 hidden states (1, 768, T_sem)
-> downsample by 2 (semantic_downsample_factor)
-> SemanticEncoder (conv blocks) (1, 768, T_frames)
-> e_acoustic (DAC encoder, 5 down-blocks) (1, 256, T_frames)
-> concat dim=1 (1, 1024, T_frames)
-> fc Linear (1024 -> 1024) (1, 1024, T_frames)
-> RVQ encode (8 codebooks residual) (1, 8, T_frames) int @ 25 fps
Decode (TTS path) :
codes [B, 8, T] int
-> RVQ decode :
for k in 0..7 :
e_k = codebook[k].embed[codes[k, :]] (B, T, 64)
p_k = e_k @ project_out[k].W.T + bias[k] (B, T, 1024)
out += p_k
-> transpose (B, 1024, T)
-> fc2 Linear (1024 -> 256) (B, 256, T)
-> acoustic_decoder DAC :
conv1 (256 -> 1024, k=7, pad=3)
for block in 0..4, ratios [8, 5, 4, 2, 3] :
snake1 (alpha)
conv_t1 (IC -> OC, k=2*r, stride=r,
padding=ceil(r/2), output_padding=r%2)
for res_unit in 0..2, dilations [1, 3, 9] :
snake1 (alpha)
conv1 (OC, OC, k=7, dil=d, pad=3*d)
snake2 (alpha)
conv2 (OC, OC, k=1)
residual add
snake1 (alpha) final
conv2 (32 -> 1, k=7, pad=3)
-> audio (B, 1, 960*T)
960x upsample = 8 * 5 * 4 * 2 * 3. T_in @ 25 fps -> T_out @ 24 kHz exact.
DAC decoder block channels
block 0 : IC=1024 OC=512 stride=8 K=16 pad=4 output_pad=0
block 1 : IC=512 OC=256 stride=5 K=10 pad=3 output_pad=1
block 2 : IC=256 OC=128 stride=4 K=8 pad=2 output_pad=0
block 3 : IC=128 OC=64 stride=2 K=4 pad=1 output_pad=0
block 4 : IC=64 OC=32 stride=3 K=6 pad=2 output_pad=1
final : 32 -> 1
PyTorch ConvTranspose1d formula :
T_out = (T_in - 1)*stride - 2*padding + dilation*(kernel - 1) + output_padding + 1
With our parameters (d=1, k=2*s, p=ceil(s/2), op=s%2) the formula
collapses to T_out = stride * T_in exactly for all five blocks.
Snake activation
DAC reference formula (Hugging Face Snake1d.forward) :
y = x + (alpha + 1e-9).reciprocal() * sin(alpha * x)^2
ggml_snake(x, a, inv_b) computes y = x + sin^2(a * x) * inv_b.
Mapping :
a = alpha (loaded direct, BF16 to F32) inv_b = 1/(alpha + 1e-9) (precomputed CPU side at load, F32)
Both stored as F32 [1, C] tensors. alpha shape in checkpoint :
(1, C, 1), ggml ne = (1, C, 1). C lives on ne[1]. Loader reads C from
mt->ne[1].
ConvTranspose1d via GEMM + col2im_1d
PyTorch nn.ConvTranspose1d(IC, OC, kernel=K, stride=s, padding=p) with
weight shape (IC, OC, K). GGML decomposition :
1. Permute weight (IC, OC, K) PyTorch -> (IC, K*OC) ggml at load time.
Layout : dst[(oc*K + k) * IC + ic] = src[ic*OC*K + oc*K + k]
This makes k vary faster than oc inside the K*OC axis, matching
what ggml_compute_forward_col2im_1d_impl expects :
col_data[(oc * K + k) + t_in * K_OC]
2. Build runtime graph :
xt = ggml_cont(ctx, ggml_transpose(ctx, x)) # [IC, T_in]
col = ggml_mul_mat(ctx, w, xt) # [K*OC, T_in]
y = ggml_col2im_1d(ctx, col, stride, OC, padding) # [T_no_op, OC]
if (output_pad > 0)
y = ggml_pad(ctx, y, output_pad, 0, 0, 0) # right-pad zeros
if (bias)
y = ggml_add(ctx, y, bias_2d)
Validated math : T_no_op = (T_in - 1)*stride + K - 2*pad. Adding
output_pad right-pad gives the PyTorch output size exactly.
RVQ codec
Per-codebook tensors (k = 0..7) :
codebook.embed (1024, 64) PyTorch -> ggml ne=(64, 1024)
project_in.weight (64, 1024) PyTorch -> ggml ne=(1024, 64) encode-only
project_in.bias (64,) encode-only
project_out.weight (1024, 64) PyTorch -> ggml ne=(64, 1024)
project_out.bias (1024,)
Decode graph (per codebook k, accumulated) :
codes_k = ggml_view_1d(codes, T, k * stride) # [T] i32
e_k = ggml_get_rows(embed[k], codes_k) # [64, T]
p_k = ggml_mul_mat(project_out_w[k], e_k) # [1024, T]
p_k = ggml_add(p_k, project_out_b[k])
acc += p_k
Encode (residual loop) :
residual = embeddings_in
for k in 0..7 :
e_k = project_in[k](residual)
codes_k = argmin_i ||e_k - codebook[k].embed[i]||^2
quantized = project_out[k](codebook[k].embed[codes_k])
residual -= quantized
HuBERT semantic encoder
12 transformer layers Pre-LN, GELU FFN, MHA 12 heads * 64 dim, biases
on all QKVO. Pre-conv feature extractor : 7 Conv1D layers, kernels
[10, 3, 3, 3, 3, 2, 2], strides [5, 2, 2, 2, 2, 2, 2], GroupNorm on
the first only, GELU between. Feature projection LayerNorm + Linear
(512 -> 768). Positional embedding via grouped Conv1D (128 kernel,
16 groups), weight_norm folded at convert time. Final LayerNorm.
Output computation :
mean(stack(all_13_hidden_states, dim=1), dim=1) # (B, T_sem, 768)
This is unusual : the encoder averages across the initial input plus the 12 transformer layer outputs, not just the last hidden state.
Long-form TTS pipeline
pipeline_tts_synthesize_long orchestrates inputs longer than the
chunking threshold. It mirrors _generate_chunked in
omnivoice/models/omnivoice.py.
1. Estimate total target tokens via duration_estimate_tokens.
2. If T_total <= chunk_threshold_sec * frame_rate, run a single shot
pipeline_tts_synthesize and skip chunking.
3. Otherwise split text on punctuation with chunk_text_punctuation,
targeting chunk_duration_sec seconds per chunk.
4. Generate chunks sequentially :
- chunk 0 with no reference (auto voice / voice design path) or
with the external reference (cloning path)
- in the auto voice case, the audio tokens of chunk 0 become the
voice prompt for chunks 1..N, locking in the speaker identity
5. Cross-fade decoded chunks with cross_fade_chunks(rate, 0.3 s).
6. Apply post-processing on the merged waveform.
A shared Philox counter ctr_lo is threaded across MaskGIT calls so
PRNG state advances continuously between chunks, matching the global
torch.cuda.manual_seed behaviour on the reference side.
Text chunking
chunk_text_punctuation(text, chunk_len, min_chunk_len) splits text on
sentence-ending punctuation (skipping abbreviation periods), then
merges sentences into chunks of at most chunk_len UTF-8 codepoints.
Undersized chunks (< min_chunk_len) are merged into a neighbour.
The function operates on UTF-8 strings and treats length as codepoints,
matching Python len(str) semantics. Per chunk character budget :
n_chars = utf8_codepoint_count(full_text)
avg_tokens_per_char = T_total / n_chars
chunk_len = (int)(chunk_duration_sec * frame_rate / avg_tokens_per_char)
add_punctuation(text) appends a terminal . (Latin) or its CJK
equivalent when missing. Used on the reference transcript when
preprocess_prompt is on.
Audio post-processing
audio-postproc.h is a strict math port of omnivoice/utils/audio.py
plus the relevant pydub.silence routines. All public functions take
and return float32 mono PCM in [-1, 1] at the pipeline rate (24 kHz).
Silence detection runs on int16 samples to match pydub bit-for-bit.
remove_silence(buf, min_silence_ms, keep_silence_ms,
seek_step_ms, threshold_dbfs)
Splits buf on contiguous silent regions where every
seek_step_ms-long frame stays below threshold_dbfs (RMS, S16,
default -50 dBFS), keeps keep_silence_ms of leading and trailing
silence around each retained segment, and concatenates the result.
cross_fade_chunks(chunks, rate, fade_seconds)
Concatenates audio chunks with a linear cross-fade of fade_seconds
at each junction.
peak_normalize_half(buf)
Scales buf so peak |x| equals 0.5. Used in pure auto voice when no
reference RMS is available.
fade_and_pad(buf, rate, fade_seconds, pad_seconds)
Applies a linear fade in / fade out and pads silence at the start
and end. Default fade 0.05 s, pad 0.05 s. Mirrors final reference
post-step.
ref_rms is plumbed end to end and decides the volume branch :
ref_rms < 0 pure auto voice, peak_normalize_half on the cross-faded waveform ref_rms < 0.1 quiet reference, rescale by ref_rms / 0.1 otherwise no-op
When a reference WAV is provided, the CLI computes its RMS on the F32 samples after optional silence trimming and passes it down. The same quantity is used on the PyTorch side.
Voice modes
auto voice no ref-wav. Chunk 0 generates with no reference,
subsequent chunks reuse chunk 0 audio tokens as the
voice prompt. peak_normalize_half on output.
voice design no ref-wav, --instruct provides one or more attribute
markers (gender, age, pitch, style, volume, emotion)
resolved by voice_design.h to the EN/ZH instruct
string the reference uses. Chunking behaves like
auto voice.
voice cloning --ref-wav and --ref-text provided. The reference is
resampled to 16 kHz, run through the audio tokenizer
encoder, and the resulting RVQ codes are reused as
the voice prompt for every chunk. The reference RMS
sets the target loudness.
Public API
Two layers, picked by use case.
Top-level public ABI : src/omnivoice.h
Single-header, plain C99, linkage extern "C". The opaque ov_context
handle aggregates the GGML backend pair, the LM pipeline, the audio
tokenizer codec, the BPE tokenizer and the voice-design vocabulary.
One init, one free, one synthesize call covers the full TTS path.
Same names, same struct layout, same calling convention from C, C++,
Python ctypes, Rust bindgen, Go cgo or any other binding generator.
#include "omnivoice.h"
struct ov_init_params iparams;
ov_init_default_params(&iparams);
iparams.model_path = "models/omnivoice-base-Q8_0.gguf";
iparams.codec_path = "models/omnivoice-tokenizer-F32.gguf";
struct ov_context * ov = ov_init(&iparams);
struct ov_tts_params params;
ov_tts_default_params(¶ms);
params.text = "Hello world.";
params.lang = "English";
struct ov_audio audio = { 0 };
enum ov_status rc = ov_synthesize(ov, ¶ms, &audio);
if (rc == OV_STATUS_OK) {
/* audio.samples is a malloc'd buffer of audio.n_samples floats
at audio.sample_rate Hz, audio.channels = 1 (mono) */
}
ov_audio_free(&audio);
ov_free(ov);
Status codes :
OV_STATUS_OK 0
OV_STATUS_INVALID_PARAMS -1 (mutually exclusive ref inputs etc.)
OV_STATUS_INSTRUCT_INVALID -2 (instruct rejected by VoiceDesign)
OV_STATUS_GENERATE_FAILED -3 (any internal generate / decode fail)
OV_STATUS_OOM -4 (output samples allocation failed)
OV_STATUS_CANCELLED -5 (cancel callback returned true)
ov_tts_params exposes cancel and cancel_user_data. The pipeline
polls between chunks of long-form output, so cancel granularity is
roughly chunk_duration_sec (15 s by default).
The MaskGIT sampler config is flattened directly into ov_tts_params
as seven mg_* fields ; ov_tts_default_params initialises them to
the reference defaults (num_step=32, guidance_scale=2.0, t_shift=0.1, layer_penalty_factor=5.0, position_temperature=5.0, class_temperature=0.0, seed=42).
ov_version() returns a static string of the form
"MAJOR.MINOR.PATCH (git-hash, date)". The macros OV_VERSION_MAJOR,
OV_VERSION_MINOR, OV_VERSION_PATCH are also available at
compile time for feature-detection.
ABI guarantee
tests/abi-c.c is built on every build with
-std=c99 -Wall -Werror -pedantic. It includes the public header,
calls every entry through its early-return path, and is wired into
the default build target. Any regression that breaks plain C
consumability fails the main build, not an opt-in step.
The static library libomnivoice-core.a is the default build
artefact. For binding consumers, configure with
-DOMNIVOICE_SHARED=ON to add a libomnivoice.so (or .dll /
.dylib) shared target that exports only the ov_* symbols ;
every internal pipeline_* and backend_* stays hidden behind
-fvisibility=hidden. Install rules follow GNUInstallDirs.
Low-level API : src/pipeline-tts.h, src/pipeline-codec.h
Direct access to the LM forward (pipeline_tts_llm_forward,
pipeline_tts_llm_forward_batched), the MaskGIT-only path
(pipeline_tts_generate), the codec encode / decode
(pipeline_codec_encode, pipeline_codec_decode), the instruct
resolver (pipeline_tts_resolve_instruct) and the manual init / free
(pipeline_tts_load, pipeline_codec_load, backend_init).
Used by --llm-test and --maskgit-test in omnivoice-tts, by
omnivoice-codec for the standalone codec roundtrip, and by the
Python cossim harness through dump files.
This layer is intentionally not part of the public ABI (C++ types in the signatures, no visibility export). It exists for the in-tree debug paths and stays available as long as the bundled CLI tools need it. The handle layer above is the recommended entry for everything else.
CLI tools
omnivoice-tts
End-to-end synthesis : text on stdin, WAV file on disk. Verbatim
--help (the binary also prints an omnivoice.cpp <hash> (<date>)
banner line first) :
Usage: omnivoice-tts --model <gguf> --codec <gguf> [options] -o <out.wav> < text.txt
Required:
--model <gguf> LLM GGUF (F32 / BF16 / Q8_0)
--codec <gguf> Codec GGUF (omnivoice-tokenizer-*.gguf)
-o <path> Output WAV (24 kHz mono). '-' streams to stdout (pipe friendly).
Input:
stdin Target text to synthesise. With -o '-', stdin is read
incrementally and synthesis starts as soon as the first
sentence boundary is reached. With -o file.wav, stdin is
read fully then synthesised in one shot.
--srt <path> Dub an SRT: synth each cue into its time slot, write one
timeline WAV ready to mux. Pairs with --ref-wav / --ref-rvq
for a cloned voice. Per cue duration comes from the SRT.
Optional:
--format <fmt> WAV output format: wav16, wav24, wav32 (default: wav16)
--lang <str> Language label (default 'None')
--instruct <str> Style instruction (default 'None')
--duration <sec> Output duration in seconds (default: estimate from text)
--no-denoise Omit the <|denoise|> prefix
--ref-wav <path> Reference WAV for voice cloning
--ref-text <path> Transcript file for the reference (required with --ref-wav / --ref-rvq)
--ref-rvq <path> Pre-encoded reference codes from omnivoice-codec (replaces --ref-wav)
--seed <int> Sampling seed (default: -1 for random)
--steps <int> MaskGIT decode steps (default: 32, fewer is faster)
--no-preprocess-prompt Skip ref-wav silence trim and ref-text terminal punctuation
--chunk-duration <sec> Long-form chunk duration (default: 15.0, <= 0 disables chunking)
--chunk-threshold <sec> Activate chunking above this estimated duration (default: 30.0)
--stream-by-line Flush synthesis at each newline, one WAV header per line (-o '-')
Debug:
--no-fa Disable flash attention
--clamp-fp16 Clamp hidden states to FP16 range
--dump <dir> Dump intermediate tensors (f32) to <dir>
--llm-test <input.bin> Full LLM forward, dump audio_logits
--maskgit-test Greedy MaskGIT decoder, dump audio_tokens [K, T]
(no codec decode, reads target text from stdin)
omnivoice-codec
Audio tokenizer round-trip : WAV to RVQ codes, RVQ codes to WAV.
Verbatim --help :
Usage: omnivoice-codec --model <gguf> -i <input>
Required:
--model <gguf> Codec GGUF (omnivoice-tokenizer-*.gguf)
-i <path> Input. WAV -> encode, .rvq -> decode
Optional:
--format <fmt> WAV output format: wav16, wav24, wav32 (default: wav16)
Output is auto-named next to input : clip.wav -> clip.rvq, clip.rvq -> clip.wav.
Encode applies the TTS reference preprocessing (RMS auto-gain, silence trim,
hop truncation); the resulting .rvq feeds omnivoice-tts --ref-rvq directly.
The .rvq file is a small binary container with shape [8, T] int32
codes plus a header carrying the sample rate and frame rate.
Module map
src/
backend.h GGML backend init, scheduler factory, env override
weight-ctx.h Generic weight context for GGUF loaders
gguf-weights.h mmap GGUF, gf_load_tensor, gf_get_*
audio-io.h WAV read, mono write (S16 / S24 / F32)
audio-resample.h Kaiser polyphase 24 kHz <-> 16 kHz
audio-postproc.h remove_silence, peak_normalize_half, fade_and_pad,
cross_fade_chunks. Strict pydub / utils.audio port.
wav.h WAV header reader (PCM16/24/F32, mono/stereo)
philox.h Philox4x32-10 counter-based PRNG, PyTorch CUDA aligned
debug.h Tensor dumper for cossim tests
bpe.h Qwen2 / GPT-2 byte-level BPE tokenizer, GGUF loader
lang-map.h Language name to ISO 639-3 ID resolution
(auto-generated from omnivoice/utils/lang_map.py)
voice-design.h Speaker attribute validation and EN / ZH instruct
resolution (mirrors voice_design.py)
text-chunker.h chunk_text_punctuation, add_punctuation, END_PUNCTUATION
duration-estimator.h RuleDurationEstimator port (per-script weights,
Unicode category fallback)
rvq-codec.h Residual VQ encode + decode (8 codebooks)
dac-decoder.h DAC acoustic decoder (5 blocks, ratios 8 5 4 2 3)
dac-encoder.h DAC acoustic encoder (mirror of decoder)
semantic-enc.h SemanticEncoder convs (768 -> 768)
hubert-enc.h HuBERT base (feature extractor + pos_conv +
12 transformer layers + final LN)
qwen3-enc.h Qwen3 transformer building blocks
omnivoice-llm.h OmniVoice TTS LLM weights and graph helpers
prompt-tts.h Prompt builder (denoise + lang + instruct + text +
ref + mask) and CFG batch stacking
maskgit-tts.h Iterative non autoregressive decoder, configurable
step count (32 default), CFG, layer penalty, gumbel
sampling, deterministic in greedy mode
pipeline-codec.{h,cpp} Audio tokenizer end-to-end (encode and decode)
pipeline-tts.{h,cpp} Full TTS orchestration, single shot and chunked,
plus low-level entries kept available for the
debug paths and the cossim test harness
omnivoice.{h,cpp} Public ABI : opaque ov_context handle, plain C99
header in extern "C", consumable from C, C++,
Python ctypes, Rust bindgen, Go cgo
tools/
omnivoice-tts.cpp CLI : text to WAV (auto / design / clone)
omnivoice-codec.cpp CLI : codes <-> WAV
quantize.cpp GGUF requantizer
version.cmake Embeds the git short hash into the binary
tests/
debug-tts-cossim.py Byte-level comparison of every pipeline stage
against the PyTorch reference, voice design path
debug-clone-cossim.py Same, voice cloning path
cross-decode.py Cross check : decode C++ tokens through PyTorch
codec and vice versa
prompt.txt Long-form English TTS sample
ref-audio.wav Voice cloning reference clip
ref-text.txt Transcript matching ref-audio.wav
abi-c.c Plain C99 smoke test for the public ABI ; built
with -Wall -Werror -pedantic on every build,
locks in C consumability and symbol linkage
GGML conventions
Tensor shape and layout
PyTorch shape (out, in) for a Linear weight stores as ggml
ne[0]=in, ne[1]=out. The GGUF tensor-shape array is reversed, so
reading reversed(t.shape) from gguf-py yields the PyTorch shape
directly.
For PyTorch Conv1d weight (OC, IC, K), ggml ne is (K, IC, OC).
The kernel axis is innermost (contiguous in memory).
For PyTorch ConvTranspose1d weight (IC, OC, K), the convert-time
permutation to ggml (IC, K*OC) rearranges
(oc*K + k) * IC + ic so that ggml_col2im_1d receives the correct
column matrix.
ggml_mul_mat(A, B) : with A.ne[0] = K (must match B.ne[0]),
A.ne[1] = M, B.ne[1] = N, output has ne = (N, M). In PyTorch terms,
A is (M, K), B is (N, K), output is (M, N), which equals
A @ B^T.
Custom GGML ops
Provided by the ServeurpersoCom/ggml fork :
ggml_snake(ctx, x, a, inv_b) : y = x + sin^2(a * x) * inv_b.
F32 / F16 / BF16 input/output. CPU + CUDA + Metal + Vulkan.
ggml_col2im_1d(ctx, a, s0, oc, p0) : scatter-add [K*OC, T_in]
columns into [T_out, OC] signal where
T_out = (T_in - 1)*s0 + K - 2*p0. Layout requires k to vary faster
than oc inside the K*OC axis. F32 / F16 / BF16. All backends. The
fork also folds the padding crop into this op via the p0 parameter,
removing a follow-up ggml_view for the typical ConvTranspose1d use
case.
Backend lifecycle
backend_init("MOD") then backend_sched_new(bp, max_nodes). Backend
handles are shared across modules in the same binary, refcounted. The
GPU backend is the default, the CPU backend is kept as a scheduler
fallback.
Validation
The reference comparison harness is tests/debug-tts-cossim.py and
tests/debug-clone-cossim.py. They run the same input through the
PyTorch reference (with TF32 disabled, eager attention) and through the
C++ binary, dump each pipeline stage to disk, and report cosine
similarity per stage. Latest run, chunked path, English long-form
prompt :
TTS chunked: Logits cos=1.000000 max 3.5e-04
Step0 pred_tokens 99.93% (2 FP flips)
Tokens 1.000000 exact 100.00%
Audio 0.999991
Clone chunked: Lf hidden cos=1.000000 max 1.9e-03
Logits cos=1.000000
Step0 pred_tokens 100.00%
Tokens 1.000000 exact 100.00%
Audio 0.999989
The few Step0 token flips are argmax ties at the FP epsilon (~2e-5 between top1 and top2 at those positions), inherent to the mixed cuBLAS vs GGML kernel arithmetic. They resorb over the 32 MaskGIT steps so the final tokens match bit for bit and decoded audio cosine is
0.9999.
Glossary
RVQ Residual Vector Quantisation. Stack of codebooks where each one quantises the residual from the previous codebook reconstruction.
DAC Descript Audio Codec. Convolutional encoder/decoder over residual VQ codes.
HuBERT Hidden-Unit BERT. Transformer encoder pretrained with masked acoustic unit prediction. Used here to extract semantic embeddings from raw audio.
Snake Periodic activation introduced in BigVGAN,
y = x + (1/alpha) * sin^2(alpha * x). Replaces
LeakyReLU in the DAC encoder/decoder.
CFG Classifier-Free Guidance. The model is run twice
(conditional and unconditional) and the outputs combined
as c + scale * (c - u) to amplify the conditional
signal.
MaskGIT Masked Generative Image Transformer (Chang et al., arXiv:2202.04200). Iterative non autoregressive decoder where masked tokens are progressively unmasked over a fixed number of steps, prioritising high-confidence positions per step. Originally introduced for image generation, adapted here to audio codes.
Philox Counter-based PRNG used by PyTorch CUDA. Thread safe and skip-ahead friendly, well suited to deterministic chunked inference.