Kokoro's phoneme encoder just took its first concrete step into CKE's generated-C lane. The step is deliberately small: one operation — three embedding tables plus a per-token LayerNorm — now lowers through the normal v8 pipeline into generated C, binds five real weight payloads, and matches a pinned PyTorch capture within 2.384185791015625e-7 maximum absolute error, at a fixed 36-token extent producing a [36,128] FP32 output.
It is worth saying out loud what this is not. This is not a complete phoneme encoder. It is not a generated waveform. CKE does not yet produce Kokoro speech from this circuit. What it is, is the first model-weight boundary of the Kokoro path traveling the exact route every other stage of the model will have to travel: BUMP weights → declared circuit → normal lowering → generated C → independent numerical comparison. The broader model survey — how Kokoro's TTS pipeline is shaped as a whole and what circuit decomposition was proposed for it — is in the earlier Kokoro post. This one zooms into the first cell of that decomposition and shows the evidence up close.
I ran the test suite for this boundary while writing this: five tests passed, one skipped (the optional exported-bundle check, which needs an asset I do not have mounted), and the evidence line the test prints put the worst error at 2.384185791015625e-7, token 0, channel 5. That is roughly two orders of magnitude inside the 2e-5 gate the comparison enforces. I re-ran the same suite at commit 75d001f43 while finalizing this draft: identical result, same worst error. The boundary did not move when the rest of the model started moving around it.
One timing note that shapes everything below. Between the merge of PR #586 and the head this post was written against (75d001f43), three follow-on boundaries merged: the ALBERT input projection (#587), the 12-head QKV prefix (#591), and the first full bidirectional attention context (#593). The embedding is no longer "the" Kokoro boundary — it is the first link of a connected chain. This post still zooms into that first link, because it is the one where the weight-binding pattern that every later link reuses was proven.

Why The Embedding Is The Right First Boundary
Kokoro's phoneme encoder is ALBERT-shaped, and the embedding layer is where model weights first touch the graph: word lookups, position lookups, token-type lookups, and a LayerNorm over the sum. Choosing it as the first generated-C boundary answers the question that every later stage inherits — how do real exported weights get bound to planned memory offsets in code the compiler emits, rather than code a human writes per model?
The kernel itself already existed. embedding_three_table_layer_norm.c is a 73-line, model-name-free, caller-owned FP32 primitive that had already passed its own oracle checks as a standalone operation. What it did not have was a circuit binding that exported its five weight references through the compiler. The merged PR #586 is small by design: seven lines registering the semantic operation in the existing v8 lowering tables, four lines teaching the checked emitter to validate planned weight offsets, one schema line allowing a single-checkpoint X-Ray profile, the circuit fixture, a 290-line test, and documentation plus nightly registration.
What The Boundary Computes
Per token t (36 tokens total, a fixed extent from the pinned reference utterance), the operation is:
x_t = word[word_ids[t]] + type[type_ids[t]] + position[t]
mu = mean(x_t) (accumulated in double)
v = var(x_t) (accumulated in double)
y_t = ((x_t - mu) / sqrt(v + epsilon)) * gamma + betaThe five weights, from the circuit's canonical references: word_embeddings at 178×128 (22,784 FP32, vocabulary size 178), position_embeddings at 512×128 (65,536), token_type_embeddings at 2×128 (256), and LayerNorm gamma/beta at 128 each. That is 88,832 FP32 weights — about 347 KiB — all caller-owned, none allocated by the kernel, and each one verified by SHA-256 before the test lets it touch the arena.
Two details matter for a CPU inference engine. First, the inputs are int32 IDs: word_ids in 0..177 and type_ids in 0..1 — the IDs are pinned fixture data, not the output of any phonemization step. Second, the LayerNorm statistics accumulate in double precision before casting back, which is exactly the kind of per-operation numerical contract CKE pins down in its kernel registry rather than leaving to compiler luck.
How The Circuit Reaches Generated C
The circuit is a version-3 bounded_graph: one block, a header operation embedding_three_table_layer_norm bound to the embedding_three_table_layer_norm_f32 kernel, external bindings for the ID inputs, an empty dense body, and a semantic checkpoint — kokoro.phoneme_encoder.embeddings, token-major, FP32, axes [token, channel]. The full fixture is in the repo at kokoro_embedding_generated_circuit.json.
From there nothing is special. The test drives the normal v8 pipeline — IR1 construction, lowering, memory layout, second lowering, call-ready IR (prefill mode) — then runs codegen_v8.py to emit the C source, then compiles it with cc -std=c11 -O2 -Wall -Wextra -Werror -pedantic alongside the kernel. For provenance, the test hashes the on-disk generated library before loading it, then captures the loaded library's identity from the runtime handle and records both — the captured path and the on-disk SHA-256 — in the X-Ray manifest. That is a provenance record tying the exercised path to the hashed artifact; it is not an attestation that every instruction executed came from that file.
Two compiler-side changes make that claim honest. The checked call emitter previously validated only planned activation buffers; it now validates planned weight offsets too, and a malformed plan — a weight offset at or past the arena size, or a duplicate buffer define — is rejected with a CheckedCallCodegenError before any code is emitted. Bad plans do not become binaries. The other change is small but consequential for single-operation subgraphs: the X-Ray parity profile now accepts a one-entry checkpoint_order, so a one-operation graph is compared as itself instead of needing an invented second checkpoint.
What The Independent Comparison Proves
The oracle is a pinned capture — Kokoro 0.9.4, Transformers 5.17.0, PyTorch 2.8.0+cpu — saved as an .npz fixture with its own SHA-256 manifest. If the fixture bytes drift, the test fails before it computes anything. That provenance line matters as much as the number: the comparison is against a specific, reproducible upstream capture, not a re-run of upstream code in the same environment as the candidate.
The test builds X-Ray manifests on both sides — the generated native side (backend ck, source generated_native, carrying runtime library identity) and the pinned capture side (backend pytorch, source pinned_capture) — and compares the single checkpoint under an FP32 contract: cosine similarity at least 0.99999, RMSE, relative RMSE, and maximum absolute error each at most 2e-5, with finiteness required. The status is pass, and my run recorded the worst absolute error at 2.384185791015625e-7, token 0, channel 5 — the same value the merged documentation reports.
What this proves: one operation of the Kokoro phoneme encoder, with its real exported weights, executed from generated C, numerically matches the pinned upstream capture within the gate. What it does not prove — and the test says so in its own printed evidence — is anything about the complete encoder (complete_kokoro_encoder: NOT_TESTED) or the waveform (generated_waveform: NOT_TESTED).
How Failure Is Supposed To Be Boring
A fair amount of the test's value is in what happens when inputs are wrong. I read through the rejection paths before believing the happy path:
- Invalid ID. Setting
word_ids[-1] = 178— the vocabulary size, one past the last valid index — returns a nonzero status, and the output buffer, pre-filled with a-777.0canary, stays exactly at that canary. No partial write, no corrupted row. - Repeated requests. The same arena is reused sequentially: an all-zero ID vector produces a different, valid output on the second call. The test exercises sequential arena reuse; nothing in it establishes concurrent re-entrancy or thread safety.
- Arena failures. An arena one byte too small, or one byte off alignment, returns
-2with the canary intact. The generated wrapper checks capacity and alignment before the kernel ever runs. - Malformed plans. A weight offset equal to the arena size, or two weights sharing a buffer define, is rejected during codegen with explicit error messages. The failure happens at build time, where it is cheap and legible.
- Read-only weights. All five payloads are SHA-256 hashed before and after the call. The kernel reads them; it does not modify them.
What Still Has To Connect Before Speech Is Heard
The Kokoro bring-up runbook is explicit about the remainder, and the wording matters: until the generated C entry point produces and verifies the waveform, the path status is unsupported / missing evidence. As of the head this draft was written against, the connected generated chain stops at the first attention context: the circuit kokoro_embedding_projection_attention_generated_boundary carries the embedding, the 768-wide input projection, the three 12-head QKV projections, and the full bidirectional attention over the pinned 36 tokens — thirteen effective BUMP weights, six compiler-declared X-Ray checkpoints. The merged attention PR reports its standalone provider within 5.125999450683594e-6 of the pinned PyTorch 2.12.1 eager capture, and the connected generated context within 7.748603820800781e-6; I have not re-run those suites, so I quote them from the merged evidence rather than claiming them as my own measurements.
After the attention context, the remainder — checked against the runbook's pinned-graph inventory, which cites the upstream source per stage — is longer than "duration, decoder, inverse STFT" suggests. The phoneme path itself still needs, inside the first shared ALBERT layer, the attention output projection, the residual normalization, the feed-forward network with the GELU variant, then the repeated shared-layer schedule for the remaining layers and the final 768→512 encoder projection, with padded or masked attention requiring its own explicit contract rather than inheriting the unmasked one. In parallel, the full Kokoro model carries a text encoder path — embedding, three weight-normalized Conv1D layers with leaky ReLU, and a bidirectional LSTM — whose output is fused with the phoneme path; the duration encoder and predictor (bidirectional LSTM with adaptive layer-norm pairs and a monotonic duration head); the prosody path (a shared bidirectional LSTM over expanded duration features with independent F0/N branches); the decoder (stride-2 convolutions, concatenation with the ASR features, and AdaIN residual blocks); the harmonic source and waveform generator (F0 upsampling, harmonic sines plus noise, voiced/unvoiced decisions); and finally the inverse STFT that yields the FP32 waveform. Voice conditioning is pinned but separate, and native G2P is explicitly a later milestone. The runbook also notes which downstream kernels already have primitive-level PyTorch oracles in isolation — duration, expansion, convolutions, adaptive layer norm, ISTFT — "candidates only" until they are connected in the circuit and compared at the graph level.
What this boundary actually demonstrates, setting aside any history of the other audio models, is a repeatable unit of work: a declared circuit with named weight references, normal lowering, generated C, and a pinned-capture comparison — with the parts that are not yet evidenced stamped NOT_TESTED instead of left implicit. The three follow-on merges (projection, QKV, attention) are the observable evidence that the unit is repeatable: each link went through the same pipeline with its own fixture, its own checkpoints, and its own quoted error budget. Whether that pace holds across the longer remaining stages is exactly what the next posts should measure, not assume.
Implementation Trail
- PR #586 — merged at b5f03325c: lowering registration, checked-emitter weight validation, single-checkpoint X-Ray profile, circuit fixture, test, docs, CI, and nightly suite registration.
- Follow-on links on the same pattern (merged after #586, before this draft's head): #587 ALBERT input projection, #591 12-head QKV prefix, #593 first attention context — one growing circuit,
kokoro_embedding_projection_attention_generated_circuit.json. - The kernel — 73 lines, no model names, validate-everything-before-write.
- The test — generated C, X-Ray parity, repeated requests, rejection paths, optional exported-BUMP verification.
- The circuit — bounded graph, one header operation, five weight references.
- The runbook — operation checklist and the new regression row; nightly key
tts_kokoro_generated_embedding.
Continue Reading
- How Kokoro TTS Works, and Where CKE Stands — the companion survey of the whole Kokoro graph, written before the projection, QKV, and attention-context boundaries merged. Read its status table as a dated snapshot.
- From Native Kernels to Generated Audio Circuits — the pattern this post extends: how CKE's audio models moved from native kernels into declared circuits and generated C.
- How CKE X-Ray Found Qwen3.6's First Bad Circuit — the comparison machinery this embedding relies on, applied to finding a real defect.
- The CKE Kernel Registry — how exact C functions and numerical contracts connect to lowering, which is what makes "the kernel the compiler picked" a claim you can check.
- How Audio Becomes Words — Whisper's kernels step by step on a CPU, for the audio-side context this post assumes.