Kokoro's phoneme encoder just took its first concrete step into CKE's generated-C lane. The step is deliberately small: one operation — three embedding tables plus a per-token LayerNorm — now lowers through the normal v8 pipeline into generated C, binds five real weight payloads, and matches a pinned PyTorch capture within 2.384185791015625e-7 maximum absolute error, at a fixed 36-token extent producing a [36,128] FP32 output.

It is worth saying out loud what this is not. This is not a complete phoneme encoder. It is not a generated waveform. CKE does not yet produce Kokoro speech from this circuit. What it is, is the first model-weight boundary of the Kokoro path traveling the exact route every other stage of the model will have to travel: BUMP weights → declared circuit → normal lowering → generated C → independent numerical comparison. The broader model survey — how Kokoro's TTS pipeline is shaped as a whole and what circuit decomposition was proposed for it — is in the earlier Kokoro post. This one zooms into the first cell of that decomposition and shows the evidence up close.

I ran the test suite for this boundary while writing this: five tests passed, one skipped (the optional exported-bundle check, which needs an asset I do not have mounted), and the evidence line the test prints put the worst error at 2.384185791015625e-7, token 0, channel 5. That is roughly two orders of magnitude inside the 2e-5 gate the comparison enforces. I re-ran the same suite at commit 75d001f43 while finalizing this draft: identical result, same worst error. The boundary did not move when the rest of the model started moving around it.

One timing note that shapes everything below. Between the merge of PR #586 and the head this post was written against (75d001f43), three follow-on boundaries merged: the ALBERT input projection (#587), the 12-head QKV prefix (#591), and the first full bidirectional attention context (#593). The embedding is no longer "the" Kokoro boundary — it is the first link of a connected chain. This post still zooms into that first link, because it is the one where the weight-binding pattern that every later link reuses was proven.

Editorial illustration of three phoneme-embedding streams meeting a measured tensor output; this is a conceptual image, not generated Kokoro audio.
A visual introduction to the embedding boundary. The measured output is a tensor, not a speech waveform.

Why The Embedding Is The Right First Boundary

Kokoro's phoneme encoder is ALBERT-shaped, and the embedding layer is where model weights first touch the graph: word lookups, position lookups, token-type lookups, and a LayerNorm over the sum. Choosing it as the first generated-C boundary answers the question that every later stage inherits — how do real exported weights get bound to planned memory offsets in code the compiler emits, rather than code a human writes per model?

Pinned phoneme IDs and three embedding tables passing through LayerNorm into a 36 by 128 output, then through normal CKE v8 lowering into generated native C, verified by a single-checkpoint X-Ray comparison against a pinned PyTorch capture.
The boundary: one Kokoro embedding operation through the normal pipeline, verified against a pinned capture. The rest of the encoder is the next series of boundaries, not this one.

The kernel itself already existed. embedding_three_table_layer_norm.c is a 73-line, model-name-free, caller-owned FP32 primitive that had already passed its own oracle checks as a standalone operation. What it did not have was a circuit binding that exported its five weight references through the compiler. The merged PR #586 is small by design: seven lines registering the semantic operation in the existing v8 lowering tables, four lines teaching the checked emitter to validate planned weight offsets, one schema line allowing a single-checkpoint X-Ray profile, the circuit fixture, a 290-line test, and documentation plus nightly registration.

Diagram of the Kokoro phoneme embedding boundary: 36 int32 word IDs and 36 int32 type IDs, five caller-owned weight buffers totaling 88,832 FP32 weights, and the 36 by 128 token-major FP32 output produced by per-token lookup sums plus LayerNorm.
Everything the boundary touches: 36 IDs in, 88,832 FP32 weights (~347 KiB) as caller-owned buffers, 4,608 output values out.

What The Boundary Computes

Per token t (36 tokens total, a fixed extent from the pinned reference utterance), the operation is:

Per-token embedding + LayerNormtext
x_t = word[word_ids[t]] + type[type_ids[t]] + position[t]
mu  = mean(x_t)      (accumulated in double)
v   = var(x_t)       (accumulated in double)
y_t = ((x_t - mu) / sqrt(v + epsilon)) * gamma + beta

The five weights, from the circuit's canonical references: word_embeddings at 178×128 (22,784 FP32, vocabulary size 178), position_embeddings at 512×128 (65,536), token_type_embeddings at 2×128 (256), and LayerNorm gamma/beta at 128 each. That is 88,832 FP32 weights — about 347 KiB — all caller-owned, none allocated by the kernel, and each one verified by SHA-256 before the test lets it touch the arena.

Two details matter for a CPU inference engine. First, the inputs are int32 IDs: word_ids in 0..177 and type_ids in 0..1 — the IDs are pinned fixture data, not the output of any phonemization step. Second, the LayerNorm statistics accumulate in double precision before casting back, which is exactly the kind of per-operation numerical contract CKE pins down in its kernel registry rather than leaving to compiler luck.

How The Circuit Reaches Generated C

The circuit is a version-3 bounded_graph: one block, a header operation embedding_three_table_layer_norm bound to the embedding_three_table_layer_norm_f32 kernel, external bindings for the ID inputs, an empty dense body, and a semantic checkpoint — kokoro.phoneme_encoder.embeddings, token-major, FP32, axes [token, channel]. The full fixture is in the repo at kokoro_embedding_generated_circuit.json.

From there nothing is special. The test drives the normal v8 pipeline — IR1 construction, lowering, memory layout, second lowering, call-ready IR (prefill mode) — then runs codegen_v8.py to emit the C source, then compiles it with cc -std=c11 -O2 -Wall -Wextra -Werror -pedantic alongside the kernel. For provenance, the test hashes the on-disk generated library before loading it, then captures the loaded library's identity from the runtime handle and records both — the captured path and the on-disk SHA-256 — in the X-Ray manifest. That is a provenance record tying the exercised path to the hashed artifact; it is not an attestation that every instruction executed came from that file.

Two compiler-side changes make that claim honest. The checked call emitter previously validated only planned activation buffers; it now validates planned weight offsets too, and a malformed plan — a weight offset at or past the arena size, or a duplicate buffer define — is rejected with a CheckedCallCodegenError before any code is emitted. Bad plans do not become binaries. The other change is small but consequential for single-operation subgraphs: the X-Ray parity profile now accepts a one-entry checkpoint_order, so a one-operation graph is compared as itself instead of needing an invented second checkpoint.

Five-stage evidence path: SHA-verified pinned PyTorch capture, bounded Kokoro circuit, normal v8 lowering, generated C with checked weight-offset validation, and single-checkpoint FP32 X-Ray comparison, with a band listing the NOT_TESTED remainder.
Every stage is a file in the repository. No stage is a black box the reader has to trust on faith.

What The Independent Comparison Proves

The oracle is a pinned capture — Kokoro 0.9.4, Transformers 5.17.0, PyTorch 2.8.0+cpu — saved as an .npz fixture with its own SHA-256 manifest. If the fixture bytes drift, the test fails before it computes anything. That provenance line matters as much as the number: the comparison is against a specific, reproducible upstream capture, not a re-run of upstream code in the same environment as the candidate.

The test builds X-Ray manifests on both sides — the generated native side (backend ck, source generated_native, carrying runtime library identity) and the pinned capture side (backend pytorch, source pinned_capture) — and compares the single checkpoint under an FP32 contract: cosine similarity at least 0.99999, RMSE, relative RMSE, and maximum absolute error each at most 2e-5, with finiteness required. The status is pass, and my run recorded the worst absolute error at 2.384185791015625e-7, token 0, channel 5 — the same value the merged documentation reports.

What this proves: one operation of the Kokoro phoneme encoder, with its real exported weights, executed from generated C, numerically matches the pinned upstream capture within the gate. What it does not prove — and the test says so in its own printed evidence — is anything about the complete encoder (complete_kokoro_encoder: NOT_TESTED) or the waveform (generated_waveform: NOT_TESTED).

How Failure Is Supposed To Be Boring

A fair amount of the test's value is in what happens when inputs are wrong. I read through the rejection paths before believing the happy path:

  • Invalid ID. Setting word_ids[-1] = 178 — the vocabulary size, one past the last valid index — returns a nonzero status, and the output buffer, pre-filled with a -777.0 canary, stays exactly at that canary. No partial write, no corrupted row.
  • Repeated requests. The same arena is reused sequentially: an all-zero ID vector produces a different, valid output on the second call. The test exercises sequential arena reuse; nothing in it establishes concurrent re-entrancy or thread safety.
  • Arena failures. An arena one byte too small, or one byte off alignment, returns -2 with the canary intact. The generated wrapper checks capacity and alignment before the kernel ever runs.
  • Malformed plans. A weight offset equal to the arena size, or two weights sharing a buffer define, is rejected during codegen with explicit error messages. The failure happens at build time, where it is cheap and legible.
  • Read-only weights. All five payloads are SHA-256 hashed before and after the call. The kernel reads them; it does not modify them.

What Still Has To Connect Before Speech Is Heard

The Kokoro bring-up runbook is explicit about the remainder, and the wording matters: until the generated C entry point produces and verifies the waveform, the path status is unsupported / missing evidence. As of the head this draft was written against, the connected generated chain stops at the first attention context: the circuit kokoro_embedding_projection_attention_generated_boundary carries the embedding, the 768-wide input projection, the three 12-head QKV projections, and the full bidirectional attention over the pinned 36 tokens — thirteen effective BUMP weights, six compiler-declared X-Ray checkpoints. The merged attention PR reports its standalone provider within 5.125999450683594e-6 of the pinned PyTorch 2.12.1 eager capture, and the connected generated context within 7.748603820800781e-6; I have not re-run those suites, so I quote them from the merged evidence rather than claiming them as my own measurements.

After the attention context, the remainder — checked against the runbook's pinned-graph inventory, which cites the upstream source per stage — is longer than "duration, decoder, inverse STFT" suggests. The phoneme path itself still needs, inside the first shared ALBERT layer, the attention output projection, the residual normalization, the feed-forward network with the GELU variant, then the repeated shared-layer schedule for the remaining layers and the final 768→512 encoder projection, with padded or masked attention requiring its own explicit contract rather than inheriting the unmasked one. In parallel, the full Kokoro model carries a text encoder path — embedding, three weight-normalized Conv1D layers with leaky ReLU, and a bidirectional LSTM — whose output is fused with the phoneme path; the duration encoder and predictor (bidirectional LSTM with adaptive layer-norm pairs and a monotonic duration head); the prosody path (a shared bidirectional LSTM over expanded duration features with independent F0/N branches); the decoder (stride-2 convolutions, concatenation with the ASR features, and AdaIN residual blocks); the harmonic source and waveform generator (F0 upsampling, harmonic sines plus noise, voiced/unvoiced decisions); and finally the inverse STFT that yields the FP32 waveform. Voice conditioning is pinned but separate, and native G2P is explicitly a later milestone. The runbook also notes which downstream kernels already have primitive-level PyTorch oracles in isolation — duration, expansion, convolutions, adaptive layer norm, ISTFT — "candidates only" until they are connected in the circuit and compared at the graph level.

What this boundary actually demonstrates, setting aside any history of the other audio models, is a repeatable unit of work: a declared circuit with named weight references, normal lowering, generated C, and a pinned-capture comparison — with the parts that are not yet evidenced stamped NOT_TESTED instead of left implicit. The three follow-on merges (projection, QKV, attention) are the observable evidence that the unit is repeatable: each link went through the same pipeline with its own fixture, its own checkpoints, and its own quoted error budget. Whether that pace holds across the longer remaining stages is exactly what the next posts should measure, not assume.

Two-column comparison: tested now lists the connected generated chain from the embedding through the input projection, QKV prefix, and first attention context with its rejection paths and X-Ray parity; still missing lists the attention feed-forward and second LayerNorm, the remaining shared ALBERT layers, duration prediction, decoder, ISTFT waveform, and phonemization, with complete encoder and waveform marked NOT_TESTED.
The honest split as of head 75d001f43: four links fully evidenced, everything downstream still open.

Implementation Trail

  • PR #586 — merged at b5f03325c: lowering registration, checked-emitter weight validation, single-checkpoint X-Ray profile, circuit fixture, test, docs, CI, and nightly suite registration.
  • Follow-on links on the same pattern (merged after #586, before this draft's head): #587 ALBERT input projection, #591 12-head QKV prefix, #593 first attention context — one growing circuit, kokoro_embedding_projection_attention_generated_circuit.json.
  • The kernel — 73 lines, no model names, validate-everything-before-write.
  • The test — generated C, X-Ray parity, repeated requests, rejection paths, optional exported-BUMP verification.
  • The circuit — bounded graph, one header operation, five weight references.
  • The runbook — operation checklist and the new regression row; nightly key tts_kokoro_generated_embedding.

Continue Reading