Jul 19, 2026
Prefill vs Decode in CKE: Why One Model Needs GEMM, GEMV, and Two Execution Plans
Follow an LLM request from tokenization through prefill, KV-cache construction and decode, then see why CPU projections become GEMM or GEMV and how CKE compares them with llama.cpp.
Read post →