← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

Byte-Exact KV-Cache Grafting Boosts Frozen Small Model Performance and Efficiency

Researchers present a method for depositing verified knowledge as a byte-exact KV-cache artifact that can be grafted into a frozen small language model, enabling bit-exact logits and zero KL divergence. On the AIME 2025 benchmark, a frozen Gemma-4-12B model's accuracy improves from 80.0% to 93.3% with grafted solutions, and token usage for recurring problems is reduced by over 6,500x. The approach also extends usable context from 32K to 2.85M tokens without additional accelerator memory.

Why it matters: This work demonstrates a significant advance in making frozen small models both more capable and efficient through exact KV-cache grafting, enabling practical knowledge reuse without retraining.

Full story at: arXiv Computation and Language