← Back to brief
ResearchOfficialPreprintarXiv Machine Learning

DGAP: Restoring Accuracy in One-Bit KV-Cache Quantization for LLMs

A new method called DGAP is proposed to address accuracy loss in low-bit KV-cache quantization for large language models (LLMs). By restoring the local distribution of top-K logits, DGAP recovers Llama-3.1-8B accuracy from 47.8% to 83.2% under one-bit quantization, with only modest decoding overhead. The approach targets structured local misranking, a key source of quality degradation in quantized inference.

Why it matters: This technique could make long-context LLM inference more memory- and bandwidth-efficient without sacrificing accuracy, improving the deployability of large models.

Full story at: arXiv Machine Learning