← Back to brief
ResearchOfficialPreprintarXiv Information Retrieval

STATIC: Efficient Constrained Decoding for LLM-Based Generative Retrieval on Accelerators

Researchers present STATIC, a method that transforms prefix trees into sparse matrices to enable efficient constrained decoding for LLM-based generative retrieval on TPUs and GPUs. STATIC achieves up to a 948x speedup over CPU trie implementations and minimal latency overhead when deployed on a large-scale video recommendation platform. The method also demonstrates improved cold-start performance on academic benchmarks.

Why it matters: STATIC enables the first production-scale deployment of strictly constrained generative retrieval, addressing a major efficiency bottleneck for industrial recommender systems that require business logic constraints.

Full story at: arXiv Information Retrieval