Fine-Tuned LLMs for Identifying Vulnerability Indicators in UK Police Incident Logs: Promise and Limitations
Researchers adapted a locally hosted, open-weight large language model (LLM) pipeline to estimate the prevalence of vulnerability indicators—such as mental ill health, substance misuse, alcohol dependence, and homelessness—in nearly 3,000 UK police incident logs. The study found that while LLMs can produce meaningful prevalence estimates at scale (e.g., mental ill health in about one in five incidents), naive deployment is unreliable: single-pass classifications are unstable and tend to over-assign indicators compared to human judgment. Achieving defensible measurements required substantial human review and statistical correction, highlighting significant uncertainty and resource demands.
Why it matters: This work demonstrates that LLMs can help extract population-level insights from unstructured police data, but their outputs require rigorous methodological safeguards to be reliable, limiting their immediate operational use.
Full story at: arXiv Computers and Society ↗