Study Finds LLMs Miss Multi-Sensor Physical Hazards Despite Single-Sensor Accuracy
A new arXiv preprint benchmarks five large language models on their ability to assess physical hazards using multi-sensor data. While the models performed nearly perfectly when a single sensor exceeded safety thresholds, they consistently failed to issue warnings when multiple sensors were simultaneously elevated but individually below their limits. This gap persisted across different models and input formats.
Why it matters: The findings highlight a significant limitation in current LLMs that could pose safety risks if these models are used for real-world physical hazard monitoring.
Full story at: arXiv AI/ML ↗