Vision-Language Models Often Rewrite Imperfect Text Instead of Transcribing Faithfully, Benchmark Finds
A new arXiv preprint introduces FaithC4, a multilingual benchmark designed to test how vision-language models (VLMs) handle imperfect text. The study finds that general-purpose VLMs frequently rewrite distorted or corrupted text into more plausible forms, rather than transcribing it faithfully, a behavior not detected by standard clean-text benchmarks. In contrast, OCR-specialized models and traditional OCR systems show much less degradation in transcription accuracy under similar perturbations.
Why it matters: This work highlights a systematic limitation in current VLMs that could impact applications requiring accurate document transcription, especially when dealing with imperfect or noisy text.
Full story at: arXiv Computation and Language ↗