SynthDocBench: The Middle of a Long Document Is Still a Blind Spot

SynthDocBench tests seven vision-language models on 1,788 questions across documents averaging 51 pages. Five of six models in its position test performed worst on evidence placed in the middle.
artificial-intelligence
data-science
Author

Kabui, Charles

Published

2026-07-26

Keywords

synthdocbench, visual-document-understanding, long-context, vision-language-models, ai-benchmarks