Saturday, August 15, 2026
Facts you can rely on·101 entities·4,805 sourced facts
Source trace. Via News points to the documents behind its reporting and shows what we drew from each — so you can check any claim. How we source
News articleBAIR Berkeley

Are We Ready for Multi-Image Reasoning? Launching VHs: The Visual Haystacks Benchmark!

View original at bair.berkeley.edu
Are We Ready for Multi-Image Reasoning? Launching VHs: The Visual Haystacks Benchmark! <!-- These are comments in HTML. The above header text is needed to format the title, authors, etc…
Opening lines of the source · BAIR Berkeley · short snapshot — read the full document at the original

What we drew from this source

The claims Via News extracted from this document. We point to the source; we don't replace it.

  • Simple captioning (LLaVA) combined with LLM aggregator (Llama3) outperforms all LMM-based methods with 5+ images, demonstrating current LMMs are inadequate for cross-image information integration

    80% confidence
  • Visual domain exhibits Lost-in-Middle phenomenon analogous to NLP, with LLaVA performing best with needle before question and proprietary models preferring needle at start

    80% confidence
  • MIRAGE retriever significantly outperforms CLIP on question-like text retrieval without efficiency loss

    80% confidence
  • All evaluated models show significant performance falloff as haystack size increases, with proprietary models failing above 1K images due to API payload limits

    80% confidence
  • Visual Haystacks is the first visual-centric NIAH benchmark, compared to prior text-based OCR retrieval approaches

    80% confidence

Cited in these Via News reports