Are We Ready for Multi-Image Reasoning? Launching VHs: The Visual Haystacks Benchmark!
View original at bair.berkeley.eduAre We Ready for Multi-Image Reasoning? Launching VHs: The Visual Haystacks Benchmark! <!-- These are comments in HTML. The above header text is needed to format the title, authors, etc…
What we drew from this source
The claims Via News extracted from this document. We point to the source; we don't replace it.
Simple captioning (LLaVA) combined with LLM aggregator (Llama3) outperforms all LMM-based methods with 5+ images, demonstrating current LMMs are inadequate for cross-image information integration
80% confidenceVisual domain exhibits Lost-in-Middle phenomenon analogous to NLP, with LLaVA performing best with needle before question and proprietary models preferring needle at start
80% confidenceMIRAGE retriever significantly outperforms CLIP on question-like text retrieval without efficiency loss
80% confidenceAll evaluated models show significant performance falloff as haystack size increases, with proprietary models failing above 1K images due to API payload limits
80% confidenceVisual Haystacks is the first visual-centric NIAH benchmark, compared to prior text-based OCR retrieval approaches
80% confidence
