Wednesday, August 26, 2026
Source trace. Via News points to the documents behind its reporting and shows what we drew from each — so you can check any claim. How we source
News articleBAIR Berkeley

Are We Ready for Multi-Image Reasoning? Launching VHs: The Visual Haystacks Benchmark!

View original at bair.berkeley.edu
Are We Ready for Multi-Image Reasoning? Launching VHs: The Visual Haystacks Benchmark! <!-- These are comments in HTML. The above header text is needed to format the title, authors, etc…
Opening lines of the source · BAIR Berkeley · short snapshot — read the full document at the original

What we drew from this source

The claims Via News extracted from this document. We point to the source; we don't replace it.

  • Simple captioning (LLaVA) combined with LLM aggregator (Llama3) outperforms all LMM-based methods with 5+ images, demonstrating current LMMs are inadequate for cross-image information integration

    80% confidence
  • Visual domain exhibits Lost-in-Middle phenomenon analogous to NLP, with LLaVA performing best with needle before question and proprietary models preferring needle at start

    80% confidence
  • MIRAGE retriever significantly outperforms CLIP on question-like text retrieval without efficiency loss

    80% confidence
  • All evaluated models show significant performance falloff as haystack size increases, with proprietary models failing above 1K images due to API payload limits

    80% confidence
  • Visual Haystacks is the first visual-centric NIAH benchmark, compared to prior text-based OCR retrieval approaches

    80% confidence

Cited in these Via News reports

What we know · the intelligence behind this page
Live from the substrate
What we're seeing
AI Leadership Exodus Rattles Investor Confidence Amid Capex Boom
High-profile departures at top AI labs — Brad Lightcap's exit from OpenAI and an unnamed researcher's departure from Alphabet/Google that triggered a share-price drop — are surfacing talent retention as a market risk factor even as hyperscalers pour record capital into AI infrastructure. The reaction shows investors treating key-person risk at frontier AI labs as material to valuation, a new fragility layered onto an otherwise bullish AI-driven capex cycle.
Our read on the data ›
Signals we're tracking
EPKINLY Regulatory-Clinical Success Cascade
High probability of expanded label indications, additional combination approvals, and competitive positioning strength in follicular lymphoma market. Predicts positive commercial uptake and potential accelerated review for related indications.
Patterns we're watching ›
Where sources disagree
Broadcom Inc.
Both facts report EPS for Broadcom Inc. for the same fiscal period (Q1 2026) observed on the same date (2026-02-01). However, they report conflicting values: 1.5 USD per share vs 2.05 USD per share. This is a 37% difference for the identical metric and time period, not a value change over time.
We flag conflicts openly ›
Recently verified
Checked against the original source
4,978
facts traced to their source — and we flag the ones that don't hold up.
101 entities tracked4,978 facts checked against source5,251 source documents archived
Query this data → isubstrate.com