Data Science Wire

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks

arXiv cs.CV2w4 min read

arXiv:2607.13305v1 Announce Type: new Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding. We audit this assumption across twenty models spanning 2-78B parameters and ten architecture families. We introduce the Visual Dependency Gap (VDG), the difference in per-question correctness between original-video and black-screen conditions. Paired McNemar tests on MVBench show that accuracy and visual dependency are separable: models differ on original video (p = 0.0003) but not on black screens (p = 0.53). Across models, task-type ra

Read the full story at arXiv cs.CV

More in AI