Data Science Wire

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning

arXiv cs.CV1mo4 min read

arXiv:2607.00465v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) rely extensively on Visual Instruction Tuning (VIT) to elicit their multimodal reasoning capabilities. However, we find a discrepancy: VIT often packs multiple language tasks about the same image for conversational, multi-turn training, whereas existing benchmarks evaluate LVLMs in isolated, single-turn scenarios. The models can suffer from visual attention decay and contextual overfitting during multi-turn training, making it hard for them to realize their full potential in the mismatched test phase. To close

Read the full story at arXiv cs.CV

More in Machine Learning