Data Science Wire

On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs

Apple Machine Learning ResearchJul 24 min read

Reinforcement learning (RL) finetuning has become a key technique for enhancing large language models (LLMs) on reasoning-intensive tasks, motivating its extension to vision language models (VLMs). While RL-tuned VLMs improve on visual reasoning benchmarks, they remain vulnerable to weak visual grounding, hallucinations, and over-reliance on textual cues. We show that simple, controlled textual perturbations—misleading captions or incorrect chain-of-thought (CoT) traces—cause substantial drops in robustness and confidence, and that these effects are more pronounced when CoT consistency is…

Read the full story at Apple Machine Learning Research

More in Machine Learning