Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
arXiv cs.CV1mo4 min read
arXiv:2607.00461v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens which can lose perceptual nuance. A promising alternative is continuous latent reasoning, where the goal is to discover implicit reasoning pathways that bridge the multimodal query and the final answer. However, this introduces a severe train-inference mismatch: a training-time posterior, conditioned on the ground-truth answer, can exploit answer-dependent shortcuts. Standard variational training the
