🪢 Compositional Learning Journal Club
Join us this week for a critical exploration of robustness in Visual Question Answering systems and the broader implications for visual–language model reliability. We’ll analyze how even subtle, meaning-preserving changes to inputs can destabilize model outputs and discuss what this means for future evaluation and model design.
🌟 This Week's Presentation
📄 Paper:
Questioning the Stability of Visual Question Answering
🧠 Abstract:
Modern Visual Language Models (VLMs) have achieved impressive performance on a wide range of visual reasoning tasks, yet fundamental questions remain about their robustness to benign input perturbations. This paper presents the first large-scale, systematic study of how VLMs respond to small, meaning-preserving changes—such as pixel shifts, light geometric transformations, padded rescaling, paraphrasing, and multilingual rewrites—that do not change the true semantics of an image–question pair.
Across multiple datasets and models, the authors find that minor visual or textual perturbations frequently lead to different predicted answers, even for state-of-the-art systems like GPT-4o and Gemini 2.0 Flash. They also show that stability under perturbations correlates strongly with correctness, and that the stability patterns of small open-source models can be used to predict when larger models will fail.
In this session, we’ll discuss:
• What kinds of input changes most disrupt VQA predictions.
• How stability can serve as a proxy for reliability and model confidence.
• Implications for evaluation benchmarks and future model development.
🎙 Presenter: Amir Kasaei
Session Details:
- 📅 Date: Tuesday, December 23rd
- 🕒 Time: 3:00 PM - 4:00 PM
- 🌐 Location: Online at vc.sharif.edu/ch/rohban
We look forward to your participation! ✌️
arXiv.org
Questioning the Stability of Visual Question Answering
Visual Language Models (VLMs) have achieved remarkable progress, yet their reliability under small, meaning-preserving input changes remains poorly understood. We present the first large-scale,...

December 23, 2025 4.1K 12