🚀 Open Research Position: Visual Reasoning in Large Vision-Language Models (LVLMs)
We are looking for motivated students to join our research on visual reasoning in Large Vision-Language Models (LVLMs) at RIML Lab.
🔍 Project Description
Large Vision-Language Models have achieved remarkable performance across a wide range of multimodal tasks. However, their ability to perform complex visual reasoning remains an open challenge. This research focuses on understanding, evaluating, and improving the reasoning capabilities of LVLMs, including multi-step reasoning, visual grounding, and reasoning over complex visual scenes.
📄 Relevant Papers
Question Aware Vision Transformer for Multimodal Reasoning
Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning
🔹 Must-Have Requirements
Strong Python programming skills
Knowledge of deep learning and machine learning fundamentals
Hands-on experience with PyTorch
Familiarity with Vision-Language Models or Large Language Models
Strong research interest and willingness to learn
Ready to start immediately
⏳ Workload
Commitment: At least 20 hours per week
📌 Note: Filling out this form does not guarantee acceptance. Only shortlisted candidates will be contacted via email.
🔗 Apply here: Form
💬 Telegram: @Arianaghamohseni
@RIMLLab
#research_position #ML_research #VisionLanguageModels #MultimodalAI #VisualReasoning #DeepLearning
June 4, 2026 3.3K 74