MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

1.The Hong Kong Polytechnic University; 2.Eastern Institute of Technology, Ningbo
3.Harbin Institute of Technology (Shenzhen);

Hierarchical distribution of the MMOOC benchmark.

Hierarchical distribution of the MMOOC benchmark.

Comparison of conventional, refusal, and our MMOOC benchmarks.

Comparison of conventional, refusal, and our MMOOC benchmarks.

Abstract

Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but they often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts, while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present \textbf{MMOOC}, a large-scale benchmark for evaluating refusal ability and robust answering in MLLMs. MMOOC contains over 28K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answerability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training improves robustness.

Examples of the eight MMOOC scenarios.

Examples of the eight MMOOC scenarios.

BibTeX