CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
Abstract
Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted reports, and spatial decisions depend on maps, scenes, or web pages. We introduce CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning. CLBench-V combines converted public benchmarks with newly constructed datasets spanning domains such as science, finance, long-document understanding, spatial reasoning, and web-based visual question answering. To reduce the cost of constructing domain-specific context-learning tasks, we further use automated construction and filtering procedures for our newly built datasets. Across 3,443 instances and six recent multimodal models, the best overall score is only 0.2847, indicating that multimodal context learning remains far from saturated. Moreover, InternVL3.5-30B-A3B performs best on context grounding and new knowledge learning, while Qwen3.5-Plus performs best on new information application. We further analyze judge reliability, context length, image count, and representative failure cases. Code is available at https://github.com/IamLihua/CLBench-V.
Community
Excited to share CLBench-V, a new benchmark evaluating whether multimodal models can truly learn from context—from grounding evidence to acquiring new knowledge.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents (2026)
- Multimodal Graph RAG for Long-range Visually Rich Document Understanding (2026)
- TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs (2026)
- CulMind: Benchmarking Multimodal Understanding and Reasoning in Chinese Cultural Heritage (2026)
- WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark (2026)
- SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards (2026)
- M2Note: Continual Evolution of Vision Language Models via Mistake Notebook Learning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.25294 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper