English
Related papers

Related papers: FoodMonitor: Benchmarking MLLMs for Explainable Co…

200 papers

The evaluation and improvement of medical large language models (LLMs) are critical for their real-world deployment, particularly in ensuring accuracy, safety, and ethical alignment. Existing frameworks inadequately dissect domain-specific…

Computation and Language · Computer Science 2025-03-11 Luyi Jiang , Jiayuan Chen , Lu Lu , Xinwei Peng , Lihao Liu , Junjun He , Jie Xu

Video Large Language Models (VideoLLMs) are increasingly deployed on numerous critical applications, where users rely on auto-generated summaries while casually skimming the video stream. We show that this interaction hides a critical…

Multimedia · Computer Science 2025-11-18 Yuxin Cao , Wei Song , Derui Wang , Jingling Xue , Jin Song Dong

Automatic vision inspection holds significant importance in industry inspection. While multimodal large language models (MLLMs) exhibit strong language understanding capabilities and hold promise for this task, their performance remains…

Information Retrieval · Computer Science 2026-04-06 Kai Zhang , Zekai Zhang , Xihe Sun , Anpeng Wang , Jingmeng Nie , Qinghui Chen , Han Hao , Jianyuan Guo , Jinglin Zhang

In this paper, we tackle the problem of how to build and benchmark a large motion model (LMM). The ultimate goal of LMM is to serve as a foundation model for versatile motion-related tasks, e.g., human motion generation, with…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Liang Xu , Shaoyang Hua , Zili Lin , Yifan Liu , Feipeng Ma , Yichao Yan , Xin Jin , Xiaokang Yang , Wenjun Zeng

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able…

Recognizing and localizing student confusion from video is an important yet challenging problem in educational AI. Existing confusion datasets suffer from noisy labels, coarse temporal annotations, and limited expert validation, which…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Lu Dong , Xiao Wang , Mark Frank , Srirangaraj Setlur , Venu Govindaraju , Ifeoma Nwogu

Watching instructional videos are often used to learn about procedures. Video captioning is one way of automatically collecting such knowledge. However, it provides only an indirect, overall evaluation of multimodal models with no…

Computation and Language · Computer Science 2020-10-12 Frank F. Xu , Lei Ji , Botian Shi , Junyi Du , Graham Neubig , Yonatan Bisk , Nan Duan

Existing object navigation benchmarks usually tell an embodied agent which object category to find, such as microwave or chair. Human-facing embodied AI is often asked something less direct: "I need something to warm this food" or "the room…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Lin Qian , Shijie Li , Sihao Lin , Xuan Zhang , Bangya Liu , Yanran Li , Hujun Yin

Large Multimodal Models (LMMs) have achieved remarkable progress across various capabilities; however, complex video reasoning in the scientific domain remains a significant and challenging frontier. Current video benchmarks predominantly…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Andong Deng , Taojiannan Yang , Shoubin Yu , Lincoln Spencer , Mohit Bansal , Chen Chen , Serena Yeung-Levy , Xiaohan Wang

Evaluating alignment in language models requires testing how they behave under realistic pressure, not just what they claim they would do. While alignment failures increasingly cause real-world harm, comprehensive evaluation frameworks with…

Artificial Intelligence · Computer Science 2026-02-25 Nora Petrova , John Burden

Vision-Language Models (VLMs) have achieved strong results in video understanding, yet a key question remains: do they truly comprehend visual content or only learn shallow correlations between vision and language? Real visual…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Zongxia Li , Xiyang Wu , Guangyao Shi , Yubin Qin , Hongyang Du , Fuxiao Liu , Tianyi Zhou , Dinesh Manocha , Jordan Lee Boyd-Graber

The demand for accurate food quantification has increased in the recent years, driven by the needs of applications in dietary monitoring. At the same time, computer vision approaches have exhibited great potential in automating tasks within…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Valasia Vlachopoulou , Ioannis Sarafis , Alexandros Papadopoulos

High-quality video datasets are foundational for training robust models in tasks like action recognition, phase detection, and event segmentation. However, many real-world video datasets suffer from annotation errors such as *mislabeling*,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Praditha Alwis , Soumyadeep Chandra , Deepak Ravikumar , Kaushik Roy

Controversial contents largely inundate the Internet, infringing various cultural norms and child protection standards. Traditional Image Content Moderation (ICM) models fall short in producing precise moderation decisions for diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Mengyang Wu , Yuzhi Zhao , Jialun Cao , Mingjie Xu , Zhongming Jiang , Xuehui Wang , Qinbin Li , Guangneng Hu , Shengchao Qin , Chi-Wing Fu

Regulatory compliance auditing across diverse industrial domains requires heightened quality assurance and traceability. Present manual and intermittent approaches to such auditing yield significant challenges, potentially leading to…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Jia Syuen Lim , Ziwei Wang , Jiajun Liu , Abdelwahed Khamis , Reza Arablouei , Robert Barlow , Ryan McAllister

As Large Language Models (LLMs) are increasingly deployed as task-oriented agents in enterprise environments, ensuring their strict adherence to complex, domain-specific operational guidelines is critical. While utilizing an LLM-as-a-Judge…

Computation and Language · Computer Science 2026-04-15 Jingbo Yang , Guanyu Yao , Bairu Hou , Xinghan Yang , Nikolai Glushnev , Iwona Bialynicka-Birula , Duo Ding , Shiyu Chang

Rapid advances in multimodal models demand benchmarks that rigorously evaluate understanding and reasoning in safety-critical, dynamic real-world settings. We present AccidentBench, a large-scale benchmark that combines vehicle accident…

Counting in long videos remains a fundamental yet underexplored challenge in computer vision. Real-world recordings often span tens of minutes or longer and contain sparse, diverse events, making long-range temporal reasoning particularly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fumihiko Tsuchiya , Taiki Miyanishi , Mahiro Ukai , Nakamasa Inoue , Shuhei Kurita , Yusuke Iwasawa , Yutaka Matsuo

With the rapid development of industrial intelligence and unmanned inspection, reliable perception and safety assessment for AI systems in complex and dynamic industrial sites has become a key bottleneck for deploying predictive maintenance…

Robotics · Computer Science 2026-01-30 Zeyi Liu , Shuang Liu , Jihai Min , Zhaoheng Zhang , Jun Cen , Pengyu Han , Songqiao Hu , Zihan Meng , Xiao He , Donghua Zhou

Accurate ground truth annotations are critical to supervised learning and evaluating the performance of autonomous vehicle systems. These vehicles are typically equipped with active sensors, such as LiDAR, which scan the environment in…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Alexandre Justo Miro , Ludvig af Klinteberg , Bogdan Timus , Aron Asefaw , Ajinkya Khoche , Thomas Gustafsson , Sina Sharif Mansouri , Masoud Daneshtalab
‹ Prev 1 3 4 5 6 7 10 Next ›