中文
相关论文

相关论文: Omni-Modal Dissonance Benchmark: Systematically Br…

200 篇论文

Despite recent progress in systematic evaluation frameworks, benchmarking the uncertainty of large language models (LLMs) remains a highly challenging task. Existing methods for benchmarking the uncertainty of LLMs face three key…

计算与语言 · 计算机科学 2025-06-05 Xunzhi Wang , Zhuowei Zhang , Gaonan Chen , Qiongyu Li , Bitong Luo , Zhixin Han , Haotian Wang , Zhiyu li , Hang Gao , Mengting Hu

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their…

Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evidence alone is enough to answer a query. We study whether current omni-modal benchmarks…

多媒体 · 计算机科学 2026-05-15 Che Liu , Lichao Ma , Xiangyu Tony Zhang , Yuxin Zhang , Haoyang Zhang , Xuerui Yang , Fei Tian

Foundation models (FMs) deployed in real-world tasks such as computer-use agents must integrate diverse modalities. How good are FMs at performing joint reasoning, simultaneously reasoning over multiple modalities, especially when the…

人工智能 · 计算机科学 2025-10-07 Chen Henry Wu , Neil Kale , Aditi Raghunathan

Learning from multiple modalities often suffers from imbalance, where information-rich modalities dominate optimization while weaker or partially missing modalities contribute less. This imbalance becomes severe in realistic settings with…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Phuong-Anh Nguyen , Tien Anh Pham , Duc-Trong Le , Cam-Van Thi Nguyen

Large Multimodal Models (LMMs), harnessing the complementarity among diverse modalities, are often considered more robust than pure Language Large Models (LLMs); yet do LMMs know what they do not know? There are three key open questions…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Ruiyang Zhang , Hu Zhang , Hao Fei , Zhedong Zheng

For Large Language Models (LLMs) to be reliably deployed in both everyday and high-stakes domains, knowing when not to answer is equally critical as answering correctly. Real-world user queries, which can be underspecified, ill-posed, or…

人工智能 · 计算机科学 2025-06-11 Polina Kirichenko , Mark Ibrahim , Kamalika Chaudhuri , Samuel J. Bell

Recent advancements in Unified Multimodal Models (UMMs) have enabled remarkable image understanding and generation capabilities. However, while models like Gemini-2.5-Flash-Image show emerging abilities to reason over multiple related…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Mingrui Wu , Hang Liu , Jiayi Ji , Xiaoshuai Sun , Rongrong Ji

Detecting mind wandering is crucial in online education, and it occurs 30% of the time, as it directly impacts learners' retention, comprehension, and overall success in self-directed learning environments. Integrating automated detection…

Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in \emph{over-refusal} of benign queries or \emph{unsafe compliance} with harmful ones. While existing…

人工智能 · 计算机科学 2026-01-27 Zhihao Zhang , Liting Huang , Guanghao Wu , Preslav Nakov , Heng Ji , Usman Naseem

Research on multi-modal learning dominantly aligns the modalities in a unified space at training, and only a single one is taken for prediction at inference. However, for a real machine, e.g., a robot, sensors could be added or removed at…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Yuanhuiyi Lyu , Xu Zheng , Dahun Kim , Lin Wang

A critical yet frequently overlooked challenge in the field of deepfake detection is the lack of a standardized, unified, comprehensive benchmark. This issue leads to unfair performance comparisons and potentially misleading results.…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Zhiyuan Yan , Yong Zhang , Xinhang Yuan , Siwei Lyu , Baoyuan Wu

Contrastive learning is effective for aligning paired views or modalities, but alignment beyond two modalities remains non-trivial and comparatively underexplored. Pairwise CLIP-style losses decompose multi-modal alignment into independent…

机器学习 · 计算机科学 2026-05-29 Tianchao Li , Shujian Yu , Xinrui Zu , Zhaolong Wei , Jeremy Gummeson , Jack C. P. Cheng , Robert Jenssen

Multimodal affective computing underpins key tasks such as sentiment analysis and emotion recognition. Standard evaluations, however, often assume that textual, acoustic, and visual modalities are equally available. In real applications,…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Tien Anh Pham , Phuong-Anh Nguyen , Duc-Trong Le , Cam-Van Thi Nguyen

Benchmarks shape scientific conclusions about model capabilities and steer model development. This creates a feedback loop: stronger benchmarks drive better models, and better models demand more discriminative benchmarks. Ensuring benchmark…

计算与语言 · 计算机科学 2025-10-01 Arda Uzunoglu , Tianjian Li , Daniel Khashabi

Large Multimodal Models(LMMs) face notable challenges when encountering multimodal knowledge conflicts, particularly under retrieval-augmented generation(RAG) frameworks where the contextual information from external sources may contradict…

While Instruction-based Image Editing (IIE) has achieved significant progress, existing benchmarks pursue task breadth via mixed evaluations. This paradigm obscures a critical failure mode crucial in professional applications: the…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Yujia Yang , Yuanxiang Wang , Zhenyu Guan , Tiankun Yang , Chenxi Bao , Haopeng Jin , Jinwen Luo , Xinyu Zuo , Lisheng Duan , Haijin Liang , Jin Ma , Xinming Wang , Ruiwen Tao , Hongzhu Yi

Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-modal large language models. Existing benchmarks fall short in…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Ruixiang Zhao , Jie Yang , Zijie Xin , Tianyi Wang , Fengyun Rao , Jing LYU , Xirong Li

We introduce OmnixR, an evaluation suite designed to benchmark SoTA Omni-modality Language Models, such as GPT-4o and Gemini. Evaluating OLMs, which integrate multiple modalities such as text, vision, and audio, presents unique challenges.…

Multi-modal 3D object detection models for automated driving have demonstrated exceptional performance on computer vision benchmarks like nuScenes. However, their reliance on densely sampled LiDAR point clouds and meticulously calibrated…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Till Beemelmanns , Quan Zhang , Christian Geller , Lutz Eckstein