中文
相关论文

相关论文: Vidi2.5: Large Multimodal Models for Video Underst…

200 篇论文

Existing sign language translation methods follow a two-stage pipeline: first converting the sign language video to a gloss sequence (i.e. Sign2Gloss) and then translating the generated gloss sequence into a spoken language sentence (i.e.…

计算与语言 · 计算机科学 2023-12-19 Liqiang Jing , Xuemeng Song , Xinxing Zu , Na Zheng , Zhongzhou Zhao , Liqiang Nie

This work investigates a fundamental question: Do Video-Language Models (VidLMs) robustly account for video content, temporal sequence, and motion? Our investigation shows that, surprisingly, they often do not. We introduce REVEAL{}, a…

We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g. Flamingo, SeViLA, or GPT-4). Compared to existing benchmarks that focus on…

Any new medium, once it emerges, is used for more than the transmission of overt content alone. The information it carries typically operates on two levels: one is the content directly presented, while the other is the subtext beneath…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Qi Li , Xinchao Wang

Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major focus in the field of video synthesis. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Ailing Zhang , Lina Lei , Dehong Kong , Zhixin Wang , Jiaqi Xu , Fenglong Song , Chun-Le Guo , Chang Liu , Fan Li , Jie Chen

Large Multimodal Models (LMMs) have achieved impressive success in visual understanding and reasoning, remarkably improving the performance of mathematical reasoning in a visual context. Yet, a challenging type of visual math lies in the…

计算机视觉与模式识别 · 计算机科学 2024-05-09 Yunxin Li , Baotian Hu , Haoyuan Shi , Wei Wang , Longyue Wang , Min Zhang

The "Reason-Then-Respond" paradigm, enhanced by Reinforcement Learning, has shown great promise in advancing Multimodal Large Language Models. However, its application to the video domain has led to specialized models that excel at either…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Xinlong Chen , Yuanxing Zhang , Yushuo Guan , Weihong Lin , Zekun Wang , Bohan Zeng , Yang Shi , Sihan Yang , Qiang Liu , Pengfei Wan , Liang Wang , Tieniu Tan

Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-context streaming video…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Zhenyu Yang , Yuhang Hu , Zemin Du , Dizhan Xue , Shengsheng Qian , Jiahong Wu , Fan Yang , Weiming Dong , Changsheng Xu

Visual transformation reasoning (VTR) is a vital cognitive capability that empowers intelligent agents to understand dynamic scenes, model causal relationships, and predict future states, and thereby guiding actions and laying the…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Yuheng Ji , Yipu Wang , Yuyang Liu , Xiaoshuai Hao , Yue Liu , Yuting Zhao , Huaihai Lyu , Xiaolong Zheng

We introduce STRIVE (SpatioTemporal Reinforcement with Importance-aware Variant Exploration), a structured reinforcement learning framework for video question answering. While group-based policy optimization methods have shown promise in…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Emad Bahrami , Olga Zatsarynna , Parth Pathak , Sunando Sengupta , Juergen Gall , Mohsen Fayyaz

Visual Chain-of-Thought (VCoT) has emerged as a promising paradigm for enhancing multimodal reasoning by integrating visual perception into intermediate reasoning steps. However, existing VCoT approaches are largely confined to static…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Junhua Liu , Zhangcheng Wang , Zhike Han , Ningli Wang , Guotao Liang , Kun Kuang

Mathematical reasoning in real-world video settings presents a fundamentally different challenge than in static images or text. It requires interpreting fine-grained visual information, accurately reading handwritten or digital text, and…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Hanoona Rasheed , Abdelrahman Shaker , Anqi Tang , Muhammad Maaz , Ming-Hsuan Yang , Salman Khan , Fahad Shahbaz Khan

We introduce Emu3.5, a large-scale multimodal world model that natively predicts the next state across vision and language. Emu3.5 is pre-trained end-to-end with a unified next-token prediction objective on a corpus of vision-language…

Comprehensive highway scene understanding and robust traffic risk inference are vital for advancing Intelligent Transportation Systems (ITS) and autonomous driving. Traditional approaches often struggle with scalability and generalization,…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Yunxiang Yang , Ningning Xu , Jidong J. Yang

Recent Video-Language Models (VLMs) achieve promising results on long-video understanding, but their performance still lags behind that achieved on tasks involving images or short videos. This has led to great interest in improving the long…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Lars Doorenbos , Federico Spurio , Juergen Gall

The video topic segmentation (VTS) task segments videos into intelligible, non-overlapping topics, facilitating efficient comprehension of video content and quick access to specific content. VTS is also critical to various downstream video…

人工智能 · 计算机科学 2024-12-31 Hai Yu , Chong Deng , Qinglin Zhang , Jiaqing Liu , Qian Chen , Wen Wang

Video Large Language Models (Video LLMs) have achieved impressive performance on video-and-language tasks, such as video question answering. However, most existing Video LLMs neglect temporal information in video data, leading to struggles…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Zi-Yuan Hu , Yiwu Zhong , Shijia Huang , Michael R. Lyu , Liwei Wang

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Shenghao Fu , Qize Yang , Yuan-Ming Li , Yi-Xing Peng , Kun-Yu Lin , Xihan Wei , Jian-Fang Hu , Xiaohua Xie , Wei-Shi Zheng

Understanding visual differences between dynamic scenes requires the comparative perception of compositional, spatial, and temporal changes--a capability that remains underexplored in existing vision-language systems. While prior work on…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Jiangtao Wu , Shihao Li , Zhaozhou Bian , Jialu Chen , Runzhe Wen , An Ping , Yiwen He , Jiakai Wang , Yuanxing Zhang , Jiaheng Liu

Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity types, failing to capture the freeform, multi-action…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Alejandro Aparcedo , Akash Kumar , Aaryan Garg , Dalton Pham , Wen-Kai Chen , Anirudh Bharadwaj , Aman Chadha , Yogesh Rawat