中文
相关论文

相关论文: ViLPAct: A Benchmark for Compositional Generalizat…

200 篇论文

Benefiting from language flexibility and compositionality, humans naturally intend to use language to command an embodied agent for complex tasks such as navigation and object manipulation. In this work, we aim to fill the blank of the last…

机器人学 · 计算机科学 2022-08-18 Kaizhi Zheng , Xiaotong Chen , Odest Chadwicke Jenkins , Xin Eric Wang

Video Internet of Things (VIoT) has shown full potential in collecting an unprecedented volume of video data. How to schedule the domain-specific perceiving models and analyze the collected videos uniformly, efficiently, and especially…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Yaoyao Zhong , Mengshi Qi , Rui Wang , Yuhan Qiu , Yang Zhang , Huadong Ma

Large Vision-Language Models (LVLMs) have made significant strides in the field of video understanding in recent times. Nevertheless, existing video benchmarks predominantly rely on text prompts for evaluation, which often require complex…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Yiming Zhao , Yu Zeng , Yukun Qi , YaoYang Liu , Xikun Bao , Lin Chen , Zehui Chen , Qing Miao , Chenxi Liu , Jie Zhao , Feng Zhao

Addressing the gap in understanding visual comprehension in Large Language Models (LLMs), we designed a challenge-response study, subjecting Google Bard and GPT-Vision to 64 visual tasks, spanning categories like "Visual Situational…

计算机视觉与模式识别 · 计算机科学 2023-10-18 David Noever , Samantha Elizabeth Miller Noever

We introduce VideoComp, a benchmark and learning framework for advancing video-text compositionality understanding, aimed at improving vision-language models (VLMs) in fine-grained temporal alignment. Unlike existing benchmarks focused on…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Dahun Kim , AJ Piergiovanni , Ganesh Mallya , Anelia Angelova

Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Ali Athar , Xueqing Deng , Liang-Chieh Chen

Recent work has highlighted the potential of modelling interactive behaviour analogously to natural language. We propose interactive behaviour summarisation as a novel computational task and demonstrate its usefulness for automatically…

人机交互 · 计算机科学 2024-10-14 Guanhua Zhang , Mohamed Ahmed , Zhiming Hu , Andreas Bulling

Translating visual data into natural language is essential for machines to understand the world and interact with humans. In this work, a comprehensive study is conducted on video paragraph captioning, with the goal to generate…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Qinyu Li , Tengpeng Li , Hanli Wang , Chang Wen Chen

As Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of performing multiple tasks simultaneously. However,…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Xin Jiang , Junwei Zheng , Ruiping Liu , Jiahang Li , Jiaming Zhang , Sven Matthiesen , Rainer Stiefelhagen

The study of complex human interactions and group activities has become a focal point in human-centric computer vision. However, progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from…

计算机视觉与模式识别 · 计算机科学 2024-05-06 Che-Jui Chang , Danrui Li , Deep Patel , Parth Goel , Honglu Zhou , Seonghyeon Moon , Samuel S. Sohn , Sejong Yoon , Vladimir Pavlovic , Mubbasir Kapadia

LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience gap. We attribute this gap to existing benchmarks' reliance on over-specified queries,…

计算与语言 · 计算机科学 2026-05-28 Xiaohongshu Inc

This paper introduces a new video-and-language dataset with human actions for multimodal logical inference, which focuses on intentional and aspectual expressions that describe dynamic human actions. The dataset consists of 200 videos,…

计算机视觉与模式识别 · 计算机科学 2021-06-29 Riko Suzuki , Hitomi Yanaka , Koji Mineshima , Daisuke Bekki

There is a large variation in the activities that humans perform in their everyday lives. We consider modeling these composite human activities which comprises multiple basic level actions in a completely unsupervised setting. Our model…

计算机视觉与模式识别 · 计算机科学 2016-03-14 Chenxia Wu , Jiemi Zhang , Ozan Sener , Bart Selman , Silvio Savarese , Ashutosh Saxena

Data visualizations are powerful tools for communicating patterns in quantitative data. Yet understanding any data visualization is no small feat -- succeeding requires jointly making sense of visual, numerical, and linguistic inputs…

人机交互 · 计算机科学 2025-05-26 Arnav Verma , Kushin Mukherjee , Christopher Potts , Elisa Kreiss , Judith E. Fan

Recent advancements in multimodal large language models have driven breakthroughs in visual question answering. Yet, a critical gap persists, `conceptualization'-the ability to recognize and reason about the same concept despite variations…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Zahra Babaiee , Peyman M. Kiasari , Daniela Rus , Radu Grosu

Recent innovations in multimodal action models represent a promising direction for developing general-purpose agentic systems, combining visual understanding, language comprehension, and action generation. We introduce MultiNet - a novel,…

机器学习 · 计算机科学 2025-06-18 Pranav Guruprasad , Yangyue Wang , Sudipta Chowdhury , Jaewoo Song , Harshvardhan Sikka

Causality helps people reason about and understand complex systems, particularly through what-if analyses that explore how interventions might alter outcomes. Although existing methods embrace causal reasoning using interventions and…

人机交互 · 计算机科学 2025-07-22 Yanming Zhang , Krishnakumar Hegde , Klaus Mueller

We address the problem of accurate capture of interactive behaviors between two people in daily scenarios. Most previous works either only consider one person or solely focus on conversational gestures of two people, assuming the body…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Leo Ho , Yinghao Huang , Dafei Qin , Mingyi Shi , Wangpok Tse , Wei Liu , Junichi Yamagishi , Taku Komura

Video agentic models have advanced challenging video-language tasks. However, most agentic approaches still heavily rely on greedy parsing over densely sampled video frames, resulting in high computational cost. We present VideoSeek, a…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Jingyang Lin , Jialian Wu , Jiang Liu , Ximeng Sun , Ze Wang , Xiaodong Yu , Jiebo Luo , Zicheng Liu , Emad Barsoum

In the era of generative AI, integrating video generation models into robotics opens new possibilities for the general-purpose robot agent. This paper introduces imitation learning with latent video planning (VILP). We propose a latent…

机器人学 · 计算机科学 2025-02-05 Zhengtong Xu , Qiang Qiu , Yu She