中文
相关论文

相关论文: GRASP: Learning to Ground Social Reasoning in Mult…

200 篇论文

Spatial reasoning, an important faculty of human cognition with many practical applications, is one of the core commonsense skills that is not purely language-based and, for satisfying (as opposed to optimal) solutions, requires some…

人工智能 · 计算机科学 2025-01-20 Zhisheng Tang , Mayank Kejriwal

Human annotation plays a core role in machine learning -- annotations for supervised models, safety guardrails for generative models, and human feedback for reinforcement learning, to cite a few avenues. However, the fact that many of these…

Large language models are increasingly deployed as automated judges to evaluate the strength of arguments. As this role expands, their legitimacy depends on consistency, transparency, and the ability to separate argumentative structure from…

机器学习 · 计算机科学 2026-05-20 Diganta Misra , Antonio Orvieto , Rediet Abebe , Volkan Cevher

Despite significant progress in robotic systems for operation within human-centric environments, existing models still heavily rely on explicit human commands to identify and manipulate specific objects. This limits their effectiveness in…

机器人学 · 计算机科学 2024-10-16 Shiyu Jin , Jinxuan Xu , Yutian Lei , Liangjun Zhang

This paper presents GRASP, a novel benchmark to evaluate the language grounding and physical understanding capabilities of video-based multimodal large language models (LLMs). This evaluation is accomplished via a two-tier approach…

计算与语言 · 计算机科学 2024-06-07 Serwan Jassim , Mario Holubar , Annika Richter , Cornelius Wolff , Xenia Ohmer , Elia Bruni

Robotic grasping is one of the most fundamental tasks in robotic manipulation, and grasp detection/generation has long been the subject of extensive research. Recently, language-driven grasp generation has emerged as a promising direction…

机器人学 · 计算机科学 2025-10-08 Haoran Zhang , Shuanghao Bai , Wanqi Zhou , Yuedi Zhang , Qi Zhang , Pengxiang Ding , Cheng Chi , Donglin Wang , Badong Chen

Generating natural human grasps necessitates consideration of not just object geometry but also semantic information. Solely depending on object shape for grasp generation confines the applications of prior methods in downstream tasks. This…

计算机视觉与模式识别 · 计算机科学 2024-04-05 Kailin Li , Jingbo Wang , Lixin Yang , Cewu Lu , Bo Dai

Social reasoning abilities are crucial for AI systems to effectively interpret and respond to multimodal human communication and interaction within social contexts. We introduce SOCIAL GENOME, the first benchmark for fine-grained, grounded…

计算与语言 · 计算机科学 2025-06-05 Leena Mathur , Marian Qian , Paul Pu Liang , Louis-Philippe Morency

The dialogue-based relation extraction (DialogRE) task aims to predict the relations between argument pairs that appear in dialogue. Most previous studies utilize fine-tuning pre-trained language models (PLMs) only with extensive features…

计算与语言 · 计算机科学 2022-09-28 Junyoung Son , Jinsung Kim , Jungwoo Lim , Heuiseok Lim

Traditional scene graphs primarily focus on spatial relationships, limiting vision-language models' (VLMs) ability to reason about complex interactions in visual scenes. This paper addresses two key challenges: (1) conventional…

计算机视觉与模式识别 · 计算机科学 2025-05-15 Dayong Liang , Changmeng Zheng , Zhiyuan Wen , Yi Cai , Xiao-Yong Wei , Qing Li

Large Language Models (LLMs) have substantially improved the conversational capabilities of social robots. Nevertheless, for an intuitive and fluent human-robot interaction, robots should be able to ground the conversation by relating…

人机交互 · 计算机科学 2026-04-09 Elisabeth Menendez , Michael Gienger , Santiago Martínez , Carlos Balaguer , Anna Belardinelli

Understanding social interaction, which encompasses perceiving numerous and subtle multimodal cues, inferring unobservable mental states and relations, and dynamically predicting others' behavior, is the foundation for achieving…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Fanqi Kong , Weiqin Zu , Xinyu Chen , Yaodong Yang , Song-Chun Zhu , Xue Feng

Quantitative Systems Pharmacology (QSP) modeling is essential for drug development but it requires significant time investment that limits the throughput of domain experts. We present \textbf{GRASP} -- a multi-agent, graph-reasoning…

机器学习 · 计算机科学 2025-12-08 Omid Bazgir , Vineeth Manthapuri , Ilia Rattsev , Mohammad Jafarnejad

Group Activity Detection (GAD) involves recognizing social groups and their collective behaviors in videos. Vision Foundation Models (VFMs), like DINOv2, offer excellent features but are pretrained on object-centric data. We find that…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Thinesh Thiyakesan Ponbagavathi , Chengzheng Yang , Alina Roitberg

Agentic retrieval improves multi-hop question answering by giving language models autonomy to iteratively gather evidence. Recent work augments these systems with knowledge graphs for structured traversal, but this combination introduces…

多智能体系统 · 计算机科学 2026-05-19 Stockton Jenkins , Ramya Korlakai Vinayak , Junjie Hu

Robotic grasping is a fundamental aspect of robot functionality, defining how robots interact with objects. Despite substantial progress, its generalizability to counter-intuitive or long-tailed scenarios, such as objects with uncommon…

机器人学 · 计算机科学 2024-02-27 Dingkun Guo , Yuqi Xiang , Shuqi Zhao , Xinghao Zhu , Masayoshi Tomizuka , Mingyu Ding , Wei Zhan

Discovering social relations in images can make machines better interpret the behavior of human beings. However, automatically recognizing social relations in images is a challenging task due to the significant gap between the domains of…

计算机视觉与模式识别 · 计算机科学 2019-01-11 Meng Zhang , Xinchen Liu , Wu Liu , Anfu Zhou , Huadong Ma , Tao Mei

Multimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes. To bridge this gap, Video Scene Graph Generation (VidSGG) has emerged to capture multi-object relationships across video frames.…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Trong-Thuan Nguyen , Pha Nguyen , Jackson Cothren , Alper Yilmaz , Khoa Luu

Answering complex questions about textual narratives requires reasoning over both stated context and the world knowledge that underlies it. However, pretrained language models (LM), the foundation of most modern QA systems, do not robustly…

Human gaze provides essential cues for interpreting attention, intention, and social interaction in visual scenes, yet gaze understanding remains largely unexplored in current vision-language models (VLMs). While recent VLMs achieve strong…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Shijing Wang , Chaoqun Cui , Yaping Huang , Hyung Jin Chang , Yihua Cheng
‹ 上一页 1 2 3 10 下一页 ›