English
Related papers

Related papers: PCIE_Interaction Solution for Ego4D Social Interac…

200 papers

Human actions involving hand manipulations are structured according to the making and breaking of hand-object contact, and human visual understanding of action is reliant on anticipation of contact as is demonstrated by pioneering work in…

Computer Vision and Pattern Recognition · Computer Science 2021-02-02 Eadom Dessalene , Chinmaya Devaraj , Michael Maynord , Cornelia Fermuller , Yiannis Aloimonos

In this report, we present our approach and empirical results of applying masked autoencoders in two egocentric video understanding tasks, namely, Object State Change Classification and PNR Temporal Localization, of Ego4D Challenge 2022. As…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Jiachen Lei , Shuang Ma , Zhongjie Ba , Sai Vemprala , Ashish Kapoor , Kui Ren

Large Vision-Language Models (VLMs) have demonstrated impressive performance on complex tasks involving visual input with natural language instructions. However, it remains unclear to what extent capabilities on natural images transfer to…

Computation and Language · Computer Science 2024-02-01 Chenhui Zhang , Sherrie Wang

In 2025, Large Language Model (LLM) services have launched a new feature -- AI video chat -- allowing users to interact with AI agents via real-time video communication (RTC), just like chatting with real people. Despite its significance,…

Networking and Internet Architecture · Computer Science 2025-10-02 Jiayang Xu , Xiangjie Huang , Zijie Li , Zili Meng

Alignment research on large language models (LLMs) increasingly depends on understanding how these systems are used in everyday contexts. Yet naturalistic interaction data is difficult to access due to privacy constraints and platform…

Human-Computer Interaction · Computer Science 2026-03-24 Cathy Mengying Fang , Sheer Karny , Chayapatr Archiwaranguprok , Yasith Samaradivakara , Pat Pataranutaporn , Pattie Maes

This study explores the capabilities of multimodal large language models (LLMs) in handling challenging multistep tasks that integrate language and vision, focusing on model steerability, composability, and the application of long-term…

Artificial Intelligence · Computer Science 2023-12-20 David Noever , Samantha Elizabeth Miller Noever

In current text-based task-oriented dialogue (TOD) systems, user emotion detection (ED) is often overlooked or is typically treated as a separate and independent task, requiring additional training. In contrast, our work demonstrates that…

Computation and Language · Computer Science 2024-07-01 Armand Stricker , Patrick Paroubek

Numerous methods have been proposed to detect, estimate, and analyze properties of people in images, including 3D pose, shape, contact, human-object interaction, and emotion. While widely applicable in vision and other areas, such methods…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Jing Lin , Yao Feng , Weiyang Liu , Michael J. Black

We present a method that automatically evaluates emotional response from spontaneous facial activity recorded by a depth camera. The automatic evaluation of emotional response, or affect, is a fascinating challenge with many applications,…

Human-Computer Interaction · Computer Science 2017-01-20 Daniel Hadar

Recovering high-quality 3D human motion in complex scenes from monocular videos is important for many applications, ranging from AR/VR to robotics. However, capturing realistic human-scene interactions, while dealing with occlusions and…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Siwei Zhang , Yan Zhang , Federica Bogo , Marc Pollefeys , Siyu Tang

AI models have made significant strides in recent years in their ability to describe and answer questions about real-world images. They have also made progress in the ability to converse with users in real-time using audio input. This…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Reza Pourreza , Rishit Dagli , Apratim Bhattacharyya , Sunny Panchal , Guillaume Berger , Roland Memisevic

Person re-identification (Re-ID) is a crucial task in computer vision, aiming to recognize individuals across non-overlapping camera views. While recent advanced vision-language models (VLMs) excel in logical reasoning and multi-task…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Ke Niu , Haiyang Yu , Mengyang Zhao , Teng Fu , Siyang Yi , Wei Lu , Bin Li , Xuelin Qian , Xiangyang Xue

We propose a new approach to determine correspondences between image pairs in the wild under large changes in illumination, viewpoint, context, and material. While other approaches find correspondences between pairs of images by treating…

Computer Vision and Pattern Recognition · Computer Science 2021-03-29 Olivia Wiles , Sebastien Ehrhardt , Andrew Zisserman

In this paper, we present our solutions for a spectrum of automation tasks in life-saving intervention procedures within the Trauma THOMPSON (T3) Challenge, encompassing action recognition, action anticipation, and Visual Question Answering…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Trinh T. L. Vuong , Doanh C. Bui , Jin Tae Kwak

Human-swarm interaction has recently gained attention due to its plethora of new applications in disaster relief, surveillance, rescue, and exploration. However, if the task difficulty increases, the performance of the human operator…

Human-Computer Interaction · Computer Science 2021-09-27 Joseph P. Distefano , Hemanth Manjunatha , Souma Chowdhury , Karthik Dantu , David Doermann , Ehsan T. Esfahani

Face recognition in collaborative learning videos presents many challenges. In collaborative learning videos, students sit around a typical table at different positions to the recording camera, come and go, move around, get partially or…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Phuong Tran , Marios Pattichis , Sylvia Celedón-Pattichis , Carlos LópezLeiva

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Haoyu Zhang , Qiaohui Chu , Meng Liu , Haoxiang Shi , Yaowei Wang , Liqiang Nie

The rapid advancement of Text-guided Image Editing (TIE) enables image modifications through text prompts. However, current TIE models still struggle to balance image quality, editing alignment, and consistency with the original image,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Zitong Xu , Huiyu Duan , Bingnan Liu , Guangji Ma , Jiarui Wang , Liu Yang , Shiqi Gao , Xiaoyu Wang , Jia Wang , Xiongkuo Min , Guangtao Zhai , Weisi Lin

Smart glass is emerging as an useful device since it provides plenty of insights under hands-busy, eyes-on-task situations. To understand the context of the wearer, 6D object pose estimation in egocentric view is becoming essential.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Taegyoon Yoon , Yegyu Han , Seojin Ji , Jaewoo Park , Sojeong Kim , Taein Kwon , Hyung-Sin Kim

Temporal action proposal generation (TAPG) is a challenging task, which requires localizing action intervals in an untrimmed video. Intuitively, we as humans, perceive an action through the interactions between actors, relevant objects, and…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Khoa Vo , Sang Truong , Kashu Yamazaki , Bhiksha Raj , Minh-Triet Tran , Ngan Le