English
Related papers

Related papers: PCIE_Interaction Solution for Ego4D Social Interac…

200 papers

People control their bodies to establish contact with the environment. To comprehensively understand actions across diverse visual contexts, it is essential to simultaneously consider \textbf{what} action is occurring and \textbf{where} it…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yuxiao Wang , Yu Lei , Wolin Liang , Weiying Xue , Zhenao Wei , Nan Zhuang , Qi Liu

Effective collaboration requires groups to strategically regulate themselves to overcome challenges. Research has shown that groups may fail to regulate due to differences in members' perceptions of challenges which may benefit from…

Computation and Language · Computer Science 2024-01-04 Wannapon Suraworachet , Jennifer Seon , Mutlu Cukurova

Face performance capture and reenactment techniques use multiple cameras and sensors, positioned at a distance from the face or mounted on heavy wearable devices. This limits their applications in mobile and outdoor environments. We present…

Computer Vision and Pattern Recognition · Computer Science 2019-05-28 Mohamed Elgharib , Mallikarjun BR , Ayush Tewari , Hyeongwoo Kim , Wentao Liu , Hans-Peter Seidel , Christian Theobalt

Text-image cross-modal retrieval is a challenging task in the field of language and vision. Most previous approaches independently embed images and sentences into a joint embedding space and compare their similarities. However, previous…

Computer Vision and Pattern Recognition · Computer Science 2019-09-13 Zihao Wang , Xihui Liu , Hongsheng Li , Lu Sheng , Junjie Yan , Xiaogang Wang , Jing Shao

Generating captions for long and complex videos is both critical and challenging, with significant implications for the growing fields of text-to-video generation and multi-modal understanding. One key challenge in long video captioning is…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Zhantao Yang , Huangji Wang , Ruili Feng , Han Zhang , Yuting Hu , Shangwen Zhu , Junyan Li , Yu Liu , Fan Cheng

Large Language Models (LLMs) are increasingly explored for educational tasks such as grading, yet their alignment with human evaluation in real classrooms remains underexamined. In this study, we investigate the feasibility of using an LLM…

Computation and Language · Computer Science 2025-11-19 Grace Byun , Swati Rajwal , Jinho D. Choi

This report presents our Le3DE2E_Occ solution for 4D Occupancy Forecasting in Argoverse Challenges at CVPR 2023 Workshop on Autonomous Driving (WAD). Our solution consists of a strong LiDAR-based Bird's Eye View (BEV) encoder with temporal…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Pengfei Zheng , Kanokphan Lertniphonphan , Feng Chen , Siwei Chen , Bingchuan Sun , Jun Xie , Zhepeng Wang

Identifying different objects (man and cup) is an important problem on its own, but identifying the relationship between them (holding) is critical for many real world use cases. This paper describes an approach to reduce a visual…

Computer Vision and Pattern Recognition · Computer Science 2018-09-27 Toshiyuki Fukuzawa

Generic Event Boundary Detection (GEBD) is a newly introduced task that aims to detect "general" event boundaries that correspond to natural human perception. In this paper, we introduce a novel contrastive learning based approach to deal…

Computer Vision and Pattern Recognition · Computer Science 2021-06-23 Hyolim Kang , Jinwoo Kim , Kyungmin Kim , Taehyun Kim , Seon Joo Kim

Accurate recognition of human emotions is critical for adaptive human-computer interaction, yet remains challenging in dynamic, conversation-like settings. This work presents a personality-aware multimodal framework that integrates…

Eye contact is a crucial non-verbal interaction modality and plays an important role in our everyday social life. While humans are very sensitive to eye contact, the capabilities of machines to capture a person's gaze are still mediocre. We…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Thorsten Hempel , Magnus Jung , Ahmed A. Abdelrahman , Ayoub Al-Hamadi

In this report, we present the ReLER@ZJU1 submission to the Ego4D Moment Queries Challenge in ECCV 2022. In this task, the goal is to retrieve and localize all instances of possible activities in egocentric videos. Ego4D dataset is…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Jiayi Shao , Xiaohan Wang , Yi Yang

Large vision-language models (LVLMs), while proficient in following instructions and responding to diverse questions, invariably generate detailed responses even when questions are ambiguous or unanswerable, leading to hallucinations and…

Computation and Language · Computer Science 2024-10-21 Shujin Wu , Yi R. Fung , Sha Li , Yixin Wan , Kai-Wei Chang , Heng Ji

We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visual streams. Our approach couples masked visual modeling with…

Artificial Intelligence · Computer Science 2026-01-07 Zeyu Ling , Xiaodong Gu , Jiangnan Tang , Changqing Zou

Large Language Models (LLMs) have so far impressed the world, with unprecedented capabilities that emerge in models at large scales. On the vision side, transformer models (i.e., ViT) are following the same trend, achieving the best…

Computer Vision and Pattern Recognition · Computer Science 2023-10-30 Mustafa Shukor , Corentin Dancette , Matthieu Cord

Artificial agents that support human group interactions hold great promise, especially in sensitive contexts such as well-being promotion and therapeutic interventions. However, current systems struggle to mediate group interactions…

Human-Computer Interaction · Computer Science 2026-03-17 Giulia Huang , Maristella Matera , Micol Spitale

This study highlights the potential of ChatGPT (specifically GPT-4o) as a competitive alternative for Face Presentation Attack Detection (PAD), outperforming several PAD models, including commercial solutions, in specific scenarios. Our…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Alain Komaty , Hatef Otroshi Shahreza , Anjith George , Sebastien Marcel

Analyzing instructional interactions between an instructor and a learner who are co-present in the same physical space is a critical problem for educational support and skill transfer. Yet such face-to-face instructional scenes have not…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Yuki Sakai , Ryosuke Furuta , Juichun Yen , Yoichi Sato

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D dataset to benchmark the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Yuxuan Li , Vijay Veerabadran , Michael L. Iuzzolino , Brett D. Roads , Asli Celikyilmaz , Karl Ridgeway

Touch contact and pressure are essential for understanding how humans interact with and manipulate objects, insights which can significantly benefit applications in mixed reality and robotics. However, estimating these interactions from an…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Yiming Zhao , Taein Kwon , Paul Streli , Marc Pollefeys , Christian Holz