中文
相关论文

相关论文: PCIE_Interaction Solution for Ego4D Social Interac…

200 篇论文

People control their bodies to establish contact with the environment. To comprehensively understand actions across diverse visual contexts, it is essential to simultaneously consider \textbf{what} action is occurring and \textbf{where} it…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yuxiao Wang , Yu Lei , Wolin Liang , Weiying Xue , Zhenao Wei , Nan Zhuang , Qi Liu

Effective collaboration requires groups to strategically regulate themselves to overcome challenges. Research has shown that groups may fail to regulate due to differences in members' perceptions of challenges which may benefit from…

计算与语言 · 计算机科学 2024-01-04 Wannapon Suraworachet , Jennifer Seon , Mutlu Cukurova

Face performance capture and reenactment techniques use multiple cameras and sensors, positioned at a distance from the face or mounted on heavy wearable devices. This limits their applications in mobile and outdoor environments. We present…

计算机视觉与模式识别 · 计算机科学 2019-05-28 Mohamed Elgharib , Mallikarjun BR , Ayush Tewari , Hyeongwoo Kim , Wentao Liu , Hans-Peter Seidel , Christian Theobalt

Text-image cross-modal retrieval is a challenging task in the field of language and vision. Most previous approaches independently embed images and sentences into a joint embedding space and compare their similarities. However, previous…

计算机视觉与模式识别 · 计算机科学 2019-09-13 Zihao Wang , Xihui Liu , Hongsheng Li , Lu Sheng , Junjie Yan , Xiaogang Wang , Jing Shao

Generating captions for long and complex videos is both critical and challenging, with significant implications for the growing fields of text-to-video generation and multi-modal understanding. One key challenge in long video captioning is…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Zhantao Yang , Huangji Wang , Ruili Feng , Han Zhang , Yuting Hu , Shangwen Zhu , Junyan Li , Yu Liu , Fan Cheng

Large Language Models (LLMs) are increasingly explored for educational tasks such as grading, yet their alignment with human evaluation in real classrooms remains underexamined. In this study, we investigate the feasibility of using an LLM…

计算与语言 · 计算机科学 2025-11-19 Grace Byun , Swati Rajwal , Jinho D. Choi

This report presents our Le3DE2E_Occ solution for 4D Occupancy Forecasting in Argoverse Challenges at CVPR 2023 Workshop on Autonomous Driving (WAD). Our solution consists of a strong LiDAR-based Bird's Eye View (BEV) encoder with temporal…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Pengfei Zheng , Kanokphan Lertniphonphan , Feng Chen , Siwei Chen , Bingchuan Sun , Jun Xie , Zhepeng Wang

Identifying different objects (man and cup) is an important problem on its own, but identifying the relationship between them (holding) is critical for many real world use cases. This paper describes an approach to reduce a visual…

计算机视觉与模式识别 · 计算机科学 2018-09-27 Toshiyuki Fukuzawa

Generic Event Boundary Detection (GEBD) is a newly introduced task that aims to detect "general" event boundaries that correspond to natural human perception. In this paper, we introduce a novel contrastive learning based approach to deal…

计算机视觉与模式识别 · 计算机科学 2021-06-23 Hyolim Kang , Jinwoo Kim , Kyungmin Kim , Taehyun Kim , Seon Joo Kim

Accurate recognition of human emotions is critical for adaptive human-computer interaction, yet remains challenging in dynamic, conversation-like settings. This work presents a personality-aware multimodal framework that integrates…

Eye contact is a crucial non-verbal interaction modality and plays an important role in our everyday social life. While humans are very sensitive to eye contact, the capabilities of machines to capture a person's gaze are still mediocre. We…

计算机视觉与模式识别 · 计算机科学 2023-11-09 Thorsten Hempel , Magnus Jung , Ahmed A. Abdelrahman , Ayoub Al-Hamadi

In this report, we present the ReLER@ZJU1 submission to the Ego4D Moment Queries Challenge in ECCV 2022. In this task, the goal is to retrieve and localize all instances of possible activities in egocentric videos. Ego4D dataset is…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Jiayi Shao , Xiaohan Wang , Yi Yang

Large vision-language models (LVLMs), while proficient in following instructions and responding to diverse questions, invariably generate detailed responses even when questions are ambiguous or unanswerable, leading to hallucinations and…

计算与语言 · 计算机科学 2024-10-21 Shujin Wu , Yi R. Fung , Sha Li , Yixin Wan , Kai-Wei Chang , Heng Ji

We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visual streams. Our approach couples masked visual modeling with…

人工智能 · 计算机科学 2026-01-07 Zeyu Ling , Xiaodong Gu , Jiangnan Tang , Changqing Zou

Large Language Models (LLMs) have so far impressed the world, with unprecedented capabilities that emerge in models at large scales. On the vision side, transformer models (i.e., ViT) are following the same trend, achieving the best…

计算机视觉与模式识别 · 计算机科学 2023-10-30 Mustafa Shukor , Corentin Dancette , Matthieu Cord

Artificial agents that support human group interactions hold great promise, especially in sensitive contexts such as well-being promotion and therapeutic interventions. However, current systems struggle to mediate group interactions…

人机交互 · 计算机科学 2026-03-17 Giulia Huang , Maristella Matera , Micol Spitale

This study highlights the potential of ChatGPT (specifically GPT-4o) as a competitive alternative for Face Presentation Attack Detection (PAD), outperforming several PAD models, including commercial solutions, in specific scenarios. Our…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Alain Komaty , Hatef Otroshi Shahreza , Anjith George , Sebastien Marcel

Analyzing instructional interactions between an instructor and a learner who are co-present in the same physical space is a critical problem for educational support and skill transfer. Yet such face-to-face instructional scenes have not…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Yuki Sakai , Ryosuke Furuta , Juichun Yen , Yoichi Sato

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D dataset to benchmark the…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Yuxuan Li , Vijay Veerabadran , Michael L. Iuzzolino , Brett D. Roads , Asli Celikyilmaz , Karl Ridgeway

Touch contact and pressure are essential for understanding how humans interact with and manipulate objects, insights which can significantly benefit applications in mixed reality and robotics. However, estimating these interactions from an…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Yiming Zhao , Taein Kwon , Paul Streli , Marc Pollefeys , Christian Holz