中文
相关论文

相关论文: Three-Stream Fusion Network for First-Person Inter…

200 篇论文

This paper presents a method to predict the future movements (location and gaze direction) of basketball players as a whole from their first person videos. The predicted behaviors reflect an individual physical space that affords to take…

计算机视觉与模式识别 · 计算机科学 2016-11-30 Shan Su , Jung Pyo Hong , Jianbo Shi , Hyun Soo Park

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Tuyen Tran , Thao Minh Le , Hung Tran , Truyen Tran

A comprehensive understanding of surgical scenes allows for monitoring of the surgical process, reducing the occurrence of accidents and enhancing efficiency for medical professionals. Semantic modeling within operating rooms, as a scene…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Diandian Guo , Manxi Lin , Jialun Pei , He Tang , Yueming Jin , Pheng-Ann Heng

This paper explores the development of a multimodal sentiment analysis model that integrates text, audio, and visual data to enhance sentiment classification. The goal is to improve emotion detection by capturing the complex interactions…

计算与语言 · 计算机科学 2025-01-15 Hui Lee , Singh Suniljit , Yong Siang Ong

Video object detection is more challenging compared to image object detection. Previous works proved that applying object detector frame by frame is not only slow but also inaccurate. Visual clues get weakened by defocus and motion blur,…

计算机视觉与模式识别 · 计算机科学 2017-12-19 Congrui Hetang , Hongwei Qin , Shaohui Liu , Junjie Yan

Spatio-temporal information is very important to capture the discriminative cues between genuine and fake faces from video sequences. To explore such a temporal feature, the fine-grained motions (e.g., eye blinking, mouth movements and head…

计算机视觉与模式识别 · 计算机科学 2019-01-18 Xiaoguang Tu , Hengsheng Zhang , Mei Xie , Yao Luo , Yuefei Zhang , Zheng Ma

In this paper, we introduce Coarse-Fine Networks, a two-stream architecture which benefits from different abstractions of temporal resolution to learn better video representations for long-term motion. Traditional Video models process…

计算机视觉与模式识别 · 计算机科学 2021-04-02 Kumara Kahatapitiya , Michael S. Ryoo

We present a new approach for video-driven animation of high-quality neural 3D head models, addressing the challenge of person-independent animation from video input. Typically, high-quality generative models are learned for specific…

计算机视觉与模式识别 · 计算机科学 2024-03-08 Wolfgang Paier , Paul Hinzer , Anna Hilsmann , Peter Eisert

We study multi-sensor fusion for 3D semantic segmentation that is important to scene understanding for many applications, such as autonomous driving and robotics. Existing fusion-based methods, however, may not achieve promising performance…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Mingkui Tan , Zhuangwei Zhuang , Sitao Chen , Rong Li , Kui Jia , Qicheng Wang , Yuanqing Li

This paper presents a concept of image pixel fusion of visual and thermal faces, which can significantly improve the overall performance of a face recognition system. Several factors affect face recognition performance including pose…

计算机视觉与模式识别 · 计算机科学 2010-07-06 Debotosh Bhattacharjee , Mrinal Kanti Bhowmik , Mita Nasipuri , Dipak Kumar Basu , Mahantapas Kundu

In this paper, we propose a novel approach to address the problem of camera and radar sensor fusion for 3D object detection in autonomous vehicle perception systems. Our approach builds on recent advances in deep learning and leverages the…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Daniel Dworak , Mateusz Komorkiewicz , Paweł Skruch , Jerzy Baranowski

We present a novel method for populating 3D indoor scenes with virtual humans that can navigate in the environment and interact with objects in a realistic manner. Existing approaches rely on training sequences that contain captured human…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Kaifeng Zhao , Yan Zhang , Shaofei Wang , Thabo Beeler , Siyu Tang

The objective of this work is human pose estimation in videos, where multiple frames are available. We investigate a ConvNet architecture that is able to benefit from temporal context by combining information across the multiple frames…

计算机视觉与模式识别 · 计算机科学 2015-11-10 Tomas Pfister , James Charles , Andrew Zisserman

Classifying videos according to content semantics is an important problem with a wide range of applications. In this paper, we propose a hybrid deep learning framework for video classification, which is able to model static spatial…

计算机视觉与模式识别 · 计算机科学 2015-04-08 Zuxuan Wu , Xi Wang , Yu-Gang Jiang , Hao Ye , Xiangyang Xue

Recent advances in deep learning have enabled the generation of videos from textual descriptions as well as the prediction of future sequences from input videos. Similarly, in human motion modeling, motions can be generated from text or…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Masato Soga , Ryuki Takebayashi

Facial action units (AUs) are essential to decode human facial expressions. Researchers have focused on training AU detectors with a variety of features and classifiers. However, several issues remain. These are spatial representation,…

计算机视觉与模式识别 · 计算机科学 2016-08-03 Wen-Sheng Chu , Fernando De la Torre , Jeffrey F. Cohn

Real-time multi-view point cloud reconstruction is a core problem in 3D vision and immersive perception, with wide applications in VR, AR, robotic navigation, digital twins, and computer interaction. Despite advances in multi-camera systems…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Chentian Sun

Scene flow describes the motion of 3D objects in real world and potentially could be the basis of a good feature for 3D action recognition. However, its use for action recognition, especially in the context of convolutional neural networks…

计算机视觉与模式识别 · 计算机科学 2017-03-28 Pichao Wang , Wanqing Li , Zhimin Gao , Yuyao Zhang , Chang Tang , Philip Ogunbona

Analysis of human interaction is one important research topic of human motion analysis. It has been studied either using first person vision (FPV) or third person vision (TPV). However, the joint learning of both types of vision has so far…

计算机视觉与模式识别 · 计算机科学 2022-09-22 Zihui Guo , Yonghong Hou , Pichao Wang , Zhimin Gao , Mingliang Xu , Wanqing Li

For the current 3D human pose estimation task, a group of methods mainly learn the rules of 2D-3D projection from spatial and temporal correlation. However, earlier methods model the global features of the entire body joint in the time…

计算机视觉与模式识别 · 计算机科学 2022-11-01 Xinwei Yu , Xiaohua Zhang