English
Related papers

Related papers: Detecting speaking persons in video

200 papers

This paper describes a novel face identification method that combines the eigenfaces theory with the Neural Nets. We use the eigenfaces methodology in order to reduce the dimensionality of the input image, and a neural net classifier that…

Computer Vision and Pattern Recognition · Computer Science 2022-04-04 Virginia Espinosa-Duro , Marcos Faundez-Zanuy

This work discusses and implements the application of speaker recognition for the detection of collaborations in YouTube videos. CATANA, an existing framework for detection and analysis of YouTube collaborations, is utilizing face…

Computer Vision and Pattern Recognition · Computer Science 2018-07-06 Moritz Lode , Michael Örtl , Christian Koch , Amr Rizk , Ralf Steinmetz

Speechreading is the task of inferring phonetic information from visually observed articulatory facial movements, and is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a…

Computer Vision and Pattern Recognition · Computer Science 2017-08-31 Ariel Ephrat , Tavi Halperin , Shmuel Peleg

Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to…

Computer Vision and Pattern Recognition · Computer Science 2021-08-29 Thanh-Dat Truong , Chi Nhan Duong , The De Vu , Hoang Anh Pham , Bhiksha Raj , Ngan Le , Khoa Luu

In this paper, we propose a methodology for early recognition of human activities from videos taken with a first-person viewpoint. Early recognition, which is also known as activity prediction, is an ability to infer an ongoing activity at…

Computer Vision and Pattern Recognition · Computer Science 2015-07-07 M. S. Ryoo , Thomas J. Fuchs , Lu Xia , J. K. Aggarwal , Larry Matthies

The task of searching certain people in videos has seen increasing potential in real-world applications, such as video organization and editing. Most existing approaches are devised to work in an offline manner, where identities can only be…

Computer Vision and Pattern Recognition · Computer Science 2020-08-11 Jiangyue Xia , Anyi Rao , Qingqiu Huang , Linning Xu , Jiangtao Wen , Dahua Lin

In this paper, we address the problem of spatio-temporal person retrieval from multiple videos using a natural language query, in which we output a tube (i.e., a sequence of bounding boxes) which encloses the person described by the query.…

Computer Vision and Pattern Recognition · Computer Science 2017-08-24 Masataka Yamaguchi , Kuniaki Saito , Yoshitaka Ushiku , Tatsuya Harada

The widespread use of cameras in everyday life situations generates a vast amount of data that may contain sensitive information about the people and vehicles moving in front of them (location, license plates, physical characteristics,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Roman Plaud , Jose-Luis Lisani

Current methods for active speak er detection focus on modeling short-term audiovisual information from a single speaker. Although this strategy can be enough for addressing single-speaker scenarios, it prevents accurate detection when the…

Computer Vision and Pattern Recognition · Computer Science 2020-05-21 Juan Leon Alcazar , Fabian Caba Heilbron , Long Mai , Federico Perazzi , Joon-Young Lee , Pablo Arbelaez , Bernard Ghanem

Given an audio clip and a reference face image, the goal of the talking head generation is to generate a high-fidelity talking head video. Although some audio-driven methods of generating talking head videos have made some achievements in…

Computer Vision and Pattern Recognition · Computer Science 2023-06-07 Jianrong Wang , Yaxin Zhao , Li Liu , Tianyi Xu , Qi Li , Sen Li

The mechanism proposed here is for real-time speaker change detection in conversations, which firstly trains a neural network text-independent speaker classifier using in-domain speaker data. Through the network, features of conversational…

Sound · Computer Science 2017-03-20 Zhenhao Ge , Ananth N. Iyer , Srinath Cheluvaraja , Aravind Ganapathiraju

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and eye blinks.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Dongchan Min , Minyoung Song , Eunji Ko , Sung Ju Hwang

The objective of this work is speaker diarisation of speech recordings 'in the wild'. The ability to determine speech segments is a crucial part of diarisation systems, accounting for a large proportion of errors. In this paper, we present…

Sound · Computer Science 2020-12-01 Youngki Kwon , Hee Soo Heo , Jaesung Huh , Bong-Jin Lee , Joon Son Chung

The topic of facial landmark detection has been widely covered for pictures of human faces, but it is still a challenge for drawings. Indeed, the proportions and symmetry of standard human faces are not always used for comics or mangas. The…

Computer Vision and Pattern Recognition · Computer Science 2018-11-09 Marco Stricker , Olivier Augereau , Koichi Kise , Motoi Iwata

Existing deep learning based facial landmark detection methods have achieved excellent performance. These methods, however, do not explicitly embed the structural dependencies among landmark points. They hence cannot preserve the geometric…

Computer Vision and Pattern Recognition · Computer Science 2022-06-29 Lisha Chen , Hui Su , Qiang Ji

Speech-driven facial animation is the process that automatically synthesizes talking characters based on speech signals. The majority of work in this domain creates a mapping from audio features to visual features. This approach often…

Computer Vision and Pattern Recognition · Computer Science 2019-06-18 Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

Video event extraction aims to detect salient events from a video and identify the arguments for each event as well as their semantic roles. Existing methods focus on capturing the overall visual scene of each frame, ignoring fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Guang Yang , Manling Li , Jiajie Zhang , Xudong Lin , Shih-Fu Chang , Heng Ji

In this paper, we present a novel benchmark for Emotion Recognition using facial landmarks extracted from realistic news videos. Traditional methods relying on RGB images are resource-intensive, whereas our approach with Facial Landmark…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Qixuan Zhang , Zhifeng Wang , Yang Liu , Zhenyue Qin , Kaihao Zhang , Sabrina Caldwell , Tom Gedeon

The ability to classify spoken speech based on the style of speaking is an important problem. With the advent of BPO's in recent times, specifically those that cater to a population other than the local population, it has become necessary…

Computation and Language · Computer Science 2015-04-08 Sunil Kopparapu , Saurabh Bhatnagar , K. Sahana , Sathyanarayana , Akhilesh Srivastava , P. V. S. Rao

Crowd counting problem aims to count the number of objects within an image or a frame in the videos and is usually solved by estimating the density map generated from the object location annotations. The values in the density map, by…

Computer Vision and Pattern Recognition · Computer Science 2019-06-21 Shengqin Jiang , Xiaobo Lu , Yinjie Lei , Lingqiao Liu
‹ Prev 1 4 5 6 7 8 10 Next ›