English
Related papers

Related papers: Detecting speaking persons in video

200 papers

A major challenge in DeepFake forgery detection is that state-of-the-art algorithms are mostly trained to detect a specific fake method. As a result, these approaches show poor generalization across different types of facial manipulations,…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Davide Cozzolino , Andreas Rössler , Justus Thies , Matthias Nießner , Luisa Verdoliva

Speaker recognition is a task of identifying persons from their voices. Recently, deep learning has dramatically revolutionized speaker recognition. However, there is lack of comprehensive reviews on the exciting progress. In this paper, we…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-06 Zhongxin Bai , Xiao-Lei Zhang

Novel text-to-speech systems can generate entirely new voices that were not seen during training. However, it remains a difficult task to efficiently create personalized voices from a high-dimensional speaker space. In this work, we use…

We propose a method for human action recognition, one that can localize the spatiotemporal regions that `define' the actions. This is a challenging task due to the subtlety of human actions in video and the co-occurrence of contextual…

Computer Vision and Pattern Recognition · Computer Science 2019-04-12 Yang Wang , Vinh Tran , Gedas Bertasius , Lorenzo Torresani , Minh Hoai

Human speech is often accompanied by hand and arm gestures. Given audio speech input, we generate plausible gestures to go along with the sound. Specifically, we perform cross-modal translation from "in-the-wild'' monologue speech of a…

Computer Vision and Pattern Recognition · Computer Science 2019-06-11 Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , Jitendra Malik

We present an approach to labeling short video clips with English verbs as event descriptions. A key distinguishing aspect of this work is that it labels videos with verbs that describe the spatiotemporal interaction between event…

Emotion detection in conversations is a necessary step for a number of applications, including opinion mining over chat history, social media threads, debates, argumentation mining, understanding consumer feedback in live conversations,…

Computation and Language · Computer Science 2019-05-28 Navonil Majumder , Soujanya Poria , Devamanyu Hazarika , Rada Mihalcea , Alexander Gelbukh , Erik Cambria

An essential goal of computational media intelligence is to support understanding how media stories -- be it news, commercial or entertainment media -- represent and reflect society and these portrayals are perceived. People are a central…

Computer Vision and Pattern Recognition · Computer Science 2022-03-23 Rahul Sharma , Shrikanth Narayanan

Personas are useful for dialogue response prediction. However, the personas used in current studies are pre-defined and hard to obtain before a conversation. To tackle this issue, we study a new task, named Speaker Persona Detection (SPD),…

Computation and Language · Computer Science 2021-09-06 Jia-Chen Gu , Zhen-Hua Ling , Yu Wu , Quan Liu , Zhigang Chen , Xiaodan Zhu

Due to a drastic improvement in the quality of internet services worldwide, there is an explosion of multilingual content generation and consumption. This is especially prevalent in countries with large multilingual audience, who are…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-14 Mudit Verma , Arun Balaji Buduru

Deep Neural Networks (DNNs) have shown to outperform traditional methods in various visual recognition tasks including Facial Expression Recognition (FER). In spite of efforts made to improve the accuracy of FER systems using DNN, existing…

Computer Vision and Pattern Recognition · Computer Science 2020-04-17 Behzad Hasani , Mohammad H. Mahoor

This paper proposes dynamic human group detection in videos. For detecting complex groups, not only the local appearance features of in-group members but also the global context of the scene are important. Such local and global appearance…

Computer Vision and Pattern Recognition · Computer Science 2025-09-08 Kaname Yokoyama , Chihiro Nakatani , Norimichi Ukita

In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling audio-visual synchronization cue, which depends on the video…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-13 Yidi Jiang , Ruijie Tao , Zexu Pan , Haizhou Li

The deaf and hard of hearing community relies on American Sign Language (ASL) as their primary mode of communication, but communication with others who do not know ASL can be difficult, especially during emergencies where no interpreter is…

Image and Video Processing · Electrical Eng. & Systems 2023-05-12 Janice Nguyen , Y. Curtis Wang

Each smile is unique: one person surely smiles in different ways (e.g., closing/opening the eyes or mouth). Given one input image of a neutral face, can we generate multiple smile videos with distinctive characteristics? To tackle this…

Computer Vision and Pattern Recognition · Computer Science 2018-03-29 Wei Wang , Xavier Alameda-Pineda , Dan Xu , Pascal Fua , Elisa Ricci , Nicu Sebe

We present a new task that predicts future locations of people observed in first-person videos. Consider a first-person video stream continuously recorded by a wearable camera. Given a short clip of a person that is extracted from the…

Computer Vision and Pattern Recognition · Computer Science 2018-03-29 Takuma Yagi , Karttikeya Mangalam , Ryo Yonetani , Yoichi Sato

With the widespread use of intelligent systems, such as smart speakers, addressee recognition has become a concern in human-computer interaction, as more and more people expect such systems to understand complicated social scenes, including…

Artificial Intelligence · Computer Science 2018-09-13 Thao Minh Le , Nobuyuki Shimizu , Takashi Miyazaki , Koichi Shinoda

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

Speaker Diarization is the problem of separating speakers in an audio. There could be any number of speakers and final result should state when speaker starts and ends. In this project, we analyze given audio file with 2 channels and 2…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-11 Vishal Sharma , Zekun Zhang , Zachary Neubert , Curtis Dyreson

Given the vast amounts of video available online, and recent breakthroughs in object detection with static images, object detection in video offers a promising new frontier. However, motion blur and compression artifacts cause substantial…

Computer Vision and Pattern Recognition · Computer Science 2016-07-20 Subarna Tripathi , Zachary C. Lipton , Serge Belongie , Truong Nguyen