English
Related papers

Related papers: BIAS: A Body-based Interpretable Active Speaker Ap…

200 papers

Speech recognition is very challenging in student learning environments that are characterized by significant cross-talk and background noise. To address this problem, we present a bilingual speech recognition system that uses an…

We introduce Brain-Artificial Intelligence Interfaces (BAIs) as a new class of Brain-Computer Interfaces (BCIs). Unlike conventional BCIs, which rely on intact cognitive capabilities, BAIs leverage the power of artificial intelligence to…

Human-Computer Interaction · Computer Science 2024-03-18 Anja Meunier , Michal Robert Žák , Lucas Munz , Sofiya Garkot , Manuel Eder , Jiachen Xu , Moritz Grosse-Wentrup

Voice Activity Detection (VAD) is an important pre-processing step in a wide variety of speech processing systems. VAD should in a practical application be able to detect speech in both noisy and noise-free environments, while not…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-06 Claus Meyer Larsen , Peter Koch , Zheng-Hua Tan

Speech signals are subjected to more acoustic interference and emotional factors than other signals. Noisy emotion-riddled speech data is a challenge for real-time speech processing applications. It is essential to find an effective way to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-25 Shibani Hamsa , Ismail Shahin , Youssef Iraqi , Ernesto Damiani , Naoufel Werghi

Video-to-speech (V2S) synthesis, the task of generating speech directly from silent video input, is inherently more challenging than other speech synthesis tasks due to the need to accurately reconstruct both speech content and speaker…

Sound · Computer Science 2025-03-10 Yifan Liu , Yu Fang , Zhouhan Lin

General accent recognition (AR) models tend to directly extract low-level information from spectrums, which always significantly overfit on speakers or channels. Considering accent can be regarded as a series of shifts relative to native…

Sound · Computer Science 2022-07-04 Qijie Shao , Jinghao Yan , Jian Kang , Pengcheng Guo , Xian Shi , Pengfei Hu , Lei Xie

Correctly recognizing the behaviors of children with Autism Spectrum Disorder (ASD) is of vital importance for the diagnosis of Autism and timely early intervention. However, the observation and recording during the treatment from the…

Computer Vision and Pattern Recognition · Computer Science 2024-01-08 Andong Deng , Taojiannan Yang , Chen Chen , Qian Chen , Leslie Neely , Sakiko Oyama

Overlapped speech detection (OSD) is critical for speech applications in scenario of multi-party conversion. Despite numerous research efforts and progresses, comparing with speech activity detection (VAD), OSD remains an open challenge and…

Sound · Computer Science 2022-09-27 Ziqing Du , Kai Liu , Xucheng Wan , Huan Zhou

Analyzing individual emotions during group conversation is crucial in developing intelligent agents capable of natural human-machine interaction. While reliable emotion recognition techniques depend on different modalities (text, audio,…

Automatic speech-based affect recognition of individuals in dyadic conversation is a challenging task, in part because of its heavy reliance on manual pre-processing. Traditional approaches frequently require hand-crafted speech features…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-24 Huili Chen , Yue Zhang , Felix Weninger , Rosalind Picard , Cynthia Breazeal , Hae Won Park

In this paper, we propose a novel method for speaker adaptation in lip reading, motivated by two observations. Firstly, a speaker's own characteristics can always be portrayed well by his/her few facial images or even a single image with…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Songtao Luo , Shuang Yang , Shiguang Shan , Xilin Chen

To build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space. The articulatory space is a promising…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-26 Peter Wu , Li-Wei Chen , Cheol Jun Cho , Shinji Watanabe , Louis Goldstein , Alan W Black , Gopala K. Anumanchipalli

Although neural rendering has made significant advances in creating lifelike, animatable full-body and head avatars, incorporating detailed expressions into full-body avatars remains largely unexplored. We present DEGAS, the first 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Zhijing Shao , Duotun Wang , Qing-Yao Tian , Yao-Dong Yang , Hengyu Meng , Zeyu Cai , Bo Dong , Yu Zhang , Kang Zhang , Zeyu Wang

We propose VASA-3D, an audio-driven, single-shot 3D head avatar generator. This research tackles two major challenges: capturing the subtle expression details present in real human faces, and reconstructing an intricate 3D head avatar from…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Sicheng Xu , Guojun Chen , Jiaolong Yang , Yizhong Zhang , Yu Deng , Steve Lin , Baining Guo

In this paper, we introduce an end-to-end machine learning-based system for classifying autism spectrum disorder (ASD) using facial attributes such as expressions, action units, arousal, and valence. Our system classifies ASD using…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Beibin Li , Sachin Mehta , Deepali Aneja , Claire Foster , Pamela Ventola , Frederick Shic , Linda Shapiro

Facial pain expression is an important modality for assessing pain, especially when the patient's verbal ability to communicate is impaired. The facial muscle-based action units (AUs), which are defined by the Facial Action Coding System…

Computer Vision and Pattern Recognition · Computer Science 2018-11-21 Zhanli Chen , Rashid Ansari , Diana Wilkie

Human motion is fundamental to understanding behavior. Despite progress on single-image 3D pose and shape estimation, existing video-based state-of-the-art methods fail to produce accurate and natural motion sequences due to a lack of…

Computer Vision and Pattern Recognition · Computer Science 2020-05-01 Muhammed Kocabas , Nikos Athanasiou , Michael J. Black

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, and where someone is…

Computer Vision and Pattern Recognition · Computer Science 2022-12-26 Rahul Sharma , Krishna Somandepalli , Shrikanth Narayanan

Previous work has established that a person's demographics and speech style affect how well speech processing models perform for them. But where does this bias come from? In this work, we present the Speech Embedding Association Test…

Computation and Language · Computer Science 2023-10-31 Isaac Slaughter , Craig Greenberg , Reva Schwartz , Aylin Caliskan

Automatic speech recognition (ASR) for conversational speech remains challenging due to the limited availability of large-scale, well-annotated multi-speaker dialogue data and the complex temporal dynamics of natural interactions.…

Sound · Computer Science 2026-02-05 Máté Gedeon , Péter Mihajlik
‹ Prev 1 8 9 10 Next ›