English
Related papers

Related papers: The VVAD-LRS3 Dataset for Visual Voice Activity De…

200 papers

Longitudinal brain MRI is essential for characterizing the progression of neurological diseases such as Alzheimer's disease assessment. However, current deep-learning tools fragment this process: classifiers reduce a scan to a label,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Zhaoyang Jiang , Zhizhong Fu , David McAllister , Yunsoo Kim , Honghan Wu

Facial expression recognition from videos in the wild is a challenging task due to the lack of abundant labelled training data. Large DNN (deep neural network) architectures and ensemble methods have resulted in better performance, but soon…

Computer Vision and Pattern Recognition · Computer Science 2021-02-26 Vikas Kumar , Shivansh Rao , Li Yu

Audio-visual speech recognition (AVSR) gains increasing attention from researchers as an important part of human-computer interaction. However, the existing available Mandarin audio-visual datasets are limited and lack the depth…

Sound · Computer Science 2023-06-06 Jianrong Wang , Yuchen Huo , Li Liu , Tianyi Xu , Qi Li , Sen Li

A challenge in speech production research is to predict future tongue movements based on a short period of past tongue movements. This study tackles speaker-dependent tongue motion prediction problem in unlabeled ultrasound videos with…

Computer Vision and Pattern Recognition · Computer Science 2019-02-20 Chaojie Zhao , Peng Zhang , Jian Zhu , Chengrui Wu , Huaimin Wang , Kele Xu

Vision-and-language models (VLMs) have been increasingly explored in the medical domain, particularly following the success of CLIP in general domain. However, unlike the relatively straightforward pairing of 2D images and text, curating…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Ziyang Zhang , Yang Yu , Xulei Yang , Si Yong Yeo

To act in the world, a model must name what it sees and know where it is in 3D. Today's vision-language models (VLMs) excel at open-ended 2D description and grounding, yet multi-object 3D detection remains largely missing from the VLM…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Yunze Man , Shihao Wang , Guowen Zhang , Johan Bjorck , Zhiqi Li , Liang-Yan Gui , Jim Fan , Jan Kautz , Yu-Xiong Wang , Zhiding Yu

Language Identification (LID) systems are used to classify the spoken language from a given audio sample and are typically the first step for many spoken language processing tasks, such as Automatic Speech Recognition (ASR) systems. Without…

Computer Vision and Pattern Recognition · Computer Science 2017-08-17 Christian Bartz , Tom Herold , Haojin Yang , Christoph Meinel

Current Active Speaker Detection (ASD) models achieve great results on AVA-ActiveSpeaker (AVA), using only sound and facial features. Although this approach is applicable in movie setups (AVA), it is not suited for less constrained…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Tiago Roxo , Joana C. Costa , Pedro R. M. Inácio , Hugo Proença

Nonverbal communication is integral to human interaction, with gestures, facial expressions, and body language conveying critical aspects of intent and emotion. However, existing large language models (LLMs) fail to effectively incorporate…

Artificial Intelligence · Computer Science 2025-06-03 Youngmin Kim , Jiwan Chung , Jisoo Kim , Sunghyun Lee , Sangkyu Lee , Junhyeok Kim , Cheoljong Yang , Youngjae Yu

The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content authenticity. Existing synthetic video datasets predominantly focus on the visual modality…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Mengxue Hu , Yunfeng Diao , Changtao Miao , Zhiqing Guo , Jianshu Li , Zhe Li , Joey Tianyi Zhou

With the goal of more natural and human-like interaction with virtual voice assistants, recent research in the field has focused on full duplex interaction mode without relying on repeated wake-up words. This requires that in scenes with…

Sound · Computer Science 2024-09-17 Anna Wang , Da Liu , Zhiyu Zhang , Shengqiang Liu , Jie Gao , Yali Li

Open-vocabulary 3D object detection methods are able to localize 3D boxes of classes unseen during training. Despite the name, existing methods rely on user-specified classes both at training and inference. We propose to study…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Haomeng Zhang , Kuan-Chuan Peng , Suhas Lohit , Raymond A. Yeh

The intelligent dialogue system, aiming at communicating with humans harmoniously with natural language, is brilliant for promoting the advancement of human-machine interaction in the era of artificial intelligence. With the gradually…

Artificial Intelligence · Computer Science 2022-07-05 Hao Wang , Bin Guo , Yating Zeng , Yasan Ding , Chen Qiu , Ying Zhang , Lina Yao , Zhiwen Yu

Voice activity detection (VAD) is essential for speech-driven applications, but remains far from perfect in noisy and resource-limited environments. Existing methods often lack robustness to noise, and their frame-wise classification losses…

Sound · Computer Science 2025-08-29 Chien-Chun Wang , En-Lun Yu , Jeih-Weih Hung , Shih-Chieh Huang , Berlin Chen

Object detection (OD) in computer vision has made significant progress in recent years, transitioning from closed-set labels to open-vocabulary detection (OVD) based on large-scale vision-language pre-training (VLP). However, current…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Yiyang Yao , Peng Liu , Tiancheng Zhao , Qianqian Zhang , Jiajia Liao , Chunxin Fang , Kyusong Lee , Qing Wang

Vision-language model (VLM) fine-tuning for application-specific visual grounding based on natural language instructions has become one of the most popular approaches for learning-enabled autonomous systems. However, such fine-tuning relies…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Joshua R. Waite , Md. Zahid Hasan , Qisai Liu , Zhanhong Jiang , Chinmay Hegde , Soumik Sarkar

Smartphones have been employed with biometric-based verification systems to provide security in highly sensitive applications. Audio-visual biometrics are getting popular due to their usability, and also it will be challenging to spoof…

Visual speech recognition (VSR), which decodes spoken words from video data, offers significant benefits, particularly when audio is unavailable. However, the high dimensionality of video data leads to prohibitive computational costs that…

Computer Vision and Pattern Recognition · Computer Science 2025-02-10 Iason Ioannis Panagos , Giorgos Sfikas , Christophoros Nikou

Translating natural language to visualization (NL2VIS) has shown great promise for visual data analysis, but it remains a challenging task that requires multiple low-level implementations, such as natural language processing and…

Human-Computer Interaction · Computer Science 2024-08-08 Nan Chen , Yuge Zhang , Jiahang Xu , Kan Ren , Yuqing Yang

With the continuously thriving popularity around the world, fitness activity analytic has become an emerging research topic in computer vision. While a variety of new tasks and algorithms have been proposed recently, there are growing…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Yansong Tang , Jinpeng Liu , Aoyang Liu , Bin Yang , Wenxun Dai , Yongming Rao , Jiwen Lu , Jie Zhou , Xiu Li