中文
相关论文

相关论文: A Multi-Purpose Audio-Visual Corpus for Multi-Moda…

200 篇论文

Grapheme-to-phoneme (G2P) conversion for Persian presents unique challenges due to its complex phonological features, particularly homographs and Ezafe, which exist in formal and informal language contexts. This paper introduces an…

计算与语言 · 计算机科学 2025-05-13 Abbas Bertina , Shahab Beirami , Hossein Biniazian , Elham Esmaeilnia , Soheil Shahi , Mahdi Pirnia

Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from audio input. Current TFS models perform well in English but unsatisfactorily in non-English languages, producing wrong mouth shapes and rigid…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Zibo Su , Kun Wei , Jiahua Li , Xu Yang , Cheng Deng

In this paper, we introduce a large-scale and high-quality audio-visual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains…

音频与语音处理 · 电气工程与系统科学 2023-12-14 Yuke Lin , Xiaoyi Qin , Guoqing Zhao , Ming Cheng , Ning Jiang , Haiyang Wu , Ming Li

The performance of Artificial Intelligence (AI) systems fundamentally depends on high-quality training data. However, low-resource languages like Arabic suffer from severe data scarcity. Moreover, the absence of child-specific speech…

计算与语言 · 计算机科学 2025-10-28 Mouhand Alkadri , Dania Desouki , Khloud Al Jallad

Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. Augmenting these systems with visual signals has the potential to improve robustness to noise. However, audio-visual…

声音 · 计算机科学 2024-08-13 HyoJung Han , Mohamed Anwar , Juan Pino , Wei-Ning Hsu , Marine Carpuat , Bowen Shi , Changhan Wang

Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal restraint and enabling discrete interaction in noise-sensitive…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Cancan Li , Fei Su , Juan Liu , Hui Bu , Yulong Wan , Hongbin Suo , Ming Li

Visual voice activity detection (V-VAD) uses visual features to predict whether a person is speaking or not. V-VAD is useful whenever audio VAD (A-VAD) is inefficient either because the acoustic signal is difficult to analyze or because it…

计算机视觉与模式识别 · 计算机科学 2020-10-19 Sylvain Guy , Stéphane Lathuilière , Pablo Mesejo , Radu Horaud

Speechreading or lipreading is the technique of understanding and getting phonetic features from a speaker's visual features such as movement of lips, face, teeth and tongue. It has a wide range of multimedia applications such as in…

As large language models (LLMs) become increasingly embedded in our daily lives, evaluating their quality and reliability across diverse contexts has become essential. While comprehensive benchmarks exist for assessing LLM performance in…

Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent…

声音 · 计算机科学 2020-11-05 Arsha Nagrani , Joon Son Chung , Andrew Zisserman

Multi-Modal automatic speech recognition (ASR) techniques aim to leverage additional modalities to improve the performance of speech recognition systems. While existing approaches primarily focus on video or contextual information, the…

声音 · 计算机科学 2023-12-27 Haoxu Wang , Fan Yu , Xian Shi , Yuezhang Wang , Shiliang Zhang , Ming Li

This paper presents the development of Rezwan, a large-scale AI-assisted Hadith corpus comprising over 1.2M narrations, extracted and structured through a fully automated pipeline. Building on digital repositories such as Maktabat Ahl…

In this work, we present a novel audio-visual dataset for active speaker detection in the wild. A speaker is considered active when his or her face is visible and the voice is audible simultaneously. Although active speaker detection is a…

计算机视觉与模式识别 · 计算机科学 2021-08-18 You Jin Kim , Hee-Soo Heo , Soyeon Choe , Soo-Whan Chung , Yoohwan Kwon , Bong-Jin Lee , Youngki Kwon , Joon Son Chung

Sentiment analysis aims to extract people's emotions and opinion from their comments on the web. It widely used in businesses to detect sentiment in social data, gauge brand reputation, and understand customers. Most of articles in this…

计算与语言 · 计算机科学 2022-12-13 Ali Nazarizadeh , Touraj Banirostam , Minoo Sayyadpour

During a conversation, our brain is responsible for combining information obtained from multiple senses in order to improve our ability to understand the message we are perceiving. Different studies have shown the importance of presenting…

计算机视觉与模式识别 · 计算机科学 2023-11-22 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

计算机视觉与模式识别 · 计算机科学 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

Grapheme-to-phoneme (G2P) conversion is critical in speech processing, particularly for applications like speech synthesis. G2P systems must possess linguistic understanding and contextual awareness of languages with polyphone words and…

计算与语言 · 计算机科学 2024-09-16 Mahta Fetrat Qharabagh , Zahra Dehghanian , Hamid R. Rabiee

Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization. Both audio- and vision-based approaches have been used for this task in…

Generating semantically coherent and visually accurate talking faces requires bridging the gap between linguistic meaning and facial articulation. Although audio-driven methods remain prevalent, their reliance on high-quality paired audio…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Xu Wang , Shengeng Tang , Fei Wang , Lechao Cheng , Dan Guo , Feng Xue , Richang Hong

We introduce "ivrit.ai", a comprehensive Hebrew speech dataset, addressing the distinct lack of extensive, high-quality resources for advancing Automated Speech Recognition (ASR) technology in Hebrew. With over 3,300 speech hours and a over…

音频与语音处理 · 电气工程与系统科学 2023-07-19 Yanir Marmor , Kinneret Misgav , Yair Lifshitz