中文
相关论文

相关论文: SonicSim: A customizable simulation platform for s…

200 篇论文

The scarcity of large-scale classroom speech data has hindered the development of AI-driven speech models for education. Classroom datasets remain limited and not publicly available, and the absence of dedicated classroom noise or Room…

声音 · 计算机科学 2025-10-03 Ahmed Adel Attia , Jing Liu , Carol Espy Wilson

As large-scale speech-to-speech models achieve high fidelity, the distinction between synthetic voices in structured environments becomes a vital area of study. This paper introduces Advosynth-500, a specialized dataset comprising 100…

计算与语言 · 计算机科学 2026-01-16 Aniket Deroy

Detecting synthetic from real speech is increasingly crucial due to the risks of misinformation and identity impersonation. While various datasets for synthetic speech analysis have been developed, they often focus on specific areas,…

声音 · 计算机科学 2025-07-18 Zhoulin Ji , Chenhao Lin , Hang Wang , Chao Shen

Clinical information extraction, which involves structuring clinical concepts from unstructured medical text, remains a challenging problem that could benefit from the inclusion of tabular background information available in electronic…

人工智能 · 计算机科学 2025-12-10 Paloma Rabaey , Stefan Heytens , Thomas Demeester

Psychiatric symptom identification on social media aims to infer fine-grained mental health symptoms from user-generated posts, allowing a detailed understanding of users' mental states. However, the construction of large-scale…

计算与语言 · 计算机科学 2026-03-24 Migyeong Kang , Jihyun Kim , Hyolim Jeon , Sunwoo Hwang , Jihyun An , Yonghoon Kim , Haewoon Kwak , Jisun An , Jinyoung Han

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy…

音频与语音处理 · 电气工程与系统科学 2024-12-30 Jaemin Jung , Junseok Ahn , Chaeyoung Jung , Tan Dat Nguyen , Youngjoon Jang , Joon Son Chung

This paper presents a residential audio dataset to support sound event detection research for smart home applications aimed at promoting wellbeing for older adults. The dataset is constructed by deploying audio recording systems in the…

声音 · 计算机科学 2024-10-07 Gabriel Bibbó , Thomas Deacon , Arshdeep Singh , Mark D. Plumbley

A large and growing amount of speech content in real-life scenarios is being recorded on consumer-grade devices in uncontrolled environments, resulting in degraded speech quality. Transforming such low-quality device-degraded speech into…

音频与语音处理 · 电气工程与系统科学 2022-03-23 Haoyu Li , Junichi Yamagishi

There is growing interest in exploring user simulation as an alternative to gathering and scoring real user-chatbot interactions for AI chatbot evaluation. For this purpose, it is important to ensure the realism of the simulation, i.e., the…

计算与语言 · 计算机科学 2026-05-05 Yu Lu Liu , Hyokun Yun , Tanya Roosta , Ziang Xiao

Many datasets have been designed to further the development of fake audio detection. However, fake utterances in previous datasets are mostly generated by altering timbre, prosody, linguistic content or channel noise of original audio.…

We introduce the Free Universal Sound Separation (FUSS) dataset, a new corpus for experiments in separating mixtures of an unknown number of sounds from an open domain of sound types. The dataset consists of 23 hours of single-source audio…

Scaling data volume and diversity is critical for generalizing embodied intelligence. While synthetic data generation offers a scalable alternative to expensive physical data acquisition, transferring robotic manipulation policies from…

Recent progress in separating the speech signals from multiple overlapping speakers using a single audio channel has brought us closer to solving the cocktail party problem. However, most studies in this area use a constrained problem…

Recent advances in foundation models have enabled audio-generative models that produce high-fidelity sounds associated with music, events, and human actions. Despite the success achieved in modern audio-generative models, the conventional…

声音 · 计算机科学 2024-08-30 Tiantian Feng , Dimitrios Dimitriadis , Shrikanth Narayanan

General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separation models are either domain-specific, designed for fixed…

As perception models continue to develop, the need for large-scale datasets increases. However, data annotation remains far too expensive to effectively scale and meet the demand. Synthetic datasets provide a solution to boost model…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Arpit Jadon , Haoran Wang , Phillip Thomas , Michael Stanley , S. Nathaniel Cibik , Rachel Laurat , Omar Maher , Lukas Hoyer , Ozan Unal , Dengxin Dai

Simulation is increasingly being used for generating large labelled datasets in many machine learning problems. Recent methods have focused on adjusting simulator parameters with the goal of maximising accuracy on a validation task, usually…

计算机视觉与模式识别 · 计算机科学 2020-08-20 Harkirat Singh Behl , Atılım Güneş Baydin , Ran Gal , Philip H. S. Torr , Vibhav Vineet

Training models to high-end performance requires availability of large labeled datasets, which are expensive to get. The goal of our work is to automatically synthesize labeled datasets that are relevant for a downstream task. We propose…

计算机视觉与模式识别 · 计算机科学 2019-04-29 Amlan Kar , Aayush Prakash , Ming-Yu Liu , Eric Cameracci , Justin Yuan , Matt Rusiniak , David Acuna , Antonio Torralba , Sanja Fidler

Passive Acoustic Monitoring (PAM) analysis is often hindered by the intensive manual effort needed to create labelled training data. This study introduces a synthetic data framework to generate large volumes of richly labelled training data…

声音 · 计算机科学 2025-07-23 Kaspar Soltero , Tadeu Siqueira , Stefanie Gutschmidt

We propose a new dataset for cinematic audio source separation (CASS) that handles non-verbal sounds. Existing CASS datasets only contain reading-style sounds as a speech stem. These datasets differ from actual movie audio, which is more…

声音 · 计算机科学 2025-06-10 Takuya Hasumi , Yusuke Fujita