中文
相关论文

相关论文: Open-Source Conversational AI with SpeechBrain 1.0

200 篇论文

We introduce OpenVoice, a versatile voice cloning approach that requires only a short audio clip from the reference speaker to replicate their voice and generate speech in multiple languages. OpenVoice represents a significant advancement…

声音 · 计算机科学 2024-08-20 Zengyi Qin , Wenliang Zhao , Xumin Yu , Xin Sun

Non-verbal behavior is essential for embodied agents like social robots, virtual avatars, and digital humans. Existing behavior authoring approaches including keyframe animation and motion capture are too expensive to use when there are…

人机交互 · 计算机科学 2021-08-11 Youngwoo Yoon , Keunwoo Park , Minsu Jang , Jaehong Kim , Geehyuk Lee

Large scale Speech Language Models have enabled voice assistants capable of understanding natural spoken queries and performing complex tasks. However, existing speech benchmarks largely focus on isolated capabilities such as transcription…

人工智能 · 计算机科学 2026-02-16 Dhruv Jain , Harshit Shukla , Gautam Rajeev , Ashish Kulkarni , Chandra Khatri , Shubham Agarwal

One challenge in technical interviews is the think-aloud process, where candidates verbalize their thought processes while solving coding tasks. Despite its importance, opportunities for structured practice remain limited. Conversational AI…

Recent advancements in multi-turn voice interaction models have improved user-model communication. However, while closed-source models effectively retain and recall past utterances, whether open-source models share this ability remains…

声音 · 计算机科学 2025-05-26 Heeseung Kim , Che Hyun Lee , Sangkwon Park , Jiheum Yeom , Nohil Park , Sangwon Yu , Sungroh Yoon

The availability of open-source software is playing a remarkable role in the popularization of speech recognition and deep learning. Kaldi, for instance, is nowadays an established framework used to develop state-of-the-art speech…

音频与语音处理 · 电气工程与系统科学 2019-02-19 Mirco Ravanelli , Titouan Parcollet , Yoshua Bengio

Silent speech interface is a promising technology that enables private communications in natural language. However, previous approaches only support a small and inflexible vocabulary, which leads to limited expressiveness. We leverage…

人机交互 · 计算机科学 2023-03-07 Zixiong Su , Shitao Fang , Jun Rekimoto

We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct mappings from audio to raw pixels or points for creating…

计算机视觉与模式识别 · 计算机科学 2021-02-26 Yang Zhou , Xintong Han , Eli Shechtman , Jose Echevarria , Evangelos Kalogerakis , Dingzeyu Li

Sanskrit, one of humanity's most ancient languages, has a vast collection of books and manuscripts on diverse topics that have been accumulated over millennia. However, its digital content (audio and text), which is vital for the training…

计算与语言 · 计算机科学 2025-01-20 Bidit Sadhukhan , Swami Punyeshwarananda

As communications are increasingly taking place virtually, the ability to present well online is becoming an indispensable skill. Online speakers are facing unique challenges in engaging with remote audiences. However, there has been a lack…

人机交互 · 计算机科学 2023-09-12 Zeyuan Huang , Qiang He , Kevin Maher , Xiaoming Deng , Yu-Kun Lai , Cuixia Ma , Sheng-feng Qin , Yong-Jin Liu , Hongan Wang

We propose a novel method for generating high-resolution videos of talking-heads from speech audio and a single 'identity' image. Our method is based on a convolutional neural network model that incorporates a pre-trained StyleGAN…

计算机视觉与模式识别 · 计算机科学 2022-09-12 Mohammed M. Alghamdi , He Wang , Andrew J. Bulpitt , David C. Hogg

Newcomers onboarding to Open Source Software (OSS) projects face many challenges. Large Language Models (LLMs), like ChatGPT, have emerged as potential resources for answering questions and providing guidance, with many developers now…

软件工程 · 计算机科学 2025-02-12 Italo Santos , Katia Romero Felizardo , Igor Steinmacher , Marco A. Gerosa

Social interactions and conversation skills separate the successful from the rest and the confident from the shy. For college students in particular, the ability to converse can be an outlet for the stress and anxiety experienced on a daily…

人机交互 · 计算机科学 2023-10-24 Hyunbae Jeon , Rhea Ramachandran , Victoria Ploerer , Yella Diekmann , Max Bagga

Open data refers to data that is freely available for reuse. Although there has been rapid increase in availability of open data to public in the last decade, this has not translated into better decision-support tools for them. We propose…

人工智能 · 计算机科学 2019-01-14 Biplav Srivastava

Developing specialized dialogue systems for mental health support requires multi-turn conversation data, which has recently garnered increasing attention. However, gathering and releasing large-scale, real-life multi-turn conversations that…

计算与语言 · 计算机科学 2025-08-05 Huachuan Qiu , Hongliang He , Shuai Zhang , Anqi Li , Zhenzhong Lan

Sketching is a widely used medium for generating and exploring early-stage design concepts. While generative AI (GenAI) chatbots are increasingly used for idea generation, designers often struggle to craft effective prompts and find it…

人机交互 · 计算机科学 2025-11-19 Weiyan Shi , Sunaya Upadhyay , Geraldine Quek , Kenny Tsu Wei Choo

We introduce Fish Audio S2, an open-sourced text-to-speech system featuring multi-speaker, multi-turn generation, and, most importantly, instruction-following control via natural-language descriptions. To scale training, we develop a…

The AI Steerability 360 toolkit is an extensible, open-source Python library for steering LLMs. Steering abstractions are designed around four model control surfaces: input (modification of the prompt), structural (modification of the…

We present ADVISER - an open-source, multi-domain dialog system toolkit that enables the development of multi-modal (incorporating speech, text and vision), socially-engaged (e.g. emotion recognition, engagement level prediction and…

Automatic Speech Recognition and Text-to-Speech systems are primarily trained in a supervised fashion and require high-quality, accurately labeled speech datasets. In this work, we examine common problems with speech data and introduce a…

音频与语音处理 · 电气工程与系统科学 2022-01-10 Evelina Bakhturina , Vitaly Lavrukhin , Boris Ginsburg