中文
相关论文

相关论文: MiniMind-O Technical Report: An Open Small-Scale S…

200 篇论文

Despite rapid progress in text-to-speech (TTS), open-source systems still lack truly instruction-following, fine-grained control over core speech attributes (e.g., pitch, speaking rate, age, emotion, and style). We present VoiceSculptor, an…

Audio-aware large language models (ALLMs) can understand the textual and non-textual information in the audio input. In this paper, we explore using ALLMs as an automatic judge to assess the speaking styles of speeches. We use ALLM judges…

音频与语音处理 · 电气工程与系统科学 2025-06-09 Cheng-Han Chiang , Xiaofei Wang , Chung-Ching Lin , Kevin Lin , Linjie Li , Radu Kopetz , Yao Qian , Zhendong Wang , Zhengyuan Yang , Hung-yi Lee , Lijuan Wang

Full-duplex multimodal large language models (LLMs) provide a unified framework for addressing diverse speech understanding and generation tasks, enabling more natural and seamless human-machine conversations. Unlike traditional modularised…

音频与语音处理 · 电气工程与系统科学 2024-11-28 Wenyi Yu , Siyin Wang , Xiaoyu Yang , Xianzhao Chen , Xiaohai Tian , Jun Zhang , Guangzhi Sun , Lu Lu , Yuxuan Wang , Chao Zhang

Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision. Whether similar principles apply to modeling brain activity remains unclear. Here we leveraged a dataset of 3.1 million neurons…

We present MiMo-7B, a large language model born for reasoning tasks, with optimization across both pre-training and post-training stages. During pre-training, we enhance the data preprocessing pipeline and employ a three-stage data mixing…

We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlike autoregressive…

计算与语言 · 计算机科学 2026-04-02 Jaeik Kim , Woojin Kim , Jihwan Hong , Yejoon Lee , Sieun Hyeon , Mintaek Lim , Yunseok Han , Dogeun Kim , Hoeun Lee , Hyunggeun Kim , Jaeyoung Do

Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question…

音频与语音处理 · 电气工程与系统科学 2025-10-28 Wenming Tu , Guanrou Yang , Ruiqi Yan , Wenxi Chen , Ziyang Ma , Yipeng Kang , Kai Yu , Xie Chen , Zilong Zheng

We introduce AudioPaLM, a large language model for speech understanding and generation. AudioPaLM fuses text-based and speech-based language models, PaLM-2 [Anil et al., 2023] and AudioLM [Borsos et al., 2022], into a unified multimodal…

Machine writing with large language models often relies on retrieval-augmented generation. However, these approaches remain confined within the boundaries of the model's predefined scope, limiting the generation of content with rich…

计算与语言 · 计算机科学 2025-11-21 Zekun Xi , Wenbiao Yin , Jizhan Fang , Jialong Wu , Runnan Fang , Yong Jiang , Pengjun Xie , Fei Huang , Huajun Chen , Ningyu Zhang

Recognizing speaking in humans is a central task towards understanding social interactions. Ideally, speaking would be detected from individual voice recordings, as done previously for meeting scenarios. However, individual voice recordings…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Jose Vargas Quiros , Chirag Raman , Stephanie Tan , Ekin Gedik , Laura Cabrera-Quiros , Hayley Hung

Modeling the interplay between external stimuli and internal neural representations is a pivotal research area for Brain-Computer Interfaces (BCIs). A major limitation of prior work is the prevailing paradigm of specialized, single-task…

人工智能 · 计算机科学 2026-05-29 Yizhuo Lu , Changde Du , Qingyu Shi , Hang Chen , Jie Peng , Liuyun Jiang , Shuangchen Zhao , Huiguang He

Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongside spoken words, and respond with empathy that aligns with…

计算与语言 · 计算机科学 2026-03-06 Li Zhou , Lutong Yu , You Lyu , Yihang Lin , Zefeng Zhao , Junyi Ao , Yuhao Zhang , Benyou Wang , Haizhou Li

Instruction-following text-to-speech (TTS) has emerged as an important capability for controllable and expressive speech generation, yet its evaluation remains underdeveloped due to limited benchmark coverage, weak diagnostic granularity,…

音频与语音处理 · 电气工程与系统科学 2026-04-21 Huakang Chen , Jingbin Hu , Liumeng Xue , Qirui Zhan , Wenhao Li , Guobin Ma , Hanke Xie , Dake Guo , Linhan Ma , Yuepeng Jiang , Bengu Wu , Pengyuan Xie , Chuan Xie , Qiang Zhang , Lei Xie

Large Language Models (LLMs) have revolutionized natural language processing, but their application to speech-based tasks remains challenging due to the complexities of integrating audio and text modalities. This paper introduces Ichigo, a…

计算与语言 · 计算机科学 2025-04-07 Alan Dao , Dinh Bach Vu , Huy Hoang Ha

The training of high-quality, robust machine learning models for speech-driven 3D facial animation requires a large, diverse dataset of high-quality audio-animation pairs. To overcome the lack of such a dataset, recent work has introduced…

图形学 · 计算机科学 2026-02-13 Zhen Han , Mattias Teye , Derek Yadgaroff , Judith Bütepage

Large language models have recently evolved from fluent text generation to advanced reasoning across diverse domains, giving rise to reasoning language models. Among these domains, mathematical reasoning serves as a representative benchmark…

Recent advances in Large Audio-Language Models (LALMs) have made real-time, streaming spoken interaction increasingly practical. In this setting, reasoning quality and responsiveness are tightly coupled: delaying reasoning until the speech…

计算与语言 · 计算机科学 2026-05-27 Zhiyuan Song , Weici Zhao , Yang Xiao , Suhao Yu , Cheng Zhu , Jiatao Gu

Improvements in training data scale and quality have led to significant advances, yet its influence in speech recognition remains underexplored. In this paper, we present a large-scale dataset, OLMoASR-Pool, and series of models, OLMoASR,…

声音 · 计算机科学 2025-08-29 Huong Ngo , Matt Deitke , Martijn Bartelds , Sarah Pratt , Josh Gardner , Matt Jordan , Ludwig Schmidt

Theory of Mind (ToM)$\unicode{x2014}$the ability to reason about the mental states of other people$\unicode{x2014}$is a key element of our social intelligence. Yet, despite their ever more impressive performance, large-scale neural language…

计算与语言 · 计算机科学 2023-06-02 Melanie Sclar , Sachin Kumar , Peter West , Alane Suhr , Yejin Choi , Yulia Tsvetkov

As the paradigm of AI shifts from text-based LLMs to Speech Language Models (SLMs), there is a growing demand for full-duplex systems capable of real-time, natural human-computer interaction. However, the development of such models is…

声音 · 计算机科学 2026-03-31 Kyudan Jung , Jihwan Kim , Soyoon Kim , Jeonghoon Kim , Jaegul Choo , Cheonbok Park