English
Related papers

Related papers: Mandarin Lombard Flavor Classification

200 papers

Spoken word recognition involves at least two basic computations. First is matching acoustic input to phonological categories (e.g. /b/, /p/, /d/). Second is activating words consistent with those phonological categories. Here we test the…

Computation and Language · Computer Science 2017-11-21 Laura Gwilliams , David Poeppel , Alec Marantz , Tal Linzen

Language models for agglutinative languages have always been hindered in past due to myriad of agglutinations possible to any given word through various affixes. We propose a method to diminish the problem of out-of-vocabulary words by…

Computation and Language · Computer Science 2017-08-21 Seunghak Yu , Nilesh Kulkarni , Haejun Lee , Jihie Kim

Reinforcement learning with verifiable rewards (RLVR) can elicit strong reasoning in large language models (LLMs), while their performance after RLVR varies dramatically across different base models. This raises a fundamental question: what…

Machine Learning · Computer Science 2025-10-22 Xuansheng Wu , Xiaoman Pan , Wenlin Yao , Jianshu Chen

This paper presents an embedding-based approach to detecting variation without relying on prior normalisation or predefined variant lists. The method trains subword embeddings on raw text and groups related forms through combined cosine and…

Computation and Language · Computer Science 2026-02-13 Anne-Marie Lutgen , Alistair Plum , Christoph Purschke

Given the great success of Deep Neural Networks(DNNs) and the black-box nature of it,the interpretability of these models becomes an important issue.The majority of previous research works on the post-hoc interpretation of a trained…

Machine Learning · Computer Science 2021-03-22 Haoyang Li , Xinggang Wang

Large language models (LLMs) increasingly exhibit human-like patterns of pragmatic and social reasoning. This paper addresses two related questions: do LLMs approximate human social meaning not only qualitatively but also quantitatively,…

Computation and Language · Computer Science 2026-04-06 Roland Mühlenbernd

Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the extent to which generated speech is acoustically diverse…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Matthieu Futeral , Andrea Agostinelli , Marco Tagliasacchi , Neil Zeghidour , Eugene Kharitonov

The wording of natural language prompts has been shown to influence the performance of large language models (LLMs), yet the role of politeness and tone remains underexplored. In this study, we investigate how varying levels of prompt…

Computation and Language · Computer Science 2025-10-07 Om Dobariya , Akhil Kumar

Major Depressive Disorder (MDD) is a severe illness that affects millions of people, and it is critical to diagnose this disorder as early as possible. Detecting depression from voice signals can be of great help to physicians and can be…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Jinhan Wang , Vijay Ravi , Jonathan Flint , Abeer Alwan

The goal of most subjective studies is to place a set of stimuli on a perceptual scale. This is mostly done directly by rating, e.g. using single or double stimulus methodologies, or indirectly by ranking or pairwise comparison. All these…

Neurons and Cognition · Quantitative Biology 2022-03-25 Pastor Andréas , Lukáš Krasula , Xiaoqing Zhu , Zhi Li , Patrick Le Callet

The cleft lip and palate (CLP) speech intelligibility is distorted due to the deformation in their articulatory system. For addressing the same, a few previous works perform phoneme specific modification in CLP speech. In CLP speech, both…

Sound · Computer Science 2021-10-05 Protima Nomo Sudro , Rohit Sinha , S. R. Mahadeva Prasanna

An ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs…

This research delves into Musculoskeletal Disorder (MSD) risk factors, using a blend of Natural Language Processing (NLP) and mode-based ranking. The aim is to refine understanding, classification, and prioritization for focused prevention…

Computation and Language · Computer Science 2024-11-06 Md Abrar Jahin , Subrata Talapatra

A multi-modal emotional speech Mandarin database including articulatory kinematics, acoustics, glottal and facial micro-expressions is designed and established, which is described in detail from the aspects of corpus design, subject…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-17 Zhu Ting , Li Liangqi , Duan Shufei , Zhang Xueying , Xiao Zhongzhe , Jia Hairng , Liang Huizhi

NLP models often degrade in performance when real world data distributions differ markedly from training data. However, existing dataset drift metrics in NLP have generally not considered specific dimensions of linguistic drift that affect…

Computation and Language · Computer Science 2023-05-29 Tyler A. Chang , Kishaloy Halder , Neha Anna John , Yogarshi Vyas , Yassine Benajiba , Miguel Ballesteros , Dan Roth

Speech language models (SLMs) have significantly extended the interactive capability of text-based Large Language Models (LLMs) by incorporating paralinguistic information. For more realistic interactive experience with customized styles,…

Computation and Language · Computer Science 2026-03-10 Haishu Zhao , Aokai Hao , Yuan Ge , Zhenqiang Hong , Tong Xiao , Jingbo Zhu

Digital audio effects are widely used by audio engineers to alter the acoustic and temporal qualities of audio data. However, these effects can have a large number of parameters which can make them difficult to learn for beginners and…

Machine Learning · Computer Science 2023-10-02 Kieran Grant

One precondition of effective oral communication is that words should be pronounced clearly, especially for non-native speakers. Word stress is the key to clear and correct English, and misplacement of syllable stress may lead to…

Sound · Computer Science 2023-11-02 Wang Weiying , Nakajima Akinori

The introduction and regulation of loudness in broadcasting and streaming brought clear benefits to the audience, e.g., a level of uniformity across programs and channels. Yet, speech loudness is frequently reported as being too low in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-28 Matteo Torcoli , Mhd Modar Halimeh , Thomas Leitz , Yannik Grewe , Michael Kratschmer , Bernhard Neugebauer , Adrian Murtaza , Harald Fuchs , Emanuël A. P. Habets

Previous accent classification research focused mainly on detecting accents with pure acoustic information without recognizing accented speech. This work combines phonetic knowledge such as vowels with acoustic information to build Guassian…

Sound · Computer Science 2016-04-28 Zhenhao Ge , Yingyi Tan , Aravind Ganapathiraju