中文
相关论文

相关论文: Affect Models Have Weak Generalizability to Atypic…

200 篇论文

Despite the success of language models using neural networks, it remains unclear to what extent neural models have the generalization ability to perform inferences. In this paper, we introduce a method for evaluating whether neural models…

计算与语言 · 计算机科学 2020-05-05 Hitomi Yanaka , Koji Mineshima , Daisuke Bekki , Kentaro Inui

Disturbances in temporality, such as desynchronization with the social environment and its unpredictability, are considered core features of autism with a deep impact on relationships. However, limitations regarding research on this issue…

计算与语言 · 计算机科学 2026-02-16 Kacper Dudzic , Karolina Drożdż , Maciej Wodziński , Anastazja Szuła , Marcin Moskalewicz

Speech rate has been shown to vary across social categories such as gender, age, and dialect, while also being conditioned by properties of speech planning. The effect of utterance length, where speech rate is faster and less variable for…

计算与语言 · 计算机科学 2024-08-14 James Tanner , Morgan Sonderegger , Jane Stuart-Smith , Tyler Kendall , Jeff Mielke , Robin Dodsworth , Erik Thomas

Machine learning models for speech emotion recognition (SER) can be trained for different tasks and are usually evaluated based on a few available datasets per task. Tasks could include arousal, valence, dominance, emotional categories, or…

音频与语音处理 · 电气工程与系统科学 2025-02-13 Anna Derington , Hagen Wierstorf , Ali Özkil , Florian Eyben , Felix Burkhardt , Björn W. Schuller

Voice privacy approaches that preserve the anonymity of speakers modify speech in an attempt to break the link with the true identity of the speaker. Current benchmarks measure speaker protection based on signal-to-signal comparisons. In…

声音 · 计算机科学 2026-03-25 Mehtab Ur Rahman , Martha Larson , Cristian Tejedor-Garcia

We consider the problem of using observational data to estimate the causal effects of linguistic properties. For example, does writing a complaint politely lead to a faster response time? How much will a positive product review increase…

计算与语言 · 计算机科学 2021-06-15 Reid Pryzant , Dallas Card , Dan Jurafsky , Victor Veitch , Dhanya Sridhar

The double empathy problem recasts the difficulty of forming empathy bonds in social interactions between autistic and neurotypical individuals as a bidirectional problem, rather than due to a deficit exclusive to the person on the…

物理与社会 · 物理学 2026-02-04 Enrique Calderoli , Maria Cristina Varriale , Flávio Kapczinski

When training data are collected from human annotators, the design of the annotation instrument, the instructions given to annotators, the characteristics of the annotators, and their interactions can impact training data. This study…

机器学习 · 统计学 2024-01-23 Christoph Kern , Stephanie Eckman , Jacob Beck , Rob Chew , Bolei Ma , Frauke Kreuter

Depression is a global health concern with a critical need for increased patient screening. Speech technology offers advantages for remote screening but must perform robustly across patients. We have described two deep learning models…

音频与语音处理 · 电气工程与系统科学 2024-12-30 Y. Lu , A. Harati , T. Rutowski , R. Oliveira , P. Chlebek , E. Shriberg

Dialogue summarization task involves summarizing long conversations while preserving the most salient information. Real-life dialogues often involve naturally occurring variations (e.g., repetitions, hesitations) and existing dialogue…

计算与语言 · 计算机科学 2023-11-16 Ankita Gupta , Chulaka Gunasekara , Hui Wan , Jatin Ganhotra , Sachindra Joshi , Marina Danilevsky

Supervised masking approaches in the time-frequency domain aim to employ deep neural networks to estimate a multiplicative mask to extract clean speech. This leads to a single estimate for each input without any guarantees or measures of…

音频与语音处理 · 电气工程与系统科学 2023-05-16 Huajian Fang , Dennis Becker , Stefan Wermter , Timo Gerkmann

Speech signals encode emotional, linguistic, and pathological information within a shared acoustic channel; however, disentanglement is typically assessed indirectly through downstream task performance. We introduce an information-theoretic…

声音 · 计算机科学 2026-02-25 Bipasha Kashyap , Björn W. Schuller , Pubudu N. Pathirana

Self-supervised Speech Models (S3Ms) have been proven successful in many speech downstream tasks, like ASR. However, how pre-training data affects S3Ms' downstream behavior remains an unexplored issue. In this paper, we study how…

音频与语音处理 · 电气工程与系统科学 2022-04-27 Yen Meng , Yi-Hui Chou , Andy T. Liu , Hung-yi Lee

Emotional text-to-speech (TTS) systems sturggle to capture the full spectrum of human emotions due to the inherent complexity of emotional expressions and the limited coverage of existing emotion labels. To address this, we propose a…

音频与语音处理 · 电气工程与系统科学 2026-01-21 Kun Zhou , You Zhang , Dianwen Ng , Shengkui Zhao , Hao Wang , Bin Ma

We propose a novel approach for spoofed speech characterization through explainable probabilistic attribute embeddings. In contrast to high-dimensional raw embeddings extracted from a spoofing countermeasure (CM) whose dimensions are not…

音频与语音处理 · 电气工程与系统科学 2024-09-18 Manasi Chhibber , Jagabandhu Mishra , Hyejin Shim , Tomi H. Kinnunen

Large audio-language models (LALMs) unify speech and text processing, but their robustness in noisy real-world settings remains underexplored. We investigate how irrelevant audio, such as silence, synthetic noise, and environmental sounds,…

声音 · 计算机科学 2026-04-28 Chen-An Li , Tzu-Han Lin , Hung-yi Lee

Emotion and intent recognition from speech is essential and has been widely investigated in human-computer interaction. The rapid development of social media platforms, chatbots, and other technologies has led to a large volume of speech…

声音 · 计算机科学 2025-07-11 Zhao Ren , Rathi Adarshi Rammohan , Kevin Scheck , Sheng Li , Tanja Schultz

Recent works have proposed neural models for dialog act classification in spoken dialogs. However, they have not explored the role and the usefulness of acoustic information. We propose a neural model that processes both lexical and…

计算与语言 · 计算机科学 2018-03-05 Daniel Ortega , Ngoc Thang Vu

Speech enhancement in hearing aids remains a difficult task in nonstationary acoustic environments, mainly because current signal processing algorithms rely on fixed, manually tuned parameters that cannot adapt in situ to different users or…

Diagnostic procedures for ASD (autism spectrum disorder) involve semi-naturalistic interactions between the child and a clinician. Computational methods to analyze these sessions require an end-to-end speech and language processing pipeline…

音频与语音处理 · 电气工程与系统科学 2019-10-28 Rimita Lahiri , Manoj Kumar , Somer Bishop , Shrikanth Narayanan