English
Related papers

Related papers: Can you hear me $\textit{now}$? Sensitive comparis…

200 papers

Human-machine interaction is increasingly dependent on speech communication. Machine Learning models are usually applied to interpret human speech commands. However, these models can be fooled by adversarial examples, which are inputs…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-15 Jon Vadillo , Roberto Santana

With the rise of machines to human-level performance in complex recognition tasks, a growing amount of work is directed towards comparing information processing in humans and machines. These studies are an exciting chance to learn about one…

Computer Vision and Pattern Recognition · Computer Science 2021-04-14 Christina M. Funke , Judy Borowski , Karolina Stosio , Wieland Brendel , Thomas S. A. Wallis , Matthias Bethge

Machine Listening, as usually formalized, attempts to perform a task that is, from our perspective, fundamentally human-performable, and performed by humans. Current automated models of Machine Listening vary from purely data-driven…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-27 Laurie M. Heller , Benjamin Elizalde , Bhiksha Raj , Soham Deshmukh

Speech is a common and effective way of communication between humans, and modern consumer devices such as smartphones and home hubs are equipped with deep learning based accurate automatic speech recognition to enable natural interaction…

Computation and Language · Computer Science 2018-01-03 Moustafa Alzantot , Bharathan Balaji , Mani Srivastava

Recognizing a basic difference between the semiotics of humans and machines presents a possibility to overcome the shortcomings of current speech assistive devices. For the machine, the meaning of a (human) utterance is defined by its own…

Computation and Language · Computer Science 2023-06-12 Peter beim Graben , Markus Huber-Liebl , Peter Klimczak , Günther Wirsching

Humans are surrounded by audio signals that include both speech and non-speech sounds. The recognition and understanding of speech and non-speech audio events, along with a profound comprehension of the relationship between them, constitute…

Sound · Computer Science 2023-12-12 Yuan Gong , Alexander H. Liu , Hongyin Luo , Leonid Karlinsky , James Glass

Natural Language Processing (NLP) models based on Machine Learning (ML) are susceptible to adversarial attacks -- malicious algorithms that imperceptibly modify input text to force models into making incorrect predictions. However,…

Computation and Language · Computer Science 2023-05-26 Salijona Dyrmishi , Salah Ghamizi , Maxime Cordy

Natural and artificial audition can in principle acquire different solutions to a given problem. The constraints of the task, however, can nudge the cognitive science and engineering of audition to qualitatively converge, suggesting that a…

Sound · Computer Science 2023-04-20 Federico Adolfi , Jeffrey S. Bowers , David Poeppel

How similar is the human mind to the sophisticated machine-learning systems that mirror its performance? Models of object categorization based on convolutional neural networks (CNNs) have achieved human-level benchmarks in assigning known…

Computer Vision and Pattern Recognition · Computer Science 2019-08-27 Zhenglong Zhou , Chaz Firestone

Current speech translation systems, while having achieved impressive accuracies, are rather static in their behavior and do not adapt to real-world situations in ways human interpreters do. In order to improve their practical usefulness and…

Computation and Language · Computer Science 2025-08-12 Matthias Sperber , Maureen de Seyssel , Jiajun Bao , Matthias Paulik

The human ability to recognize when an object belongs or does not belong to a particular vision task outperforms all open set recognition algorithms. Human perception as measured by the methods and procedures of visual psychophysics from…

Computer Vision and Pattern Recognition · Computer Science 2023-04-26 Jin Huang , Derek Prijatelj , Justin Dulay , Walter Scheirer

In this paper, we present a data set and methods to compare speech processing models and human behaviour on a phone discrimination task. We provide Perceptimatic, an open data set which consists of French and English speech stimuli, as well…

Computation and Language · Computer Science 2020-10-14 Juliette Millet , Ewan Dunbar

This research explores how the information of prompts interacts with the high-performing speech recognition model, Whisper. We compare its performances when prompted by prompts with correct information and those corrupted with incorrect…

Computation and Language · Computer Science 2024-09-17 Chih-Kai Yang , Kuan-Po Huang , Hung-yi Lee

Recent years have seen a boom in interest in machine learning systems that can provide a human-understandable rationale for their predictions or decisions. However, exactly what kinds of explanation are truly human-interpretable remains…

Artificial Intelligence · Computer Science 2018-02-05 Menaka Narayanan , Emily Chen , Jeffrey He , Been Kim , Sam Gershman , Finale Doshi-Velez

Recently, adversarial machine learning attacks have posed serious security threats against practical audio signal classification systems, including speech recognition, speaker recognition, and music copyright detection. Previous studies…

Sound · Computer Science 2022-07-28 Rui Duan , Zhe Qu , Shangqing Zhao , Leah Ding , Yao Liu , Zhuo Lu

Computational and human perception are often considered separate approaches for studying sound changes over time; few works have touched on the intersection of both. To fill this research gap, we provide a pioneering review contrasting…

Computation and Language · Computer Science 2024-07-09 Siqi He , Wei Zhao

Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). However, whether these models can…

Sound · Computer Science 2025-12-01 Chen Li , Peiji Yang , Yicheng Zhong , Jianxing Yu , Zhisheng Wang , Zihao Gou , Wenqing Chen , Jian Yin

Many speech processing methods based on deep learning require an automatic and differentiable audio metric for the loss function. The DPAM approach of Manocha et al. learns a full-reference metric trained directly on human judgments, and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-11 Pranay Manocha , Zeyu Jin , Richard Zhang , Adam Finkelstein

Recent work in automatic recognition of conversational telephone speech (CTS) has achieved accuracy levels comparable to human transcribers, although there is some debate how to precisely quantify human performance on this task, using the…

Computation and Language · Computer Science 2022-02-22 Andreas Stolcke , Jasha Droppo

With the advent of deep learning methods, Neural Machine Translation (NMT) systems have become increasingly powerful. However, deep learning based systems are susceptible to adversarial attacks, where imperceptible changes to the input can…

Computation and Language · Computer Science 2023-06-27 Vyas Raina , Mark Gales
‹ Prev 1 2 3 10 Next ›