English
Related papers

Related papers: Are words equally surprising in audio and audio-vi…

200 papers

Transformer-based language models have shown an excellent ability to effectively capture and utilize contextual information. Although various analysis techniques have been used to quantify and trace the contribution of single contextual…

Computation and Language · Computer Science 2024-10-07 Hamidreza Amirzadeh , Afra Alishahi , Hosein Mohebbi

This work investigates whether modern speech models are sensitive to prosodic emphasis - whether they encode emphasized and neutral words in systematically different ways. Prior work typically relies on isolated acoustic correlates (e.g.,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-18 Shaun Cassini , Thomas Hain , Anton Ragni

In daily interactions, emotions are frequently conveyed and triggered through verbal exchanges. Sometimes, we must modulate our emotional reactions to align with societal norms. Among the emotional words, taboo words represent a specific…

Neurons and Cognition · Quantitative Biology 2024-10-29 Parisa Ahmadi Ghomroudi , Michele Scaltritti , Bianca Monachesi , Peera Wongupparaj , Remo Job , Alessandro Grecucci

Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems. However, existing methods are either not entirely end-to-end or do not train joint…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-23 Jiachen Lian , Alexei Baevski , Wei-Ning Hsu , Michael Auli

Language models (LM) are very powerful in lipreading systems. Language models built upon the ground truth utterances of datasets learn grammar and structure rules of words and sentences (the latter in the case of continuous speech).…

Audio and Speech Processing · Electrical Eng. & Systems 2018-09-19 Helen L Bear

Language interfaces with many other cognitive domains. This paper explores how interactions at these interfaces can be studied with deep learning methods, focusing on the relation between language emergence and visual perception. To model…

Social and Information Networks · Computer Science 2023-01-11 Xenia Ohmer , Michael Marino , Michael Franke , Peter König

Interpreting the effects of variants within the human genome and proteome is essential for analysing disease risk, predicting medication response, and developing personalised health interventions. Due to the intrinsic similarities between…

Computation and Language · Computer Science 2025-03-17 Megha Hegde , Jean-Christophe Nebel , Farzana Rahman

Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. Similar alignment has been reported in Whisper-style speech encoders, based on spoken translation retrieval using representational…

Computation and Language · Computer Science 2026-04-07 Ryan Soh-Eun Shim , Domenico De Cristofaro , Chengzhi Martin Hu , Alessandro Vietti , Barbara Plank

We introduce a seemingly impossible task: given only an audio clip of someone speaking, decide which of two face images is the speaker. In this paper we study this, and a number of related cross-modal tasks, aimed at answering the question:…

Computer Vision and Pattern Recognition · Computer Science 2018-04-04 Arsha Nagrani , Samuel Albanie , Andrew Zisserman

Acoustic word embeddings are fixed-dimensional representations of variable-length speech segments. In settings where unlabelled speech is the only available resource, such embeddings can be used in "zero-resource" speech search, indexing…

Computation and Language · Computer Science 2020-02-24 Herman Kamper , Yevgen Matusevych , Sharon Goldwater

Multimodal learning allows us to leverage information from multiple sources (visual, acoustic and text), similar to our experience of the real world. However, it is currently unclear to what extent auxiliary modalities improve performance…

Computation and Language · Computer Science 2020-01-01 Tejas Srinivasan , Ramon Sanabria , Florian Metze

We investigate how word meanings are represented in the transformer language models. Specifically, we focus on whether transformer models employ something analogous to a lexical store - where each word has an entry that contains semantic…

Computation and Language · Computer Science 2025-08-19 Jumbly Grindrod , Peter Grindrod

Recently multimodal transformer models have gained popularity because their performance on language and vision tasks suggest they learn rich visual-linguistic representations. Focusing on zero-shot image retrieval tasks, we study three…

Computation and Language · Computer Science 2021-02-02 Lisa Anne Hendricks , John Mellor , Rosalia Schneider , Jean-Baptiste Alayrac , Aida Nematzadeh

In this paper, we study the associations between human faces and voices. Audiovisual integration, specifically the integration of facial and vocal information is a well-researched area in neuroscience. It is shown that the overlapping…

Computer Vision and Pattern Recognition · Computer Science 2018-11-05 Changil Kim , Hijung Valentina Shin , Tae-Hyun Oh , Alexandre Kaspar , Mohamed Elgharib , Wojciech Matusik

Recent studies suggest that transformer-based vision-language models (VLMs) capture the multimodality of concept processing in the human brain. However, a systematic evaluation exploring different types of VLM architectures and the role…

Computation and Language · Computer Science 2026-01-23 Anna Bavaresco , Marianne de Heer Kloots , Sandro Pezzelle , Raquel Fernández

Transformer-based models have gained increasing popularity achieving state-of-the-art performance in many research fields including speech translation. However, Transformer's quadratic complexity with respect to the input sequence length…

Computation and Language · Computer Science 2023-10-19 Sara Papi , Marco Gaido , Matteo Negri , Marco Turchi

Recent advances in multimodal LLMs, have led to several video-text models being proposed for critical video-related tasks. However, most of the previous works support visual input only, essentially muting the audio signal in the video. Few…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Shivprasad Sagare , Hemachandran S , Kinshuk Sarabhai , Prashant Ullegaddi , Rajeshkumar SA

In this work, we explore the dependencies between speaker recognition and emotion recognition. We first show that knowledge learned for speaker recognition can be reused for emotion recognition through transfer learning. Then, we show the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-13 Raghavendra Pappagari , Tianzi Wang , Jesus Villalba , Nanxin Chen , Najim Dehak

The abundance and ease of utilizing sound, along with the fact that auditory clues reveal a plethora of information about what happens in a scene, make the audio-visual space an intuitive choice for representation learning. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Mahdi M. Kalayeh , Shervin Ardeshir , Lingyi Liu , Nagendra Kamath , Ashok Chandrashekar

Predictive coding theory suggests that the brain continuously anticipates upcoming words to optimize language processing, but the neural mechanisms remain unclear, particularly in naturalistic speech. Here, we simultaneously recorded EEG…

Neurons and Cognition · Quantitative Biology 2025-06-11 Nikola Kölbl , Konstantin Tziridis , Andreas Maier , Thomas Kinfe , Ricardo Chavarriaga , Achim Schilling , Patrick Krauss
‹ Prev 1 8 9 10 Next ›