English
Related papers

Related papers: A Unified Quantitative Model of Vision and Auditio…

200 papers

Sequential scientific data span many resolutions and domains, and unifying them into a common representation is a key step toward developing foundation models for the sciences. Astronomical spectra exemplify this challenge: massive surveys…

Conventional audio-visual models have independent audio and video branches. In this work, we unify the audio and visual branches by designing a Unified Audio-Visual Model (UAVM). The UAVM achieves a new state-of-the-art audio-visual event…

Computer Vision and Pattern Recognition · Computer Science 2023-02-17 Yuan Gong , Alexander H. Liu , Andrew Rouditchenko , James Glass

Music Sight Reading is a complex process in which when it is occurred in the brain some learning attributes would be emerged. Besides giving a model based on actor-critic method in the Reinforcement Learning, the agent is considered to have…

Machine Learning · Computer Science 2011-11-21 Keyvan Yahya , Pouyan Rafiei Fard

We present a novel approach to multilingual audio-visual speech recognition tasks by introducing a single model on a multilingual dataset. Motivated by a human cognitive system where humans can intuitively distinguish different languages…

Multimedia · Computer Science 2023-10-24 Joanna Hong , Se Jin Park , Yong Man Ro

Vision-language large models are moving toward the unification of visual understanding and visual generation tasks. However, whether generation can enhance understanding is still under-explored on large data scale. In this work, we analysis…

Computation and Language · Computer Science 2026-01-01 Fengjiao Chen , Minhao Jing , Weitao Lu , Yan Feng , Xiaoyu Li , Xuezhi Cao

The sound of crashing waves, the roar of fast-moving cars -- sound conveys important information about the objects in our surroundings. In this work, we show that ambient sounds can be used as a supervisory signal for learning visual…

Computer Vision and Pattern Recognition · Computer Science 2017-12-21 Andrew Owens , Jiajun Wu , Josh H. McDermott , William T. Freeman , Antonio Torralba

Adaptive streaming of 360-degree video relies on viewport prediction to allocate bandwidth efficiently. Current approaches predominantly use visual saliency or historical gaze patterns, neglecting the role of spatial audio in guiding user…

Multimedia · Computer Science 2026-01-07 Arman Nik Khah , Ravi Prakash

Respiratory diseases remain major global health challenges, and traditional auscultation is often limited by subjectivity, environmental noise, and inter-clinician variability. This study presents an explainable multimodal deep learning…

Sound · Computer Science 2025-12-02 S M Asiful Islam Saky , Md Rashidul Islam , Md Saiful Arefin , Shahaba Alam

Current generative models are able to generate high-quality artefacts but have been shown to struggle with compositional reasoning, which can be defined as the ability to generate complex structures from simpler elements. In this paper, we…

Machine Learning · Computer Science 2024-08-20 Giovanni Bindi , Philippe Esling

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak

Auditory perception involves cues in the monaural auditory pathways as well as binaural cues based on differences between the ears. So far auditory models have often focused on either monaural or binaural experiments in isolation. Although…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-24 Thomas Biberger , Stephan D. Ewert

Human auditory perception is compositional in nature -- we identify auditory streams from auditory scenes with multiple sound events. However, such auditory scenes are typically represented using clip-level representations that do not…

Sound · Computer Science 2025-03-04 Sripathi Sridhar , Mark Cartwright

Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on…

Referring expressions are natural language constructions used to identify particular objects within a scene. In this paper, we propose a unified framework for the tasks of referring expression comprehension and generation. Our model is…

Computer Vision and Pattern Recognition · Computer Science 2017-04-19 Licheng Yu , Hao Tan , Mohit Bansal , Tamara L. Berg

This work reviews the human auditory system, elucidating some of the specialized mechanisms and non-linear pathways along the chain of events between physical sound and its perception. Customary relationships between frequency, time, and…

Neurons and Cognition · Quantitative Biology 2023-08-01 Milind N. Kunchur

We propose a new framework for extracting visual information about a scene only using audio signals. Audio-based methods can overcome some of the limitations of vision-based methods i.e., they do not require "line-of-sight", are robust to…

Computer Vision and Pattern Recognition · Computer Science 2022-09-14 Fabrizio Pedersoli , Dryden Wiebe , Amin Banitalebi , Yong Zhang , George Tzanetakis , Kwang Moo Yi

Longitudinal imaging is capable of capturing the static ana\-to\-mi\-cal structures and the dynamic changes of the morphology resulting from aging or disease progression. Self-supervised learning allows to learn new representation from…

Image and Video Processing · Electrical Eng. & Systems 2019-10-25 Antoine Rivail , Ursula Schmidt-Erfurth , Wolf-Dieter Vogl , Sebastian M. Waldstein , Sophie Riedl , Christoph Grechenig , Zhichao Wu , Hrvoje Bogunović

Current models for audio--sheet music retrieval via multimodal embedding space learning use convolutional neural networks with a fixed-size window for the input audio. Depending on the tempo of a query performance, this window captures more…

Sound · Computer Science 2018-09-18 Matthias Dorfer , Jan Hajič , Gerhard Widmer

Recognition and reasoning are two pillars of visual understanding. However, these tasks have an imbalance in focus; whereas recent advances in neural networks have shown strong empirical performance in visual recognition, there has been…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Calvin Luo , Boqing Gong , Ting Chen , Chen Sun

We present a novel approach to inspecting galaxy spectra using sound, via their direct audio representation ('spectral audification'). We discuss the potential of this as a complement to (or stand-in for) visual approaches. We surveyed 58…

Instrumentation and Methods for Astrophysics · Physics 2023-06-21 James W. Trayford , C. M. Harrison , R. C. Hinz , M. Kavanagh Blatt , S. Dougherty , A. Girdhar