English
Related papers

Related papers: Learning in Audio-visual Context: A Review, Analys…

200 papers

Audio-text relevance learning refers to learning the shared semantic properties of audio samples and textual descriptions. The standard approach uses binary relevances derived from pairs of audio samples and their human-provided captions,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-28 Huang Xie , Khazar Khorrami , Okko Räsänen , Tuomas Virtanen

Acoustic event detection and scene classification are major research tasks in environmental sound analysis, and many methods based on neural networks have been proposed. Conventional methods have addressed these tasks separately; however,…

Sign language visual recognition from continuous multi-modal streams is still one of the most challenging fields. Recent advances in human actions recognition are exploiting the ascension of GPU-based learning from massive data, and are…

Computer Vision and Pattern Recognition · Computer Science 2020-09-23 Bassem Seddik , Najoua Essoukri Ben Amara

Perceptual video quality assessment plays a vital role in the field of video processing due to the existence of quality degradations introduced in various stages of video signal acquisition, compression, transmission and display. With the…

Multimedia · Computer Science 2024-02-07 Xiongkuo Min , Huiyu Duan , Wei Sun , Yucheng Zhu , Guangtao Zhai

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

Sound · Computer Science 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

In the context of Audio Visual Question Answering (AVQA) tasks, the audio visual modalities could be learnt on three levels: 1) Spatial, 2) Temporal, and 3) Semantic. Existing AVQA methods suffer from two major shortcomings; the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Asmar Nadeem , Adrian Hilton , Robert Dawes , Graham Thomas , Armin Mustafa

In this paper, we propose to make a systematic study on machines multisensory perception under attacks. We use the audio-visual event recognition task against multimodal adversarial attacks as a proxy to investigate the robustness of…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Yapeng Tian , Chenliang Xu

Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate, and interact in the multimodal real world. In the era of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 You Qin , Kai Liu , Shengqiong Wu , Kai Wang , Shijian Deng , Yapeng Tian , Junbin Xiao , Yazhou Xing , Yinghao Ma , Bobo Li , Roger Zimmermann , Lei Cui , Furu Wei , Jiebo Luo , Hao Fei

Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short video creation, and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Zhaokai Wang , Chenxi Bao , Le Zhuo , Jingrui Han , Yang Yue , Yihong Tang , Victor Shea-Jay Huang , Yue Liao

Active visual perception refers to the ability of a system to dynamically engage with its environment through sensing and action, allowing it to modify its behavior in response to specific goals or uncertainties. Unlike passive systems that…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Yian Li , Xiaoyu Guo , Hao Zhang , Shuiwang Li , Xiaowei Dai

Text-level discourse parsing aims to unmask how two sentences in the text are related to each other. We propose the task of Visual Discourse Parsing, which requires understanding discourse relations among scenes in a video. Here we use the…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Arjun R. Akula , Song-Chun Zhu

The remarkable success of deep learning in various domains relies on the availability of large-scale annotated datasets. However, obtaining annotations is expensive and requires great effort, which is especially challenging for videos.…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Madeline C. Schiappa , Yogesh S. Rawat , Mubarak Shah

Our experience of the world is multisensory, spanning a synthesis of language, sight, sound, touch, taste, and smell. Yet, artificial intelligence has primarily advanced in digital modalities like text, vision, and audio. This paper…

Machine Learning · Computer Science 2026-01-14 Paul Pu Liang

Developing embodied agents in simulation has been a key research topic in recent years. Exciting new tasks, algorithms, and benchmarks have been developed in various simulators. However, most of them assume deaf agents in silent…

Robotics · Computer Science 2023-09-19 Ruohan Gao , Hao Li , Gokul Dharan , Zhuzhu Wang , Chengshu Li , Fei Xia , Silvio Savarese , Li Fei-Fei , Jiajun Wu

Imitating how humans move their gaze in a visual scene is a vital research problem for both visual understanding and psychology, kindling crucial applications such as building alive virtual characters. Previous studies aim to predict gaze…

Multimedia · Computer Science 2025-03-03 Xiaochuan Liu , Xin Cheng , Yuchong Sun , Xiaoxue Wu , Ruihua Song , Hao Sun , Denghao Zhang

Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on…

Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information. However, a thorough understanding of videos significantly…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yudong Yang , Jimin Zhuang , Guangzhi Sun , Changli Tang , Yixuan Li , Peihan Li , Yifan Jiang , Wei Li , Zejun Ma , Chao Zhang

Animal vocalisations and natural soundscapes are fascinating objects of study, and contain valuable evidence about animal behaviours, populations and ecosystems. They are studied in bioacoustics and ecoacoustics, with signal processing and…

Sound · Computer Science 2024-02-01 Dan Stowell

Semantic image parsing, which refers to the process of decomposing images into semantic regions and constructing the structure representation of the input, has recently aroused widespread interest in the field of computer vision. The recent…

Computer Vision and Pattern Recognition · Computer Science 2018-10-11 Lili Huang , Jiefeng Peng , Ruimao Zhang , Guanbin Li , Liang Lin

Computational and human perception are often considered separate approaches for studying sound changes over time; few works have touched on the intersection of both. To fill this research gap, we provide a pioneering review contrasting…

Computation and Language · Computer Science 2024-07-09 Siqi He , Wei Zhao
‹ Prev 1 8 9 10 Next ›