中文
相关论文

相关论文: Structure from Silence: Learning Scene Structure f…

200 篇论文

While the ambient intelligence (AmI) systems we encounter in our daily lives, including security monitoring and energy-saving systems, typically serve pragmatic purposes, we wonder how we can design and implement ambient artificial…

人机交互 · 计算机科学 2025-11-04 Chengzhi Zhang , Dashiel Carrera , Daksh Kapoor , Jasmine Kaur , Jisu Kim , Brian Magerko

With the recent advancements in AI, Intelligent Virtual Assistants (IVA) have become a ubiquitous part of every home. Going forward, we are witnessing a confluence of vision, speech and dialog system technologies that are enabling the IVAs…

计算与语言 · 计算机科学 2018-12-21 Shachi H Kumar , Eda Okur , Saurav Sahay , Juan Jose Alvarado Leanos , Jonathan Huang , Lama Nachman

The recent surge in popularity of diffusion models for image generation has brought new attention to the potential of these models in other areas of media generation. One area that has yet to be fully explored is the application of…

声音 · 计算机科学 2023-02-01 Flavio Schneider

In this study, we aim to determine if generalized sounds and music can share a common emotional space, improving predictions of emotion in terms of arousal and valence. We propose the use of multiple datasets as a multi-domain learning…

声音 · 计算机科学 2024-08-15 Federico Simonetta , Francesca Certo , Stavros Ntalampiras

Automatic image synthesis research has been rapidly growing with deep networks getting more and more expressive. In the last couple of years, we have observed images of digits, indoor scenes, birds, chairs, etc. being automatically…

计算机视觉与模式识别 · 计算机科学 2016-12-02 Levent Karacan , Zeynep Akata , Aykut Erdem , Erkut Erdem

Imagine being able to listen to the birds chirping in a park without hearing the chatter from other hikers, or being able to block out traffic noise on a busy street while still being able to hear emergency sirens and car honks. We…

声音 · 计算机科学 2023-11-02 Bandhav Veluri , Malek Itani , Justin Chan , Takuya Yoshioka , Shyamnath Gollakota

Modeling imaging sensor noise is a fundamental problem for image processing and computer vision applications. While most previous works adopt statistical noise models, real-world noise is far more complicated and beyond what these models…

计算机视觉与模式识别 · 计算机科学 2020-10-20 Ke-Chi Chang , Ren Wang , Hung-Jin Lin , Yu-Lun Liu , Chia-Ping Chen , Yu-Lin Chang , Hwann-Tzong Chen

Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainly explored the problem from a localization perspective.…

计算机视觉与模式识别 · 计算机科学 2023-09-20 Arda Senocak , Hyeonggon Ryu , Junsik Kim , Tae-Hyun Oh , Hanspeter Pfister , Joon Son Chung

Robotic perception is becoming a key technology for navigation aids, especially helping individuals with visual impairments through spatial sonification. This paper introduces a mapping representation that accurately captures scene geometry…

机器人学 · 计算机科学 2025-04-18 Lan Wu , Craig Jin , Monisha Mushtary Uttsha , Teresa Vidal-Calleja

In this paper, we investigate how to learn rich and robust feature representations for audio classification from visual data and acoustic images, a novel audio data modality. Former models learn audio representations from raw signals or…

计算机视觉与模式识别 · 计算机科学 2020-02-12 Andrés F. Pérez , Valentina Sanguineti , Pietro Morerio , Vittorio Murino

While experimentation with synthetic stimuli in abstracted listening situations has a long standing and successful history in hearing research, an increased interest exists on closing the remaining gap towards real-life listening by…

Acoustic event detection and scene classification are major research tasks in environmental sound analysis, and many methods based on neural networks have been proposed. Conventional methods have addressed these tasks separately; however,…

We introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple audio sources…

计算机视觉与模式识别 · 计算机科学 2021-08-27 Sagnik Majumder , Ziad Al-Halah , Kristen Grauman

The abundance and ease of utilizing sound, along with the fact that auditory clues reveal a plethora of information about what happens in a scene, make the audio-visual space an intuitive choice for representation learning. In this paper,…

计算机视觉与模式识别 · 计算机科学 2022-05-03 Mahdi M. Kalayeh , Shervin Ardeshir , Lingyi Liu , Nagendra Kamath , Ashok Chandrashekar

Semantic understanding of scenes in three-dimensional space (3D) is a quintessential part of robotics oriented applications such as autonomous driving as it provides geometric cues such as size, orientation and true distance of separation…

计算机视觉与模式识别 · 计算机科学 2019-11-01 Kartik Srivastava , Akash Kumar Singh , Guruprasad M. Hegde

In this work, we investigate the knowledge learned in the embeddings of multimodal-BERT models. More specifically, we probe their capabilities of storing the grammatical structure of linguistic data and the structure learned over objects in…

计算与语言 · 计算机科学 2022-03-18 Victor Milewski , Miryam de Lhoneux , Marie-Francine Moens

Monitoring remote forests is a global challenge central to climate mitigation and biodiversity conservation, yet satellite observations are frequently limited by weather, dense canopies, and solar dependency. Here we show that passive…

Procedural noise is a fundamental component of computer graphics pipelines, offering a flexible way to generate textures that exhibit "natural" random variation. Many different types of noise exist, each produced by a separate algorithm. In…

Scene Text Recognition (STR) models have achieved high performance in recent years on benchmark datasets where text images are presented with minimal noise. Traditional STR recognition pipelines take a cropped image as sole input and…

计算机视觉与模式识别 · 计算机科学 2022-10-21 Joshua Cesare Placidi , Yishu Miao , Zixu Wang , Lucia Specia

Humans exhibit incredibly high levels of multi-modal understanding - combining visual cues with read, or heard knowledge comes easy to us and allows for very accurate interaction with the surrounding environment. Various simulation…

机器人学 · 计算机科学 2022-06-22 Michal Nazarczuk , Tony Ng , Krystian Mikolajczyk