中文
相关论文

相关论文: Scene Recognition Through Visual and Acoustic Cues…

200 篇论文

We present a novel problem setting in zero-shot learning, zero-shot object recognition and detection in the context. Contrary to the traditional zero-shot learning methods, which simply infers unseen categories by transferring knowledge…

计算机视觉与模式识别 · 计算机科学 2019-04-25 Ruotian Luo , Ning Zhang , Bohyung Han , Linjie Yang

Environmental sound recognition (ESR) is an emerging research topic in audio pattern recognition. Many tasks are presented to resort to computational models for ESR in real-life applications. However, current models are usually designed for…

音频与语音处理 · 电气工程与系统科学 2023-11-22 Jisheng Bai , Jianfeng Chen , Mou Wang , Muhammad Saad Ayub

Continual learning aims to refine model parameters for new tasks while retaining knowledge from previous tasks. Recently, prompt-based learning has emerged to leverage pre-trained models to be prompted to learn subsequent tasks without the…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Jisu Han , Jaemin Na , Wonjun Hwang

Navigation is a rich and well-grounded problem domain that drives progress in many different areas of research: perception, planning, memory, exploration, and optimisation in particular. Historically these challenges have been separately…

Grounded Situation Recognition (GSR) is capable of recognizing and interpreting visual scenes in a contextually intuitive way, yielding salient activities (verbs) and the involved entities (roles) depicted in images. In this work, we focus…

计算机视觉与模式识别 · 计算机科学 2023-07-18 Ruiping Liu , Jiaming Zhang , Kunyu Peng , Junwei Zheng , Ke Cao , Yufan Chen , Kailun Yang , Rainer Stiefelhagen

The concept of 3D scene graphs is increasingly recognized as a powerful semantic and hierarchical representation of the environment. Current approaches often address this at a coarse, object-level resolution. In contrast, our goal is to…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Dennis Rotondi , Fabio Scaparro , Hermann Blum , Kai O. Arras

Semantic segmentation and activity classification are key components to creating intelligent surgical systems able to understand and assist clinical workflow. In the Operating Room, semantic segmentation is at the core of creating robots…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Idris Hamoud , Alexandros Karargyris , Aidean Sharghi , Omid Mohareri , Nicolas Padoy

Upper-limb exoskeletons are primarily designed to provide assistive support by accurately interpreting and responding to human intentions. In home-care scenarios, exoskeletons are expected to adapt their assistive configurations based on…

机器人学 · 计算机科学 2025-08-15 Yu Chen , Shu Miao , Chunyu Wu , Jingsong Mu , Bo OuYang , Xiang Li

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

音频与语音处理 · 电气工程与系统科学 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram

A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over possible actions in context. However, most existing benchmarks for multimodal foundation models…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Chuhan Li , Ruilin Han , Joy Hsu , Yongyuan Liang , Rajiv Dhawan , Jiajun Wu , Ming-Hsuan Yang , Xin Eric Wang

Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of objects to determine interactable regions. However, extending…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Nghia Vu , Tuong Do , Khang Nguyen , Baoru Huang , Nhat Le , Binh Xuan Nguyen , Erman Tjiputra , Quang D. Tran , Ravi Prakash , Te-Chuan Chiu , Anh Nguyen

Multimodal scene search of conversations is essential for unlocking valuable insights into social dynamics and enhancing our communication. While experts in conversational analysis have their own knowledge and skills to find key scenes, a…

人机交互 · 计算机科学 2024-02-20 Riku Arakawa , Kiyosu Maeda , Hiromu Yakura

The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, existing approaches…

多媒体 · 计算机科学 2023-08-21 Sung Jin Um , Dongjin Kim , Jung Uk Kim

In this study, we try to address the problem of leveraging visual signals to improve Automatic Speech Recognition (ASR), also known as visual context-aware ASR (VC-ASR). We explore novel VC-ASR approaches to leverage video and text…

音频与语音处理 · 电气工程与系统科学 2020-11-10 Shahram Ghorbani , Yashesh Gaur , Yu Shi , Jinyu Li

In computational reinforcement learning, a growing body of work seeks to construct an agent's perception of the world through predictions of future sensations; predictions about environment observations are used as additional input features…

机器学习 · 计算机科学 2022-06-15 Alexandra Kearney , Anna Koop , Johannes Günther , Patrick M. Pilarski

Robots typically possess sensors of different modalities, such as colour cameras, inertial measurement units, and 3D laser scanners. Often, solving a particular problem becomes easier when more than one modality is used. However, while…

计算机视觉与模式识别 · 计算机科学 2017-01-10 Charika De Alvis , Lionel Ott , Fabio Ramos

Being able to explore an environment and understand the location and type of all objects therein is important for indoor robotic platforms that must interact closely with humans. However, it is difficult to evaluate progress in this area…

机器人学 · 计算机科学 2020-09-14 David Hall , Ben Talbot , Suman Raj Bista , Haoyang Zhang , Rohan Smith , Feras Dayoub , Niko Sünderhauf

Healthcare robotics requires robust multimodal perception and reasoning to ensure safety in dynamic clinical environments. Current Vision-Language Models (VLMs) demonstrate strong general-purpose capabilities but remain limited in temporal…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Saurav Jha , Stefan K. Ehrlich

Deep learning models have achieved state-of-the- art performance in recognizing human activities, but often rely on utilizing background cues present in typical computer vision datasets that predominantly have a stationary camera. If these…

机器人学 · 计算机科学 2017-09-20 Fahimeh Rezazadegan , Sareh Shirazi , Ben Upcroft , Michael Milford

Building models that can understand and reason about 3D scenes is difficult owing to the lack of data sources for 3D supervised training and large-scale training regimes. In this work we ask - How can the knowledge in a pre-trained language…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Shivam Chandhok
‹ 上一页 1 8 9 10 下一页 ›