中文
相关论文

相关论文: Multimodal Contextualized Semantic Parsing from Sp…

200 篇论文

Audio-visual speech enhancement (AVSE) has been found to be particularly useful at low signal-to-noise (SNR) ratios due to the immunity of the visual features to acoustic noise. However, a significant gap exists in AVSE methods tailored to…

音频与语音处理 · 电气工程与系统科学 2025-10-21 Danielle Yaffe , Ferdinand Campe , Prachi Sharma , Dorothea Kolossa , Boaz Rafaely

Due to the difficulty of automatically mapping visual features with semantic descriptors, state-of-the-art frameworks have exhibited poor performance in terms of coverage and effectiveness for indexing the visual content. This prompted us…

多媒体 · 计算机科学 2020-04-28 M. Belkhatir

Healthcare robotics requires robust multimodal perception and reasoning to ensure safety in dynamic clinical environments. Current Vision-Language Models (VLMs) demonstrate strong general-purpose capabilities but remain limited in temporal…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Saurav Jha , Stefan K. Ehrlich

The intelligent dialogue system, aiming at communicating with humans harmoniously with natural language, is brilliant for promoting the advancement of human-machine interaction in the era of artificial intelligence. With the gradually…

人工智能 · 计算机科学 2022-07-05 Hao Wang , Bin Guo , Yating Zeng , Yasan Ding , Chen Qiu , Ying Zhang , Lina Yao , Zhiwen Yu

Detecting text in natural scenes remains challenging, particularly for diverse scripts and arbitrarily shaped instances where visual cues alone are often insufficient. Existing methods do not fully leverage semantic context. This paper…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Mohammed-En-Nadhir Zighem , Abdenour Hadid

The usage of automatic speech recognition (ASR) systems are becoming omnipresent ranging from personal assistant to chatbots, home, and industrial automation systems, etc. Modern robots are also equipped with ASR capabilities for…

音频与语音处理 · 电气工程与系统科学 2022-10-25 Pradip Pramanick , Chayan Sarkar

In this paper, we present WISE, an open-source audiovisual search engine which integrates a range of multimodal retrieval capabilities into a single, practical tool accessible to users without machine learning expertise. WISE supports…

信息检索 · 计算机科学 2026-02-16 Prasanna Sridhar , Horace Lee , David M. S. Pinto , Andrew Zisserman , Abhishek Dutta

Advancements at the intersection of computer vision and natural language processing are crucial for applications like assistive tech, multimedia querying, and robotics. This dissertation proposes novel architectures to improve intelligent…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Van Quang Nguyen

General scene perception has progressed from object recognition toward open-vocabulary grounding, part localization, and affordance prediction. Yet these capabilities are often realized as isolated predictions that localize objects, parts,…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Pengxin Xu , Xincheng Lin , Luping Xiao , Qing Jiang , Meishan Zhang , Hao Fei , Shanghang Zhang , Xingyu Chen

Multimodal scene search of conversations is essential for unlocking valuable insights into social dynamics and enhancing our communication. While experts in conversational analysis have their own knowledge and skills to find key scenes, a…

人机交互 · 计算机科学 2024-02-20 Riku Arakawa , Kiyosu Maeda , Hiromu Yakura

Most existing remote sensing instance segmentation approaches are designed for close-vocabulary prediction, limiting their ability to recognize novel categories or generalize across datasets. This restricts their applicability in diverse…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Shiqi Huang , Shuting He , Huaiyuan Qin , Bihan Wen

Automatic video commentary systems are widely used on multimedia social media platforms to extract factual information about video content. However, current systems may overlook essential para-linguistic cues, including emotion and…

人机交互 · 计算机科学 2025-06-23 Qixin Wang , Songtao Zhou , Zeyu Jin , Chenglin Guo , Shikun Sun , Xiaoyu Qin

We present the Verse library with the aim of making hybrid system verification more usable for multi-agent scenarios. In Verse, decision making agents move in a map and interact with each other through sensors. The decision logic for each…

软件工程 · 计算机科学 2023-01-24 Yangge Li , Haoqing Zhu , Katherine Braught , Keyi Shen , Sayan Mitra

The visual world is fundamentally compositional. Visual scenes are defined by the composition of objects and their relations. Hence, it is essential for computer vision systems to reflect and exploit this compositionality to achieve robust…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Shuhao Fu , Andrew Jun Lee , Anna Wang , Ida Momennejad , Trevor Bihl , Hongjing Lu , Taylor W. Webb

Semantic communication (SemCom) powered by generative artificial intelligence enables highly efficient and reliable information transmission. However, it still necessitates the transmission of substantial amounts of data when dealing with…

信息论 · 计算机科学 2025-06-17 Guojun Huang , Jiancheng An , Lu Gan , Dusit Niyato , Mérouane Debbah , Tie Jun Cui

In VR interactions with embodied conversational agents, users' emotional intent is often conveyed more by how something is said than by what is said. However, most VR agent pipelines rely on speech-to-text processing, discarding prosodic…

人机交互 · 计算机科学 2026-03-11 SangYeop Jeong , Yeongseo Na , Seung Gyu Jeong , Jin-Woo Jeong , Seong-Eun Kim

This work presents iMiGUE-Speech, an extension of the iMiGUE dataset that provides a spontaneous affective corpus for studying emotional and affective states. The new release focuses on speech and enriches the original dataset with…

音频与语音处理 · 电气工程与系统科学 2026-02-26 Sofoklis Kakouros , Fang Kang , Haoyu Chen

Semantic communications have been utilized to execute numerous intelligent tasks by transmitting task-related semantic information instead of bits. In this article, we propose a semantic-aware speech-to-text transmission system for the…

音频与语音处理 · 电气工程与系统科学 2024-10-08 Zhenzi Weng , Zhijin Qin , Huiqiang Xie , Xiaoming Tao , Khaled B. Letaief

With the widespread use of intelligent systems, such as smart speakers, addressee recognition has become a concern in human-computer interaction, as more and more people expect such systems to understand complicated social scenes, including…

人工智能 · 计算机科学 2018-09-13 Thao Minh Le , Nobuyuki Shimizu , Takashi Miyazaki , Koichi Shinoda

Despite recent advances in multimodal content generation enabled by vision-language models (VLMs), their ability to reason about and generate structured 3D scenes remains largely underexplored. This limitation constrains their utility in…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang