中文
相关论文

相关论文: Cross-Task Transfer for Geotagged Audiovisual Aeri…

200 篇论文

The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting and annotating…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Francesco Barbato , Matteo Caligiuri , Pietro Zanuttigh

Despite recent advancements in computer vision research, object detection in aerial images still suffers from several challenges. One primary challenge to be mitigated is the presence of multiple types of variation in aerial images, for…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Sungjune Park , Hyunjun Kim , Beomchan Park , Yong Man Ro

Remote sensing scene classification aims to assign a specific semantic label to a remote sensing image. Recently, convolutional neural networks have greatly improved the performance of remote sensing scene classification. However, some…

计算机视觉与模式识别 · 计算机科学 2022-05-24 Zhang Yue , Zheng Xiangtao , Lu Xiaoqiang

In a recent acoustic scene classification (ASC) research field, training and test device channel mismatch have become an issue for the real world implementation. To address the issue, this paper proposes a channel domain conversion using…

声音 · 计算机科学 2018-12-06 Seongkyu Mun , Suwon Shon

The demand for unmanned aerial vehicle (UAV)-based image acquisition and analysis has surged, with UAVs increasingly utilized for semantic segmentation tasks. To meet the real-time analysis requirements of UAV remote sensing missions,…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Anqi Lu , Yun Cheng , Youbing Hu , Zhiqiang Cao , Jie Liu , Zhijun Li

Existing image perception methods based on VLMs generally follow a paradigm wherein models extract and analyze image content based on user-provided textual task prompts. However, such methods face limitations when applied to UAV imagery,…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Mingning Guo , Mengwei Wu , Shaoxian Li , Haifeng Li , Chao Tao

This paper introduces a novel task in generative speech processing, Acoustic Scene Transfer (AST), which aims to transfer acoustic scenes of speech signals to diverse environments. AST promises an immersive experience in speech perception…

音频与语音处理 · 电气工程与系统科学 2024-06-19 Miseul Kim , Soo-Whan Chung , Youna Ji , Hong-Goo Kang , Min-Seok Choi

Connecting current observations with prior experiences helps robots adapt and plan in new, unseen 3D environments. Recently, 3D scene analogies have been proposed to connect two 3D scenes, which are smooth maps that align scene regions with…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Junho Kim , Young Min Kim

Scene understanding plays a critical role in enabling intelligence and autonomy in robotic systems. Traditional approaches often face challenges, including occlusions, ambiguous boundaries, and the inability to adapt attention based on…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Guodong Sun , Junjie Liu , Gaoyang Zhang , Bo Wu , Yang Zhang

Aerial imagery is increasingly used in Earth science and natural resource management as a complement to labor-intensive ground-based surveys. Aerial systems can collect overlapping images that provide multiple views of each location from…

计算机视觉与模式识别 · 计算机科学 2024-05-16 David Russell , Ben Weinstein , David Wettergreen , Derek Young

The Detection and Classification of Acoustic Scenes and Events (DCASE) 2019 challenge focuses on audio tagging, sound event detection and spatial localisation. DCASE 2019 consists of five tasks: 1) acoustic scene classification, 2) audio…

声音 · 计算机科学 2019-04-16 Qiuqiang Kong , Yin Cao , Turab Iqbal , Yong Xu , Wenwu Wang , Mark D. Plumbley

We propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a single model. Most previous 3D segmentation approaches are…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Wei Xu , Chunsheng Shi , Sifan Tu , Xin Zhou , Dingkang Liang , Xiang Bai

Object detection in aerial images is an important task in environmental, economic, and infrastructure-related tasks. One of the most prominent applications is the detection of vehicles, for which deep learning approaches are increasingly…

计算机视觉与模式识别 · 计算机科学 2021-04-08 Immanuel Weber , Jens Bongartz , Ribana Roscher

To better understand scene images in the field of remote sensing, multi-label annotation of scene images is necessary. Moreover, to enhance the performance of deep learning models for dealing with semantic scene understanding tasks, it is…

计算机视觉与模式识别 · 计算机科学 2020-10-02 Xiaoman Qi , PanPan Zhu , Yuebin Wang , Liqiang Zhang , Junhuan Peng , Mengfan Wu , Jialong Chen , Xudong Zhao , Ning Zang , P. Takis Mathiopoulos

It is undeniable that aerial/satellite images can provide useful information for a large variety of tasks. But, since these images are always looking from above, some applications can benefit from complementary information provided by other…

计算机视觉与模式识别 · 计算机科学 2020-08-05 Gabriel Machado , Edemir Ferreira , Keiller Nogueira , Hugo Oliveira , Pedro Gama , Jefersson A. dos Santos

The goal of acoustic (or sound) events detection (AED or SED) is to predict the temporal position of target events in given audio segments. This task plays a significant role in safety monitoring, acoustic early warning and other scenarios.…

音频与语音处理 · 电气工程与系统科学 2019-11-26 Wenhao Ding , Liang He

Emotion recognition in conversations is essential for ensuring advanced human-machine interactions. However, creating robust and accurate emotion recognition systems in real life is challenging, mainly due to the scarcity of emotion…

计算与语言 · 计算机科学 2023-08-30 Théo Deschamps-Berger , Lori Lamel , Laurence Devillers

Acoustic Scene Classification (ASC) is a challenging task, as a single scene may involve multiple events that contain complex sound patterns. For example, a cooking scene may contain several sound sources including silverware clinking,…

音频与语音处理 · 电气工程与系统科学 2019-09-20 Weimin Wang , Weiran Wang , Ming Sun , Chao Wang

Visual events are usually accompanied by sounds in our daily lives. We pose the question: Can the machine learn the correspondence between visual scene and the sound, and localize the sound source only by observing sound and visual scene…

计算机视觉与模式识别 · 计算机科学 2019-02-18 Arda Senocak , Tae-Hyun Oh , Junsik Kim , Ming-Hsuan Yang , In So Kweon

Action recognition and anticipation are key to the success of many computer vision applications. Existing methods can roughly be grouped into those that extract global, context-aware representations of the entire image or sequence, and…

计算机视觉与模式识别 · 计算机科学 2016-11-21 Mohammad Sadegh Aliakbarian , Fatemehsadat Saleh , Basura Fernando , Mathieu Salzmann , Lars Petersson , Lars Andersson