中文
相关论文

相关论文: BlabberSeg: Real-Time Embedded Open-Vocabulary Aer…

200 篇论文

Open-vocabulary semantic segmentation strives to distinguish pixels into different semantic groups from an open set of categories. Most existing methods explore utilizing pre-trained vision-language models, in which the key is to adopt the…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Bin Xie , Jiale Cao , Jin Xie , Fahad Shahbaz Khan , Yanwei Pang

In recent years, unmanned aerial vehicles (UAVs) have played an increasingly crucial role in supporting disaster emergency response efforts by analyzing aerial images. While current deep-learning models focus on improving accuracy, they…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Lemeng Zhao , Junjie Hu , Jianchao Bi , Yanbing Bai , Erick Mas , Shunichi Koshimura

Instance segmentation of planar regions in indoor scenes benefits visual SLAM and other applications such as augmented reality (AR) where scene understanding is required. Existing methods built upon two-stage frameworks show satisfactory…

计算机视觉与模式识别 · 计算机科学 2021-03-30 Yaxu Xie , Jason Rambach , Fangwen Shu , Didier Stricker

Open-vocabulary semantic segmentation attempts to classify and outline objects in an image using arbitrary text labels, including those unseen during training. Self-supervised learning resolves numerous visual and linguistic processing…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Muhammad Atta ur Rahman , Dooseop Choi , Seung-Ik Lee , KyoungWook Min

The development of online high-definition maps is significant since they provide real-time, accurate, and updatable geographic information for location-based applications, such as autonomous driving and intelligent transportation, thus…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Mingchao Jiang , Yin Cheng , Linghai Liu

Vision-language models (VLMs) excel in visual understanding but often lack reliable grounding capabilities and actionable inference rates. Integrating them with open-vocabulary object detection (OVD), instance segmentation, and tracking…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Bastian Pätzold , Jan Nogga , Sven Behnke

Accurate ultrasound image segmentation is a prerequisite for precise biometrics and accurate assessment. Relying on manual delineation introduces significant errors and is time-consuming. However, existing segmentation models are designed…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Jie He , Minglang Chen , Minying Lu , Bocheng Liang , Junming Wei , Guiyan Peng , Jiaxi Chen , Ying Tan

Individuals with hearing impairments face challenges in their ability to comprehend speech, particularly in noisy environments. The aim of this study is to explore the effectiveness of audio-visual speech enhancement (AVSE) in enhancing the…

音频与语音处理 · 电气工程与系统科学 2025-10-07 Richard Lee Lai , Jen-Cheng Hou , I-Chun Chern , Kuo-Hsuan Hung , Yi-Ting Chen , Mandar Gogate , Tughrul Arslan , Amir Hussain , Yu Tsao

Despite recent progress in large-scale sound event detection (SED) systems capable of handling hundreds of sound classes, existing multi-class classification frameworks remain fundamentally limited. They cannot process free-text sound…

音频与语音处理 · 电气工程与系统科学 2025-09-24 Jiarui Hai , Helin Wang , Weizhe Guo , Mounya Elhilali

Segmentation of organs or lesions from medical images plays an essential role in many clinical applications such as diagnosis and treatment planning. Though Convolutional Neural Networks (CNN) have achieved the state-of-the-art performance…

计算机视觉与模式识别 · 计算机科学 2021-05-27 Xiangde Luo , Guotai Wang , Tao Song , Jingyang Zhang , Michael Aertsen , Jan Deprest , Sebastien Ourselin , Tom Vercauteren , Shaoting Zhang

As a novel and challenging task, referring segmentation combines computer vision and natural language processing to localize and segment objects based on textual descriptions. While referring image segmentation (RIS) has been extensively…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Rui Li , Xiaowei Zhao

Recently, Vision-Language Models (VLMs) have advanced segmentation techniques by shifting from the traditional segmentation of a closed-set of predefined object classes to open-vocabulary segmentation (OVS), allowing users to segment novel…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Gonca Yilmaz , Songyou Peng , Marc Pollefeys , Francis Engelmann , Hermann Blum

Audio-visual speech recognition (AVSR) is a multimodal extension of automatic speech recognition (ASR), using video as a complement to audio. In AVSR, considerable efforts have been directed at datasets for facial features such as…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Hao Wang , Shuhei Kurita , Shuichiro Shimizu , Daisuke Kawahara

This paper introduces BEV-VLM, a novel approach for trajectory planning in autonomous driving that leverages Vision-Language Models (VLMs) with Bird's-Eye View (BEV) feature maps as visual input. Unlike conventional trajectory planning…

机器人学 · 计算机科学 2026-03-02 Guancheng Chen , Sheng Yang , Tong Zhan , Jian Wang

Recently, groundbreaking results have been presented on open-vocabulary semantic image segmentation. Such methods segment each pixel in an image into arbitrary categories provided at run-time in the form of text prompts, as opposed to a…

机器人学 · 计算机科学 2023-03-21 Kenneth Blomqvist , Francesco Milano , Jen Jen Chung , Lionel Ott , Roland Siegwart

The ability to perform semantic segmentation in real-time capable applications with limited hardware is of great importance. One such application is the interpretation of the visual bird's-eye view, which requires the semantic segmentation…

计算机视觉与模式识别 · 计算机科学 2018-11-30 Timo Sämann , Karl Amende , Stefan Milz , Christian Witt , Martin Simon , Johannes Petzold

The field-of-view is an important metric when designing a model for semantic segmentation. To obtain a large field-of-view, previous approaches generally choose to rapidly downsample the resolution, usually with average poolings or stride 2…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Roland Gao

Real-time audio-visual speech enhancement (AVSE) is a key enabler for immersive and interactive multimedia services, yet its performance is tightly constrained by network latency, uplink capacity, and computational delay. This paper…

声音 · 计算机科学 2026-04-29 Anis Hamadouche , Haifeng Luo , Mathini Sellathurai , Amir Hussain , Tharm Ratnarajah

Accurate far-field speech datasets are critical for tasks such as automatic speech recognition (ASR), dereverberation, speech enhancement, and source separation. However, current datasets are limited by the trade-off between acoustic…

音频与语音处理 · 电气工程与系统科学 2025-10-28 Sarabeth S. Mullins , Georg Götz , Eric Bezzam , Steven Zheng , Daniel Gert Nielsen

Transformer-based Large Language Models (LLMs) have made a significant impact on various domains. However, LLMs' efficiency suffers from both heavy computation and memory overheads. Compression techniques like sparsification and…

‹ 上一页 1 8 9 10 下一页 ›