中文
相关论文

相关论文: You Only Hear Once: A YOLO-like Algorithm for Audi…

200 篇论文

Automatic modulation recognition (AMR) is a crucial step in wireless communication systems, which identifies the modulation scheme from detected signals to provide key information for further processing. However, previous work has mainly…

信号处理 · 电气工程与系统科学 2025-12-01 Yunpeng Qu , Yazhou Sun , Bingyu Hui , Jian Wang

Efficient and accurate annotation of datasets remains a significant challenge for deploying object detection models such as You Only Look Once (YOLO) in real-world applications, particularly in agriculture where rapid decision-making is…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Mohamed Abdallah Salem , Ahmed Harb Rabia

Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes…

We propose UOLO, a novel framework for the simultaneous detection and segmentation of structures of interest in medical images. UOLO consists of an object segmentation module which intermediate abstract representations are processed and…

计算机视觉与模式识别 · 计算机科学 2019-06-18 Teresa Araújo , Guilherme Aresta , Adrian Galdran , Pedro Costa , Ana Maria Mendonça , Aurélio Campilho

Laughter is a social non-vocalization that is universal across cultures and languages, and is crucial for human communication, including social bonding and communication signaling. However, detecting laughter in audio is a challenging task,…

计算与语言 · 计算机科学 2026-05-14 Sofia Callejas , Nahuel Gomez , Catherine Pelachaud , Brian Ravenet , Valentin Barriere

Spatiotemporal action recognition is the task of locating and classifying actions in videos. Our project applies this task to analyzing video footage of restaurant workers preparing food, for which potential applications include automated…

计算机视觉与模式识别 · 计算机科学 2020-08-26 Akshat Gupta , Milan Desai , Wusheng Liang , Magesh Kannan

Efficient audio quality assessment is vital for streamlining audio codec development. Objective assessment tools have been developed over time to algorithmically predict quality ratings from subjective assessments, the gold standard for…

音频与语音处理 · 电气工程与系统科学 2024-11-28 Pablo M. Delgado , Jürgen Herre

Machine hearing or listening represents an emerging area. Conventional approaches rely on the design of handcrafted features specialized to a specific audio task and that can hardly generalized to other audio fields. For example,…

计算机视觉与模式识别 · 计算机科学 2018-12-13 Imad Rida , Romain Hérault , Gilles Gasso

Outdoor LiDAR point cloud 3D instance segmentation is a crucial task in autonomous driving. However, it requires laborious human efforts to annotate the point cloud for training a segmentation model. To address this challenge, we propose a…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Guangfeng Jiang , Jun Liu , Yongxuan Lv , Yuzhi Wu , Xianfei Li , Wenlong Liao , Tao He , Pai Peng

Object detection and segmentation are widely employed in computer vision applications, yet conventional models like YOLO series, while efficient and accurate, are limited by predefined categories, hindering adaptability in open scenarios.…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Ao Wang , Lihao Liu , Hui Chen , Zijia Lin , Jungong Han , Guiguang Ding

Multimodal fusion is a multimedia technique that has become popular in the wide range of tasks where image information is accompanied by a signal/audio. The latter may not convey highly semantic information, such as speech or music, but…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Alexey Zhukov , Jenny Benois-Pineau , Amira Youssef , Akka Zemmari , Mohamed Mosbah , Virginie Taillandier

This research work dives into an in-depth evaluation of the YOLOv8 (You Only Look Once) algorithm's efficiency in object detection, specially focusing on Barcode and QR code recognition. Utilizing the real-time detection abilities of…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Kushagra Pandya , Heli Hathi , Het Buch , Ravikumar R N , Shailendrasinh Chauhan , Sushil Kumar Singh

Visual object detection utilizing deep learning plays a vital role in computer vision and has extensive applications in transportation engineering. This paper focuses on detecting pavement marking quality during daytime using the You Only…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Gian Antariksa , Rohit Chakraborty , Shriyank Somvanshi , Subasish Das , Mohammad Jalayer , Deep Rameshkumar Patel , David Mills

For years, the YOLO series has been the de facto industry-level standard for efficient object detection. The YOLO community has prospered overwhelmingly to enrich its use in a multitude of hardware platforms and abundant scenarios. In this…

Advances in deep learning have resulted in state-of-the-art performance for many audio classification tasks but, unlike humans, these systems traditionally require large amounts of data to make accurate predictions. Not every person or…

音频与语音处理 · 电气工程与系统科学 2020-12-04 Piper Wolters , Chris Careaga , Brian Hutchinson , Lauren Phillips

Multispectral imaging and deep learning have emerged as powerful tools supporting diverse use cases from autonomous vehicles, to agriculture, infrastructure monitoring and environmental assessment. The combination of these technologies has…

计算机视觉与模式识别 · 计算机科学 2024-09-23 James E. Gallagher , Edward J. Oughton

Overtaking is a critical maneuver in driving that requires accurate information about the location and distance of other vehicles on the road. This study suggests a real-time overtaking assistance system that uses a combination of the You…

计算机视觉与模式识别 · 计算机科学 2023-05-01 Chanthini Bhaskar , Bharath Manoj Nair , Dev Mehta

In the realm of Tiny AI, we introduce ``You Only Look at Interested Cells" (YOLIC), an efficient method for object localization and classification on edge devices. Through seamlessly blending the strengths of semantic segmentation and…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Kai Su , Yoichi Tomioka , Qiangfu Zhao , Yong Liu

YOLO is a deep neural network (DNN) model presented for robust real-time object detection following the one-stage inference approach. It outperforms other real-time object detectors in terms of speed and accuracy by a wide margin.…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Mohammadamin Baghbanbashi , Mohsen Raji , Behnam Ghavami

Lifelong audio feature extraction involves learning new sound classes incrementally, which is essential for adapting to new data distributions over time. However, optimizing the model only on new data can lead to catastrophic forgetting of…

音频与语音处理 · 电气工程与系统科学 2024-02-08 Xilin Jiang , Yinghao Aaron Li , Nima Mesgarani