English
Related papers

Related papers: DoReMi: First glance at a universal OMR dataset

200 papers

Labeled data is a fundamental component in training supervised deep learning models for computer vision tasks. However, the labeling process, especially for ordinal image classification where class boundaries are often ambiguous, is prone…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Alireza Sedighi Moghaddam , Mohammad Reza Mohammadi

Developing open-source foundation models is essential for advancing research in music audio understanding and ensuring access to powerful, multipurpose representations for music information retrieval. We present OMAR-RQ, a model trained…

Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short video creation, and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Zhaokai Wang , Chenxi Bao , Le Zhuo , Jingrui Han , Yang Yue , Yihong Tang , Victor Shea-Jay Huang , Yue Liao

The goal of score following is to track a musical performance, usually in the form of audio, in a corresponding score representation. Established methods mainly rely on computer-readable scores in the form of MIDI or MusicXML and achieve…

Machine Learning · Computer Science 2019-10-17 Florian Henkel , Rainer Kelz , Gerhard Widmer

Pretraining on large-scale datasets can boost the performance of object detectors while the annotated datasets for object detection are hard to scale up due to the high labor cost. What we possess are numerous isolated filed-specific…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Jing Hao , Song Chen , Xiaodi Wang , Shumin Han

This paper tackles two key challenges: detecting small, dense, and overlapping objects (a major hurdle in computer vision) and improving the quality of noisy images, especially those encountered in industrial environments. [1, 2]. Our focus…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Oussama Messai , Abbass Zein-Eddine , Abdelouahid Bentamou , Mickaël Picq , Nicolas Duquesne , Stéphane Puydarrieux , Yann Gavet

With the advancement of drone technology, the volume of video data increases rapidly, creating an urgent need for efficient semantic retrieval. We are the first to systematically propose and study the drone video-text retrieval (DVTR) task.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Jinghao Huang , Yaxiong Chen , Ganchao Liu

This paper introduces a newly collected and novel dataset (StereoMSI) for example-based single and colour-guided spectral image super-resolution. The dataset was first released and promoted during the PIRM2018 spectral image…

Computer Vision and Pattern Recognition · Computer Science 2019-05-02 Mehrdad Shoeiby , Antonio Robles-Kelly , Ran Wei , Radu Timofte

Recent multi-camera 3D object detectors usually leverage temporal information to construct multi-view stereo that alleviates the ill-posed depth estimation. However, they typically assume all the objects are static and directly aggregate…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Qing Lian , Tai Wang , Dahua Lin , Jiangmiao Pang

Language-based object detection is a promising direction towards building a natural interface to describe objects in images that goes far beyond plain category names. While recent methods show great progress in that direction, proper…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Samuel Schulter , Vijay Kumar B G , Yumin Suh , Konstantinos M. Dafnis , Zhixing Zhang , Shiyu Zhao , Dimitris Metaxas

We present the first version of DDMD (Digital Drug Music Detector), a binary classifier that distinguishes digital drug music from normal music. In the literature, digital drug music is primarily explored regarding its psychological,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-01 Mohamed Gharzouli

Open-set image recognition (OSR) aims to both classify known-class samples and identify unknown-class samples in the testing set, which supports robust classifiers in many realistic applications, such as autonomous driving, medical…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Jiayin Sun , Qiulei Dong

Sensing emerges as a critical challenge in 6G networks, which require simultaneous communication and target sensing capabilities. State-of-the-art super-resolution techniques for the direction of arrival (DoA) estimation encounter…

Signal Processing · Electrical Eng. & Systems 2025-06-17 Murat Babek Salman , Emil Björnson

In the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding. This is evident in the…

Emotional information is essential for enhancing human-computer interaction and deepening image understanding. However, while deep learning has advanced image recognition, the intuitive understanding and precise control of emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Junjie Xu , Xingjiao Wu , Tanren Yao , Zihao Zhang , Jiayang Bei , Wu Wen , Liang He

The astounding success made by artificial intelligence (AI) in healthcare and other fields proves that AI can achieve human-like performance. However, success always comes with challenges. Deep learning algorithms are data-dependent and…

Image and Video Processing · Electrical Eng. & Systems 2021-06-25 Johann Li , Guangming Zhu , Cong Hua , Mingtao Feng , BasheerBennamoun , Ping Li , Xiaoyuan Lu , Juan Song , Peiyi Shen , Xu Xu , Lin Mei , Liang Zhang , Syed Afaq Ali Shah , Mohammed Bennamoun

Healthcare data are generated in many different formats, which makes it difficult to integrate and reuse across institutions and studies. Standardisation is required to enable consistent large-scale analysis. The OMOP-CDM, developed by the…

Quantitative Methods · Quantitative Biology 2025-11-13 Jacob Desmond , Ryan Wartmann , Chng Wei Lau , Steven Thomas , Paul M. Middleton , Jeewani Anupama Ginige

Foundation models have shown remarkable capabilities in various domains, but their performance on complex, multimodal engineering problems remains largely unexplored. We introduce SoM-1K, the first large-scale multimodal benchmark dataset…

Computation and Language · Computer Science 2025-09-26 Qixin Wan , Zilong Wang , Jingwen Zhou , Wanting Wang , Ziheng Geng , Jiachen Liu , Ran Cao , Minghui Cheng , Lu Cheng

With the great achievement of artificial intelligence, vehicle technologies have advanced significantly from human centric driving towards fully automated driving. An intelligent vehicle should be able to understand the driver's perception…

Human-Computer Interaction · Computer Science 2019-03-12 Yang Zheng , Izzat H. Izzat , John H. L. Hansen

Multimodal remote sensing image (MRSI) matching is pivotal for cross-modal fusion, localization, and object detection, but it faces severe challenges due to geometric, radiometric, and viewpoint discrepancies across imaging modalities.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Peihao Wu , Yongxiang Yao , Wenfei Zhang , Dong Wei , Yi Wan , Yansheng Li , Yongjun Zhang
‹ Prev 1 8 9 10 Next ›