English
Related papers

Related papers: GIRAFE: Glottal Imaging Dataset for Advanced Segme…

200 papers

Group Relative Policy Optimization (GRPO) methods for video generation like FlowGRPO remain far less reliable than their counterparts for language models and images. This gap arises because video generation has a complex solution space, and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Mingzhe Zheng , Weijie Kong , Yue Wu , Dengyang Jiang , Yue Ma , Xuanhua He , Bin Lin , Kaixiong Gong , Zhao Zhong , Liefeng Bo , Qifeng Chen , Harry Yang

Open-vocabulary segmentation aims to identify and segment specific regions and objects based on text-based descriptions. A common solution is to leverage powerful vision-language models (VLMs), such as CLIP, to bridge the gap between vision…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Bingyu Li , Da Zhang , Zhiyuan Zhao , Junyu Gao , Xuelong Li

Traditional video-induced physiological datasets usually rely on whole-trial labels, which introduce temporal label noise in dynamic emotion recognition. We present FIRMED, a peak-centered multimodal dataset based on an immediate-recall…

Human-Computer Interaction · Computer Science 2026-04-01 Hao Tang , Songyun Xie , Xinzhou Xie , Can Liao , Bohan Li , Zhongyu Tian , Dalu Zheng

We introduce the Expanded Groove MIDI dataset (E-GMD), an automatic drum transcription (ADT) dataset that contains 444 hours of audio from 43 drum kits, making it an order of magnitude larger than similar datasets, and the first with…

Sound · Computer Science 2020-12-02 Lee Callender , Curtis Hawthorne , Jesse Engel

While deep learning technologies are now capable of generating realistic images confusing humans, the research efforts are turning to the synthesis of images for more concrete and application-specific purposes. Facial image generation based…

Computer Vision and Pattern Recognition · Computer Science 2020-06-11 Yeqi Bai , Tao Ma , Lipo Wang , Zhenjie Zhang

The growing demand for prenatal ultrasound imaging has intensified a global shortage of trained sonographers, creating barriers to essential fetal health monitoring. Deep learning has the potential to enhance sonographers' efficiency and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Hussain Alasmawi , Numan Saeed , Mohammad Yaqub

Automatic Singing Assessment and Singing Information Processing have evolved over the past three decades to support singing pedagogy, performance analysis, and vocal training. While the first approach objectively evaluates a singer's…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Arthur N. dos Santos , Bruno S. Masiero

The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality labels, previous methods often suffer from limited realism…

Computer Vision and Pattern Recognition · Computer Science 2024-01-31 Qingcheng Zhao , Pengyu Long , Qixuan Zhang , Dafei Qin , Han Liang , Longwen Zhang , Yingliang Zhang , Jingyi Yu , Lan Xu

Surgical tool segmentation in endoscopic videos is an important component of computer assisted interventions systems. Recent success of image-based solutions using fully-supervised deep learning approaches can be attributed to the…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Manish Sahu , Ronja Strömsdörfer , Anirban Mukhopadhyay , Stefan Zachow

3D scene segmentation based on neural implicit representation has emerged recently with the advantage of training only on 2D supervision. However, existing approaches still requires expensive per-scene optimization that prohibits…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Hanlin Chen , Chen Li , Mengqi Guo , Zhiwen Yan , Gim Hee Lee

Head-related transfer functions (HRTFs) with dense spatial grids are desired for immersive binaural audio generation, but their recording is time-consuming. Although HRTF spatial upsampling has shown remarkable progress with neural fields,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-23 Yoshiki Masuyama , Gordon Wichern , François G. Germain , Christopher Ick , Jonathan Le Roux

Whole abdominal organ segmentation is important in diagnosing abdomen lesions, radiotherapy, and follow-up. However, oncologists' delineating all abdominal organs from 3D volumes is time-consuming and very expensive. Deep learning-based…

Image and Video Processing · Electrical Eng. & Systems 2023-02-14 Xiangde Luo , Wenjun Liao , Jianghong Xiao , Jieneng Chen , Tao Song , Xiaofan Zhang , Kang Li , Dimitris N. Metaxas , Guotai Wang , Shaoting Zhang

Integrating real-time artificial intelligence (AI) systems in clinical practices faces challenges such as scalability and acceptance. These challenges include data availability, biased outcomes, data quality, lack of transparency, and…

Image and Video Processing · Electrical Eng. & Systems 2023-08-21 Debesh Jha , Vanshali Sharma , Neethi Dasu , Nikhil Kumar Tomar , Steven Hicks , M. K. Bhuyan , Pradip K. Das , Michael A. Riegler , Pål Halvorsen , Ulas Bagci , Thomas de Lange

Active Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Yu Wang , Juhyung Ha , Frangil M. Ramirez , Yuchen Wang , David J. Crandall

Surgical instrument segmentation (SIS) on endoscopic images stands as a long-standing and essential task in the context of computer-assisted interventions for boosting minimally invasive surgery. Given the recent surge of deep learning…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Mingyu Sheng , Jianan Fan , Dongnan Liu , Ron Kikinis , Weidong Cai

Recently, deep learning enabled the accurate segmentation of various diseases in medical imaging. These performances, however, typically demand large amounts of manual voxel annotations. This tedious process for volumetric data becomes more…

Diffusion-based generative models have achieved state-of-the-art performance for perceptual quality in speech enhancement (SE). However, their iterative nature requires numerous Neural Function Evaluations (NFEs), posing a challenge for…

Deployment of deep learning models in robotics as sensory information extractors can be a daunting task to handle, even using generic GPU cards. Here, we address three of its most prominent hurdles, namely, i) the adaptation of a single…

Computer Vision and Pattern Recognition · Computer Science 2019-02-28 Vladimir Nekrasov , Thanuja Dharmasiri , Andrew Spek , Tom Drummond , Chunhua Shen , Ian Reid

We propose an Explicit Conditional Multimodal Variational Auto-Encoder (ECMVAE) for audio-visual segmentation (AVS), aiming to segment sound sources in the video sequence. Existing AVS methods focus on implicit feature fusion strategies,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Yuxin Mao , Jing Zhang , Mochu Xiang , Yiran Zhong , Yuchao Dai

Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and generation are often treated as distinct tasks, hindering the…

‹ Prev 1 4 5 6 7 8 10 Next ›