English
Related papers

Related papers: Multivariate Gaussian Representation Learning for …

200 papers

This paper contributes a novel realtime multi-person motion capture algorithm using multiview video inputs. Due to the heavy occlusions in each view, joint optimization on the multiview images and multiple temporal frames is indispensable,…

Computer Vision and Pattern Recognition · Computer Science 2020-03-02 Yuxiang Zhang , Liang An , Tao Yu , Xiu Li , Kun Li , Yebin Liu

Recently, learned video compression has achieved exciting performance. Following the traditional hybrid prediction coding framework, most learned methods generally adopt the motion estimation motion compensation (MEMC) method to remove…

Image and Video Processing · Electrical Eng. & Systems 2023-10-20 Yiming Wang , Qian Huang , Bin Tang , Huashan Sun , Xing Li

Learning representations of multimodal data that are both informative and robust to missing modalities at test time remains a challenging problem due to the inherent heterogeneity of data obtained from different channels. To address it, we…

Machine Learning · Computer Science 2022-11-21 Petra Poklukar , Miguel Vasco , Hang Yin , Francisco S. Melo , Ana Paiva , Danica Kragic

Document Visual Question Answering (DocVQA) faces dual challenges in processing lengthy multimodal documents (text, images, tables) and performing cross-modal reasoning. Current document retrieval-augmented generation (DocRAG) methods…

Information Retrieval · Computer Science 2025-11-10 Kuicai Dong , Yujing Chang , Shijie Huang , Yasheng Wang , Ruiming Tang , Yong Liu

Gait recognition is emerging as a promising technology and an innovative field within computer vision, with a wide range of applications in remote human identification. However, existing methods typically rely on complex architectures to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Zhengxian Wu , Chuanrui Zhang , Shenao Jiang , Hangrui Xu , Zirui Liao , Luyuan Zhang , Huaqiu Li , Peng Jiao , Haoqian Wang

Acquiring properly annotated data is expensive in the medical field as it requires experts, time-consuming protocols, and rigorous validation. Active learning attempts to minimize the need for large annotated samples by actively sampling…

Image and Video Processing · Electrical Eng. & Systems 2023-06-22 Bidur Khanal , Binod Bhattarai , Bishesh Khanal , Danail Stoyanov , Cristian A. Linte

Multi-label image classification demands adaptive training strategies to navigate complex, evolving visual-semantic landscapes, yet conventional methods rely on static configurations that falter in dynamic settings. We propose MAT-Agent, a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Jusheng Zhang , Kaitong Cai , Yijia Fan , Ningyuan Liu , Keze Wang

Medical Visual Question Answering (MedVQA) is a promising field for developing clinical decision support systems, yet progress is often limited by the available datasets, which can lack clinical complexity and visual diversity. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Sushant Gautam , Michael A. Riegler , Pål Halvorsen

Self-supervised, multi-modal learning has been successful in holistic representation of complex scenarios. This can be useful to consolidate information from multiple modalities which have multiple, versatile uses. Its application in…

Computer Vision and Pattern Recognition · Computer Science 2020-11-03 Aniruddha Tamhane , Jie Ying Wu , Mathias Unberath

We propose a novel multi-modal and multi-task architecture for simultaneous low level gesture and surgical task classification in Robot Assisted Surgery (RAS) videos.Our end-to-end architecture is based on the principles of a long…

Computer Vision and Pattern Recognition · Computer Science 2018-05-03 Duygu Sarikaya , Khurshid A. Guru , Jason J. Corso

Fine-grained video classification requires understanding complex spatio-temporal and semantic cues that often exceed the capacity of a single modality. In this paper, we propose a multimodal framework that fuses video, image, and text…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Namho Kim , Junhwa Kim

Learning efficient representations for concepts has been proven to be an important basis for many applications such as machine translation or document classification. Proper representations of medical concepts such as diagnosis, medication,…

Machine Learning · Computer Science 2016-02-18 Edward Choi , Mohammad Taha Bahadori , Elizabeth Searles , Catherine Coffey , Jimeng Sun

Obtaining large-scale human-labeled datasets to train acoustic representation models is a very challenging task. On the contrary, we can easily collect data with machine-generated labels. In this work, we propose to exploit…

Computer Vision and Pattern Recognition · Computer Science 2020-01-03 Shaoyong Jia , Xin Shu , Yang Yang , Dawei Liang , Qiyue Liu , Junhui Liu

Recent years have witnessed remarkable progress in multimodal learning within computational pathology. Existing models primarily rely on vision and language modalities; however, language alone lacks molecular specificity and offers limited…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Minghao Han , Dingkang Yang , Linhao Qu , Zizhi Chen , Gang Li , Han Wang , Jiacong Wang , Lihua Zhang

Video-based gait analysis has become a promising approach for assessing motor impairment in children with cerebral palsy (CP). However, existing methods usually rely on either pose sequences or handcrafted gait features alone, making it…

Image and Video Processing · Electrical Eng. & Systems 2026-03-25 Kaiyuan Yang , Xupeng Chen , Jiangpeng He

Training multimodal models requires a large amount of labeled data. Active learning (AL) aim to reduce labeling costs. Most AL methods employ warm-start approaches, which rely on sufficient labeled data to train a well-calibrated model that…

Multimedia · Computer Science 2024-12-13 Meng Shen , Yake Wei , Jianxiong Yin , Deepu Rajan , Di Hu , Simon See

Surgical workflow anticipation can give predictions on what steps to conduct or what instruments to use next, which is an essential part of the computer-assisted intervention system for surgery, e.g. workflow reasoning in robotic surgery.…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Xiatian Zhang , Noura Al Moubayed , Hubert P. H. Shum

Estimating physical properties for visual data is a crucial task in computer vision, graphics, and robotics, underpinning applications such as augmented reality, physical simulation, and robotic grasping. However, this area remains…

Constructing vivid 3D head avatars for given subjects and realizing a series of animations on them is valuable yet challenging. This paper presents GaussianHead, which models the actional human head with anisotropic 3D Gaussians. In our…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Jie Wang , Jiu-Cheng Xie , Xianyan Li , Feng Xu , Chi-Man Pun , Hao Gao

In medical visual question answering (Med-VQA), achieving accurate responses relies on three critical steps: precise perception of medical imaging data, logical reasoning grounded in visual input and textual questions, and coherent answer…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Songtao Jiang , Yuan Wang , Ruizhe Chen , Yan Zhang , Ruilin Luo , Bohan Lei , Sibo Song , Yang Feng , Jimeng Sun , Jian Wu , Zuozhu Liu