English
Related papers

Related papers: Contrastive Multi-Modal Hypergraph Reasoning for 3…

200 papers

Holistic 3D human-scene reconstruction is a crucial and emerging research area in robot perception. A key challenge in holistic 3D human-scene reconstruction is to generate a physically plausible 3D scene from a single monocular RGB image.…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Sandika Biswas , Kejie Li , Biplab Banerjee , Subhasis Chaudhuri , Hamid Rezatofighi

Multi-contrast MRI images provide complementary contrast information about the characteristics of anatomical structures and are commonly used in clinical practice. Recently, a multi-flip-angle (FA) and multi-echo GRE method (MULTIPLEX MRI)…

Image and Video Processing · Electrical Eng. & Systems 2021-05-19 Eric Z. Chen , Yongquan Ye , Xiao Chen , Jingyuan Lyu , Zhongqi Zhang , Yichen Hu , Terrence Chen , Jian Xu , Shanhui Sun

We introduce MetricHMSR, a novel framework for recovering metric human meshes and 3D scenes from a single monocular image. Existing methods struggle to recover metric scale due to monocular scale ambiguity and weak-perspective camera…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Chentao Song , He Zhang , Haolei Yuan , Haozhe Lin , Jianhua Tao , Hongwen Zhang , Tao Yu

Multi-modal semantic understanding requires integrating information from different modalities to extract users' real intention behind words. Most previous work applies a dual-encoder structure to separately encode image and text, but fails…

Computation and Language · Computer Science 2024-03-12 Ming Zhang , Ke Chang , Yunfang Wu

We consider the problem of obtaining dense 3D reconstructions of humans from single and partially occluded views. In such cases, the visual evidence is usually insufficient to identify a 3D reconstruction uniquely, so we aim at recovering…

Computer Vision and Pattern Recognition · Computer Science 2020-11-03 Benjamin Biggs , Sébastien Ehrhadt , Hanbyul Joo , Benjamin Graham , Andrea Vedaldi , David Novotny

Hypergraphs, describing networks where interactions take place among any number of units, are a natural tool to model many real-world social and biological systems. In this work we propose a principled framework to model the organization of…

Social and Information Networks · Computer Science 2023-10-25 Nicolò Ruggeri , Martina Contisciani , Federico Battiston , Caterina De Bacco

Recently multi-view crowd counting using deep neural networks has been proposed to enable counting in large and wide scenes using multiple cameras. The current methods project the camera-view features to the average-height plane of the 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Qi Zhang , Antoni B. Chan

The image-text retrieval task aims to retrieve relevant information from a given image or text. The main challenge is to unify multimodal representation and distinguish fine-grained differences across modalities, thereby finding similar…

Multimedia · Computer Science 2024-05-20 Ziyu Gong , Chengcheng Mai , Yihua Huang

3D Human Body Reconstruction from a monocular image is an important problem in computer vision with applications in virtual and augmented reality platforms, animation industry, en-commerce domain, etc. While several of the existing works…

Computer Vision and Pattern Recognition · Computer Science 2019-08-20 Abbhinav Venkat , Chaitanya Patel , Yudhik Agrawal , Avinash Sharma

Recently, leveraging different channels to model social semantic information and using self-supervised learning tasks to boost recommendation performance has been proven to be a very promising work. However, how to deeply dig out the…

Information Retrieval · Computer Science 2022-09-27 Yundong Sun , Dongjie Zhu , Haiwen Du , Zhaoshuo Tian

Learning representations of multimodal data that are both informative and robust to missing modalities at test time remains a challenging problem due to the inherent heterogeneity of data obtained from different channels. To address it, we…

Machine Learning · Computer Science 2022-11-21 Petra Poklukar , Miguel Vasco , Hang Yin , Francisco S. Melo , Ana Paiva , Danica Kragic

We present a method for inferring diverse 3D models of human-object interactions from images. Reasoning about how humans interact with objects in complex scenes from a single 2D image is a challenging task given ambiguities arising from the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Xi Wang , Gen Li , Yen-Ling Kuo , Muhammed Kocabas , Emre Aksan , Otmar Hilliges

Monocular 3D clothed human reconstruction aims to generate a complete and realistic textured 3D avatar from a single image. Existing methods are commonly trained under multi-view supervision with annotated geometric priors, and during…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Nanjie Yao , Gangjian Zhang , Wenhao Shen , Jian Shu , Yu Feng , Hao Wang

Recent advancements in Graph Contrastive Learning (GCL) have demonstrated remarkable effectiveness in improving graph representations. However, relying on predefined augmentations (e.g., node dropping, edge perturbation, attribute masking)…

Machine Learning · Computer Science 2025-02-27 Khaled Mohammed Saifuddin , Shihao Ji , Esra Akbas

Object-centric reconstruction seeks to recover the 3D structure of a scene through composition of independent objects. While this independence can simplify modeling, it discards strong signals that could improve reconstruction, notably…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Qirui Wu , Yawar Siddiqui , Duncan Frost , Samir Aroudj , Armen Avetisyan , Richard Newcombe , Angel X. Chang , Jakob Engel , Henry Howard-Jenkins

Crowd counting is a challenging yet critical task in computer vision with applications ranging from public safety to urban planning. Recent advances using Convolutional Neural Networks (CNNs) that estimate density maps have shown…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Abhinav Sagar

Image-Text Retrieval (ITR) is challenging in bridging visual and lingual modalities. Contrastive learning has been adopted by most prior arts. Except for limited amount of negative image-text pairs, the capability of constrastive learning…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Haoran Wang , Dongliang He , Wenhao Wu , Boyang Xia , Min Yang , Fu Li , Yunlong Yu , Zhong Ji , Errui Ding , Jingdong Wang

Multi-view representation learning is essential for many multi-view tasks, such as clustering and classification. However, there are two challenging problems plaguing the community: i)how to learn robust multi-view representation from mass…

Computer Vision and Pattern Recognition · Computer Science 2023-02-28 Guanzhou Ke , Yongqi Zhu , Yang Yu

Recent methods for dynamic human reconstruction have attained promising reconstruction results. Most of these methods rely only on RGB color supervision without considering explicit geometric constraints. This leads to existing human…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Junhui Yin , Wei Yin , Hao Chen , Xuqian Ren , Zhanyu Ma , Jun Guo , Yifan Liu

Magnetic resonance imaging (MRI) can present multi-contrast images of the same anatomical structures, enabling multi-contrast super-resolution (SR) techniques. Compared with SR reconstruction using a single-contrast, multi-contrast SR…

Image and Video Processing · Electrical Eng. & Systems 2022-03-29 Guangyuan Li , Jun Lv , Yapeng Tian , Qi Dou , Chengyan Wang , Chenliang Xu , Jing Qin
‹ Prev 1 3 4 5 6 7 10 Next ›