中文
相关论文

相关论文: Contrastive Multi-Modal Hypergraph Reasoning for 3…

200 篇论文

Due to the mutual occlusion, severe scale variation, and complex spatial distribution, the current multi-person mesh recovery methods cannot produce accurate absolute body poses and shapes in large-scale crowded scenes. To address the…

计算机视觉与模式识别 · 计算机科学 2023-08-31 Buzhen Huang , Jingyi Ju , Zhihao Li , Yangang Wang

This is a technical report for the GigaCrowd challenge. Reconstructing 3D crowds from monocular images is a challenging problem due to mutual occlusions, server depth ambiguity, and complex spatial distribution. Since no large-scale 3D…

计算机视觉与模式识别 · 计算机科学 2023-10-11 Buzhen Huang , Jingyi Ju , Yangang Wang

In this study, we focus on the problem of 3D human mesh recovery from a single image under obscured conditions. Most state-of-the-art methods aim to improve 2D alignment technologies, such as spatial averaging and 2D joint sampling.…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Jiahao Li , Zongxin Yang , Xiaohan Wang , Jianxin Ma , Chang Zhou , Yi Yang

Reconstructing multi-human body mesh from a single monocular image is an important but challenging computer vision problem. In addition to the individual body mesh models, we need to estimate relative 3D positions among subjects to generate…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Chenyan Wu , Yandong Li , Xianfeng Tang , James Wang

Single-view 3D human reconstruction has garnered significant attention in recent years. Despite numerous advancements, prior research has concentrated on reconstructing 3D models from clear, close-up images of individual subjects, often…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Yizheng Song , Yiyu Zhuang , Qipeng Xu , Haixiang Wang , Jiahe Zhu , Jing Tian , Siyu Zhu , Hao Zhu

Multi-modal recommender system focuses on utilizing rich modal information ( i.e., images and textual descriptions) of items to improve recommendation performance. The current methods have achieved remarkable success with the powerful…

信息检索 · 计算机科学 2025-08-20 Shouxing Ma , Yawen Zeng , Shiqing Wu , Guandong Xu

The burgeoning presence of multimodal content-sharing platforms propels the development of personalized recommender systems. Previous works usually suffer from data sparsity and cold-start problems, and may fail to adequately explore…

信息检索 · 计算机科学 2025-04-24 Xu Guo , Tong Zhang , Fuyun Wang , Xudong Wang , Xiaoya Zhang , Xin Liu , Zhen Cui

3D reconstruction of dynamic crowds in large scenes has become increasingly important for applications such as city surveillance and crowd analysis. However, current works attempt to reconstruct 3D crowds from a static image, causing a lack…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Hao Wen , Hongbo Kang , Jian Ma , Jing Huang , Yuanwang Yang , Haozhe Lin , Yu-Kun Lai , Kun Li

Reconstructing textured 3D human models from a single image is fundamental for AR/VR and digital human applications. However, existing methods mostly focus on single individuals and thus fail in multi-human scenes, where naive composition…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Gwanghyun Kim , Junghun James Kim , Suh Yoon Jeon , Jason Park , Se Young Chun

Multimodal emotion recognition in conversation (MERC) seeks to identify the speakers' emotions expressed in each utterance, offering significant potential across diverse fields. The challenge of MERC lies in balancing speaker modeling and…

多媒体 · 计算机科学 2025-07-25 Zijian Yi , Ziming Zhao , Zhishu Shen , Tiehua Zhang

We consider the problem of recovering a single person's 3D human mesh from in-the-wild crowded scenes. While much progress has been in 3D human mesh estimation, existing methods struggle when test input has crowded scenes. The first reason…

计算机视觉与模式识别 · 计算机科学 2022-09-20 Hongsuk Choi , Gyeongsik Moon , JoonKyu Park , Kyoung Mu Lee

Reconstructing clothed humans from a single image is a fundamental task in computer vision with wide-ranging applications. Although existing monocular clothed human reconstruction solutions have shown promising results, they often rely on…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Arindam Dutta , Meng Zheng , Zhongpai Gao , Benjamin Planche , Anwesha Choudhuri , Terrence Chen , Amit K. Roy-Chowdhury , Ziyan Wu

Motion reasoning serves as the cornerstone of multi-object tracking (MOT), as it enables consistent association of targets across frames. However, existing motion estimation approaches face two major limitations: (1) instability caused by…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Zikai Song , Junqing Yu , Yi-Ping Phoebe Chen , Wei Yang , Xinchao Wang

Epipolar constraints are at the core of feature matching and depth estimation in current multi-person multi-camera 3D human pose estimation methods. Despite the satisfactory performance of this formulation in sparser crowd scenes, its…

计算机视觉与模式识别 · 计算机科学 2020-07-22 He Chen , Pengfei Guo , Pengfei Li , Gim Hee Lee , Gregory Chirikjian

Image-based multi-person reconstruction in wide-field large scenes is critical for crowd analysis and security alert. However, existing methods cannot deal with large scenes containing hundreds of people, which encounter the challenges of…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Hao Wen , Jing Huang , Huili Cui , Haozhe Lin , YuKun Lai , Lu Fang , Kun Li

In this work, we address the problem of multi-person 3D pose estimation from a single image. A typical regression approach in the top-down setting of this problem would first detect all humans and then reconstruct each one of them…

计算机视觉与模式识别 · 计算机科学 2020-06-16 Wen Jiang , Nikos Kolotouros , Georgios Pavlakos , Xiaowei Zhou , Kostas Daniilidis

Perspective distortions and crowd variations make crowd counting a challenging task in computer vision. To tackle it, many previous works have used multi-scale architecture in deep neural networks (DNNs). Multi-scale branches can be either…

计算机视觉与模式识别 · 计算机科学 2023-07-19 Zhipeng Du , Miaojing Shi , Jiankang Deng , Stefanos Zafeiriou

Cross-modal 3D retrieval is a critical yet challenging task, aiming to achieve bi-directional retrieval between 3D and text modalities. Current methods predominantly rely on a certain 3D representation (e.g., point cloud), with few…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Junlong Ren , Hao Wang

Multi-person 3D pose estimation is a challenging task because of occlusion and depth ambiguity, especially in the cases of crowd scenes. To solve these problems, most existing methods explore modeling body context cues by enhancing feature…

计算机视觉与模式识别 · 计算机科学 2022-08-09 Zhongwei Qiu , Qiansheng Yang , Jian Wang , Dongmei Fu

Crowd understanding has aroused the widespread interest in vision domain due to its important practical significance. Unfortunately, there is no effort to explore crowd understanding in multi-modal domain that bridges natural language and…

计算机视觉与模式识别 · 计算机科学 2022-06-17 Heqian Qiu , Hongliang Li , Taijin Zhao , Lanxiao Wang , Qingbo Wu , Fanman Meng
‹ 上一页 1 2 3 10 下一页 ›