中文
相关论文

相关论文: AnyCrowd: Instance-Isolated Identity-Pose Binding …

200 篇论文

Identity-consistent generation has become an important focus in text-to-image research, with recent models achieving notable success in producing images aligned with a reference identity. Yet, the scarcity of large-scale paired datasets…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Hengyuan Xu , Wei Cheng , Peng Xing , Yixiao Fang , Shuhan Wu , Rui Wang , Xianfang Zeng , Daxin Jiang , Gang Yu , Xingjun Ma , Yu-Gang Jiang

Sign Language Recognition (SLR) has garnered significant attention from researchers in recent years, particularly the intricate domain of Continuous Sign Language Recognition (CSLR), which presents heightened complexity compared to Isolated…

计算机视觉与模式识别 · 计算机科学 2024-02-23 Razieh Rastgoo , Kourosh Kiani , Sergio Escalera

We present Lynx, a high-fidelity model for personalized video synthesis from a single input image. Built on an open-source Diffusion Transformer (DiT) foundation model, Lynx introduces two lightweight adapters to ensure identity fidelity.…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Shen Sang , Tiancheng Zhi , Tianpei Gu , Jing Liu , Linjie Luo

Character image animation, which generates high-quality videos from a reference image and target pose sequence, has seen significant progress in recent years. However, most existing methods only apply to human figures, which usually do not…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Shuai Tan , Biao Gong , Zhuoxin Liu , Yan Wang , Xi Chen , Yifan Feng , Hengshuang Zhao

Multiple object tracking (MOT) is a crucial task in computer vision society. However, most tracking-by-detection MOT methods, with available detected bounding boxes, cannot effectively handle static, slow-moving and fast-moving camera…

计算机视觉与模式识别 · 计算机科学 2020-06-25 Jiarui Cai , Yizhou Wang , Haotian Zhang , Hung-Min Hsu , Chengqian Ma , Jenq-Neng Hwang

Image-conditioned generation methods, such as depth- and canny-conditioned approaches, have demonstrated remarkable abilities for precise image synthesis. However, existing models still struggle to accurately control the content of multiple…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Dewei Zhou , Mingwei Li , Zongxin Yang , Yi Yang

Portrait animation from a single source image and a driving video is a long-standing problem. Recent approaches tend to adopt diffusion-based image/video generation models for realistic and expressive animation. However, none of these…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Yuxiang Shi , Zhe Li , Yanwen Wang , Hao Zhu , Xun Cao , Ligang Liu

Recent advances in diffusion models have greatly improved pose-driven character animation. However, existing methods are limited to spatially aligned reference-pose pairs with matched skeletal structures. Handling reference-pose…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Shijun Shi , Jing Xu , Zhihang Li , Chunli Peng , Xiaoda Yang , Lijing Lu , Kai Hu , Jiangning Zhang

Facial expression recognition (FER) is a challenging problem because the expression component is always entangled with other irrelevant factors, such as identity and head pose. In this work, we propose an identity and pose disentangled…

计算机视觉与模式识别 · 计算机科学 2022-08-18 Jing Jiang , Weihong Deng

Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video editing (RV2AV) and…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Xu Guo , Fulong Ye , Qichao Sun , Liyang Chen , Bingchuan Li , Pengze Zhang , Jiawei Liu , Songtao Zhao , Qian He , Xiangwang Hou

Action recognition is a fundamental task in video understanding. Existing methods typically extract unified features to process all actions in one video, which makes it challenging to model the interactions between different objects in…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Tianci Wu , Guangming Zhu , Jiang Lu , Siyuan Wang , Ning Wang , Nuoye Xiong , Zhang Liang

Recent advances in image-to-video (I2V) generation have achieved remarkable progress in synthesizing high-quality, temporally coherent videos from static images. Among all the applications of I2V, human-centric video generation includes a…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Liao Shen , Wentao Jiang , Yiran Zhu , Jiahe Li , Tiezheng Ge , Zhiguo Cao , Bo Zheng

Recent Video Large Language Models (Video-LLMs) have demonstrated strong capability in video understanding, yet they still suffer from hallucinations. Existing mitigation methods typically rely on training, input modification, auxiliary…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Zijian Liu , Sihan Cao , Pengcheng Zheng , Kuien Liu , Caiyan Qin , Xiaolin Qin , Jiwei Wei , Chaoning Zhang

Several factors contribute to the appearance of an object in a visual scene, including pose, illumination, and deformation, among others. Each factor accounts for a source of variability in the data, while the multiplicative interactions of…

计算机视觉与模式识别 · 计算机科学 2019-02-26 Mengjiao Wang , Zhixin Shu , Shiyang Cheng , Yannis Panagakis , Dimitris Samaras , Stefanos Zafeiriou

We introduce CharacterGAN, a generative model that can be trained on only a few samples (8 - 15) of a given character. Our model generates novel poses based on keypoint locations, which can be modified in real time while providing…

计算机视觉与模式识别 · 计算机科学 2022-01-13 Tobias Hinz , Matthew Fisher , Oliver Wang , Eli Shechtman , Stefan Wermter

Objects in a scene are not always related. The execution efficiency of the one-stage scene graph generation approaches are quite high, which infer the effective relation between entity pairs using sparse proposal sets and a few queries.…

计算机视觉与模式识别 · 计算机科学 2022-12-20 Yuxiang Zhang , Zhenbo Liu , Shuai Wang

The field of controllable human-centric video generation has witnessed remarkable progress, particularly with the advent of diffusion models. However, achieving precise and localized control over human motion in videos, such as replacing or…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Xiang Wang , Shiwei Zhang , Haonan Qiu , Ruihang Chu , Zekun Li , Yingya Zhang , Changxin Gao , Yuehuan Wang , Chunhua Shen , Nong Sang

Video diffusion transformers (DiTs) suffer from prohibitive inference latency due to quadratic attention complexity. Existing sparse attention methods either overlook semantic similarity or fail to adapt to heterogeneous token distributions…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Haoyue Tan , Shengnan Wang , Yulin Qiao , Juncheng Zhang , Youhui Bai , Ping Gong , Zewen Jin , Cheng Li

The diversity, quantity, and quality of manipulation data are critical for training effective robot policies. However, due to hardware and physical setup constraints, collecting large-scale real-world manipulation data remains difficult to…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Boyang Wang , Haoran Zhang , Shujie Zhang , Jinkun Hao , Mingda Jia , Qi Lv , Yucheng Mao , Zhaoyang Lyu , Jia Zeng , Xudong Xu , Jiangmiao Pang

Autonomous driving relies on robust models trained on large-scale, high-quality multi-view driving videos. Although world models provide a cost-effective solution for generating realistic driving data, they often suffer from identity drift,…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Zhuoran Yang , Yanyong Zhang