English
Related papers

Related papers: OmniHuman: A Large-scale Dataset and Benchmark for…

200 papers

Current datasets for long-form video understanding often fall short of providing genuine long-form comprehension challenges, as many tasks derived from these datasets can be successfully tackled by analyzing just one or a few random frames…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Ruchit Rawal , Khalid Saifullah , Miquel Farré , Ronen Basri , David Jacobs , Gowthami Somepalli , Tom Goldstein

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data,…

Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting flexibility for…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Wendong Bu , Kaihang Pan , Yuze Lin , Jiacheng Li , Kai Shen , Wenqiao Zhang , Juncheng Li , Jun Xiao , Siliang Tang

Text-to-image diffusion models have significantly advanced in conditional image generation. However, these models usually struggle with accurately rendering images featuring humans, resulting in distorted limbs and other anomalies. This…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Guian Fang , Wenbiao Yan , Yuanfan Guo , Jianhua Han , Zutao Jiang , Hang Xu , Shengcai Liao , Xiaodan Liang

Fine-grained perception of multimodal information is critical for advancing human-AI interaction. With recent progress in audio-visual technologies, Omni Language Models (OLMs), capable of processing audio and video signals in parallel,…

Computation and Language · Computer Science 2026-03-17 Ziyang Ma , Ruiyang Xu , Zhenghao Xing , Yunfei Chu , Yuxuan Wang , Jinzheng He , Jin Xu , Pheng-Ann Heng , Kai Yu , Junyang Lin , Eng Siong Chng , Xie Chen

Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generative fields, most…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Jing Lin , Ruisi Wang , Junzhe Lu , Ziqi Huang , Guorui Song , Ailing Zeng , Xian Liu , Chen Wei , Wanqi Yin , Qingping Sun , Zhongang Cai , Lei Yang , Ziwei Liu

Despite the promising progress in subject-driven image generation, current models often deviate from the reference identities and struggle in complex scenes with multiple subjects. To address this challenge, we introduce OpenSubject, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Yexin Liu , Manyuan Zhang , Yueze Wang , Hongyu Li , Dian Zheng , Weiming Zhang , Changsheng Lu , Xunliang Cai , Yan Feng , Peng Pei , Harry Yang

Video generation has emerged as a promising tool for world simulation, leveraging visual data to replicate real-world environments. Within this context, egocentric video generation, which centers on the human perspective, holds significant…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Xiaofeng Wang , Kang Zhao , Feng Liu , Jiayu Wang , Guosheng Zhao , Xiaoyi Bao , Zheng Zhu , Yingya Zhang , Xingang Wang

Story visualization aims to generate coherent image sequences that faithfully represent a narrative and match given character references. Despite progress in generative models, existing benchmarks remain narrow in scope, often limited to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Cailin Zhuang , Ailin Huang , Yaoqi Hu , Jingwei Wu , Wei Cheng , Jiaqi Liao , Hongyuan Wang , Xinyao Liao , Weiwei Cai , Hengyuan Xu , Xuanyang Zhang , Xianfang Zeng , Zhewei Huang , Gang Yu , Chi Zhang

Unified large multimodal models (LMMs) have achieved remarkable progress in general-purpose multimodal understanding and generation. However, they still operate under a ``one-size-fits-all'' paradigm and struggle to model user-specific…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Yu Zhong , Tianwei Lin , Ruike Zhu , Yuqian Yuan , Haoyu Zheng , Liang Liang , Wenqiao Zhang , Feifei Shao , Haoyuan Li , Wanggui He , Hao Jiang , Yueting Zhuang

Controllable human video generation aims to produce realistic videos of humans with explicitly guided motions and appearances,serving as a foundation for digital humans, animation, and embodied AI.However, the scarcity of largescale,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Yuanchen Fei , Yude Zou , Zejian Kang , Ming Li , Jiaying Zhou , Xiangru Huang

Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existing OLM benchmarks remain anchored to static, accuracy-centric tasks, leaving a critical gap in…

Artificial Intelligence · Computer Science 2026-03-18 Tianyu Xie , Jinfa Huang , Yuexiao Ma , Rongfang Luo , Yan Yang , Wang Chen , Yuhui Zeng , Ruize Fang , Yixuan Zou , Xiawu Zheng , Jiebo Luo , Rongrong Ji

The rapid advancement of GenAI technology over the past few years has significantly contributed towards highly realistic deepfake content generation. Despite ongoing efforts, the research community still lacks a large-scale and reasoning…

Multimedia · Computer Science 2025-06-17 Parul Gupta , Shreya Ghosh , Tom Gedeon , Thanh-Toan Do , Abhinav Dhall

Most existing video tasks related to "human" focus on the segmentation of salient humans, ignoring the unspecified others in the video. Few studies have focused on segmenting and tracking all humans in a complex video, including pedestrians…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Ran Yu , Chenyu Tian , Weihao Xia , Xinyuan Zhao , Haoqian Wang , Yujiu Yang

Whole-body humanoid teleoperation enables humans to remotely control humanoid robots, serving as both a real-time operational tool and a scalable engine for collecting demonstrations for autonomous learning. Despite recent advances,…

Robotics · Computer Science 2026-03-17 Yixuan Li , Le Ma , Yutang Lin , Yushi Du , Mengya Liu , Kaizhe Hu , Jieming Cui , Yixin Zhu , Wei Liang , Baoxiong Jia , Siyuan Huang

Video generation models are rapidly improving in their ability to synthesize human actions in novel contexts, holding the potential to serve as high-level planners for contextual robot control. To realize this potential, a key research…

Robotics · Computer Science 2025-12-12 James Ni , Zekai Wang , Wei Lin , Amir Bar , Yann LeCun , Trevor Darrell , Jitendra Malik , Roei Herzig

Improving visual text synthesis has long been a challenging and evolving frontier for image generation models. While recent state-of-the-art (SOTA) models have made remarkable strides in text generation capabilities, existing benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Peirong Zhang , Haowei Xu , Jiaxin Zhang , Xuhan Zheng , Guitao Xu , Yuyi Zhang , Junle Liu , Zhenhua Yang , Wei Zhou , Lianwen Jin

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for video applications sets higher requirements for high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Zhucun Xue , Jiangning Zhang , Teng Hu , Haoyang He , Yinan Chen , Yuxuan Cai , Yabiao Wang , Chengjie Wang , Yong Liu , Xiangtai Li , Dacheng Tao

End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject conditions in a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Zhenzhi Wang , Jiaqi Yang , Jianwen Jiang , Chao Liang , Gaojie Lin , Zerong Zheng , Ceyuan Yang , Yuan Zhang , Mingyuan Gao , Dahua Lin

Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio-visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective…

Computation and Language · Computer Science 2026-01-21 Qian Chen , Jinlan Fu , Changsong Li , See-Kiong Ng , Xipeng Qiu
‹ Prev 1 8 9 10 Next ›