中文
相关论文

相关论文: OmniShow: Unifying Multimodal Conditions for Human…

200 篇论文

Current multimodal large language models (MLLMs) have demonstrated remarkable capabilities in short-form video understanding, yet translating long-form cinematic videos into detailed, temporally grounded scripts remains a significant…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Junfu Pu , Yuxin Chen , Teng Wang , Ying Shan

Character image animation, which synthesizes videos of reference characters driven by pose sequences, has advanced rapidly but remains largely limited to single-human settings. Existing methods struggle to generalize to multi-humanoid…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Xirui Hu , Yanbo Ding , Jiahao Wang , Tingting Shi , Yali Wang , Guo Zhi Zhi , Weizhan Zhang

Most existing video diffusion models (VDMs) are limited to mere text conditions. Thereby, they are usually lacking in control over visual appearance and geometry structure of the generated videos. This work presents Moonshot, a new video…

计算机视觉与模式识别 · 计算机科学 2024-01-04 David Junhao Zhang , Dongxu Li , Hung Le , Mike Zheng Shou , Caiming Xiong , Doyen Sahoo

Modeling 4D human-object interaction (HOI) is a compelling challenge in computer vision and an essential technology powering virtual and mixed-reality applications. While existing works have achieved promising results on specific HOI…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Mengfei Zhang , Jinlu Zhang , Zhigang Tu

Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality,…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Hui Li , Mingwang Xu , Yun Zhan , Shan Mu , Jiaye Li , Kaihui Cheng , Yuxuan Chen , Tan Chen , Mao Ye , Jingdong Wang , Siyu Zhu

Understanding the human-object interactions (HOIs) from a video is essential to fully comprehend a visual scene. This line of research has been addressed by detecting HOIs from images and lately from videos. However, the video-based HOI…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Zhifan Ni , Esteve Valls Mascaró , Hyemin Ahn , Dongheui Lee

Most methods for conditional video synthesis use a single modality as the condition. This comes with major limitations. For example, it is problematic for a model conditioned on an image to generate a specific motion trajectory desired by…

计算机视觉与模式识别 · 计算机科学 2022-03-08 Ligong Han , Jian Ren , Hsin-Ying Lee , Francesco Barbieri , Kyle Olszewski , Shervin Minaee , Dimitris Metaxas , Sergey Tulyakov

This study focuses on a challenging yet promising task, Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text conditions, meanwhile ensuring both modalities are aligned with text. Despite…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Kaisi Guan , Xihua Wang , Zhengfeng Lai , Xin Cheng , Peng Zhang , XiaoJiang Liu , Ruihua Song , Meng Cao

We present Uni-Inter, a unified framework for human motion generation that supports a wide range of interaction scenarios: including human-human, human-object, and human-scene-within a single, task-agnostic architecture. In contrast to…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Sheng Liu , Yuanzhi Liang , Jiepeng Wang , Sidan Du , Chi Zhang , Xuelong Li

Human-Object Interaction (HOI) detection aims to understand the interactions between humans and objects, which plays a curtail role in high-level semantic understanding tasks. However, most works pursue designing better architectures to…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Shuman Fang , Shuai Liu , Jie Li , Guannan Jiang , Xianming Lin , Rongrong Ji

We tackle the dual challenges of video understanding and controllable video generation within a unified diffusion framework. Our key insights are two-fold: geometry-only cues (e.g., depth, edges) are insufficient: they specify layout but…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Dianbing Xi , Jiepeng Wang , Yuanzhi Liang , Xi Qiu , Jialun Liu , Hao Pan , Yuchi Huo , Rui Wang , Haibin Huang , Chi Zhang , Xuelong Li

Generating a video given the first several static frames is challenging as it anticipates reasonable future frames with temporal coherence. Besides video prediction, the ability to rewind from the last frame or infilling between the head…

计算机视觉与模式识别 · 计算机科学 2023-03-23 Tsu-Jui Fu , Licheng Yu , Ning Zhang , Cheng-Yang Fu , Jong-Chyi Su , William Yang Wang , Sean Bell

Recently, interactive digital human video generation has attracted widespread attention and achieved remarkable progress. However, building such a practical system that can interact with diverse input signals in real time remains…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Ming Chen , Liyuan Cui , Wenyuan Zhang , Haoxian Zhang , Yan Zhou , Xiaohan Li , Songlin Tang , Jiwen Liu , Borui Liao , Hejia Chen , Xiaoqiang Liu , Pengfei Wan

A long-standing objective in humanoid robotics is the realization of versatile agents capable of following diverse multimodal instructions with human-level flexibility. Despite advances in humanoid control, bridging high-level multimodal…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Nan Jiang , Zimo He , Wanhe Yu , Lexi Pang , Yunhao Li , Hongjie Li , Jieming Cui , Yuhan Li , Yizhou Wang , Yixin Zhu , Siyuan Huang

Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for embodied intelligence. A major challenge in this setting is…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Yiren Song , Xiyao Deng , Pei Yang , Yihan Wang , Mike Zheng Shou

We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a shared diffusion…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Junyi Chen , Tong He , Zhoujie Fu , Pengfei Wan , Kun Gai , Weicai Ye

We present Omni-Video 2, a scalable and computationally efficient model that connects pretrained multimodal large-language models (MLLMs) with video diffusion models for unified video generation and editing. Our key idea is to exploit the…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Hao Yang , Zhiyu Tan , Jia Gong , Luozheng Qin , Hesen Chen , Xiaomeng Yang , Yuqing Sun , Yuetan Lin , Mengping Yang , Hao Li

In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with existing benchmarks, our OmniEval has several distinctive…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Yiman Zhang , Ziheng Luo , Qiangyu Yan , Wei He , Borui Jiang , Xinghao Chen , Kai Han

Storytelling in real-world videos often unfolds through multiple shots -- discontinuous yet semantically connected clips that together convey a coherent narrative. However, existing multi-shot video generation (MSV) methods struggle to…

Human image animation involves generating videos from a character photo, allowing user control and unlocking the potential for video and movie production. While recent approaches yield impressive results using high-quality training data,…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Zhenzhi Wang , Yixuan Li , Yanhong Zeng , Youqing Fang , Yuwei Guo , Wenran Liu , Jing Tan , Kai Chen , Tianfan Xue , Bo Dai , Dahua Lin