中文
相关论文

相关论文: MagicAnime: A Hierarchically Annotated, Multimodal…

200 篇论文

Face animation has received a lot of attention from researchers in recent years due to its wide range of promising applications. Many face animation models based on optical flow or deep neural networks have achieved great success. However,…

计算机视觉与模式识别 · 计算机科学 2023-04-13 Zhaoying Pan , Jinge Ma

Recently, artificial intelligence techniques for education have been received increasing attentions, while it still remains an open problem to design the effective music instrument instructing systems. Although key presses can be directly…

声音 · 计算机科学 2025-02-26 Qijun Gan , Song Wang , Shengtao Wu , Jianke Zhu

We present an efficient text-to-video generation framework based on latent diffusion models, termed MagicVideo. MagicVideo can generate smooth video clips that are concordant with the given text descriptions. Due to a novel and efficient 3D…

计算机视觉与模式识别 · 计算机科学 2023-05-12 Daquan Zhou , Weimin Wang , Hanshu Yan , Weiwei Lv , Yizhe Zhu , Jiashi Feng

Dense pixel-specific representation learning at scale has been bottlenecked due to the unavailability of large-scale multi-view datasets. Current methods for building effective pretraining datasets heavily rely on annotated 3D meshes, point…

3D characters are essential to modern creative industries, but making them animatable often demands extensive manual work in tasks like rigging and skinning. Existing automatic rigging tools face several limitations, including the necessity…

图形学 · 计算机科学 2025-03-12 Zhiyang Guo , Jinxu Xiang , Kai Ma , Wengang Zhou , Houqiang Li , Ran Zhang

Research on diffusion model-based video generation has advanced rapidly. However, limitations in object fidelity and generation length hinder its practical applications. Additionally, specific domains like animated wallpapers require…

计算机视觉与模式识别 · 计算机科学 2024-04-17 Fanyi Wang , Peng Liu , Haotian Hu , Dan Meng , Jingwen Su , Jinjin Xu , Yanhao Zhang , Xiaoming Ren , Zhiwang Zhang

This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language models (MLLMs) to predict the emotions evoked by images. Recent user studies report an…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Tianwei Chen , Takuya Furusawa , Yuki Hirakawa , Ryotaro Shimizu , Mo Fan , Takashi Wada

Multi-instance Repetitive Action Counting (MRAC) aims to estimate the number of repetitive actions performed by multiple instances in untrimmed videos, commonly found in human-centric domains like sports and exercise. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2024-09-09 Yin Tang , Wei Luo , Jinrui Zhang , Wei Huang , Ruihai Jing , Deyu Zhang

Contrastive pre-training on large-scale image-text pair datasets has driven major advances in vision-language representation learning. Recent work shows that pretraining on global data followed by language or culture specific fine-tuning is…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Issa Sugiura , Shuhei Kurita , Yusuke Oda , Daisuke Kawahara , Yasuo Okabe , Naoaki Okazaki

Multi-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury. To enable representation learning for medical images at scale, we turn to YouTube, a platform with a…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Wisdom O. Ikezogwo , Kevin Zhang , Mehmet Saygin Seyfioglu , Fatemeh Ghezloo , Linda Shapiro , Ranjay Krishna

Head avatars animated by visual signals have gained popularity, particularly in cross-driving synthesis where the driver differs from the animated character, a challenging but highly practical approach. The recently presented MegaPortraits…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Nikita Drobyshev , Antoni Bigata Casademunt , Konstantinos Vougioukas , Zoe Landgraf , Stavros Petridis , Maja Pantic

Recent advances in Generative AI (GenAI) have led to significant improvements in the quality of generated visual content. As AI-generated visual content becomes increasingly indistinguishable from real content, the challenge of detecting…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Keerthi Veeramachaneni , Praveen Tirupattur , Amrit Singh Bedi , Mubarak Shah

Dynamical systems theory and reinforcement learning view world evolution as latent-state dynamics driven by actions, with visual observations providing partial information about the state. Recent video world models attempt to learn this…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Zhen Li , Zian Meng , Shuwei Shi , Wenshuo Peng , Yuwei Wu , Bo Zheng , Chuanhao Li , Kaipeng Zhang

In multi-modal dialogue systems, it is important to allow the use of images as part of a multi-turn conversation. Training such dialogue systems generally requires a large-scale dataset consisting of multi-turn dialogues that involve…

计算与语言 · 计算机科学 2021-07-20 Nyoungwoo Lee , Suwon Shin , Jaegul Choo , Ho-Jin Choi , Sung-Hyun Myaeng

The field has made significant progress in synthesizing realistic human motion driven by various modalities. Yet, the need for different methods to animate various body parts according to different control signals limits the scalability of…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Zixiang Zhou , Yu Wan , Baoyuan Wang

We present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users. Our objective is to optimize for a set of high-quality data samples…

Comprehensive evaluation of Multimodal Large Language Models (MLLMs) has recently garnered widespread attention in the research community. However, we observe that existing benchmarks present several common barriers that make it difficult…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Yi-Fan Zhang , Huanyu Zhang , Haochen Tian , Chaoyou Fu , Shuangqing Zhang , Junfei Wu , Feng Li , Kun Wang , Qingsong Wen , Zhang Zhang , Liang Wang , Rong Jin , Tieniu Tan

Recent advancements in Large Video-Language Models (LVLMs) have led to promising results in multimodal video understanding. However, it remains unclear whether these models possess the cognitive capabilities required for high-level tasks,…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Chenglin Li , Qianglong Chen , Zhi Li , Feng Tao , Yin Zhang

Deep video action recognition models have been highly successful in recent years but require large quantities of manually annotated data, which are expensive and laborious to obtain. In this work, we investigate the generation of synthetic…

计算机视觉与模式识别 · 计算机科学 2019-10-16 César Roberto de Souza , Adrien Gaidon , Yohann Cabon , Naila Murray , Antonio Manuel López

Creating 3D character animations traditionally requires significant time and effort from the animator. Advancements in generative methods now enable easy creation of multiple character animation variations for use or further editing.…

人机交互 · 计算机科学 2026-05-05 Ludwig Sidenmark , Qian Zhou , George Fitzmaurice , Fraser Anderson
‹ 上一页 1 8 9 10 下一页 ›