中文
相关论文

相关论文: ActAvatar: Temporally-Aware Precise Action Control…

200 篇论文

Vision-Language-Action (VLA) models have demonstrated remarkable capabilities and generalization in embodied manipulation. However, their decision-making relies on a fast, instinctive process that lacks deliberation. This strategy often…

机器人学 · 计算机科学 2026-05-29 Wenhao Li , Xiu Su , Yichao Cao , Hongyan Xu , Xiaobo Xia , Shan You , Yi Chen , Chang Xu

Text-to-Motion (T2M) generation aims to synthesize realistic human motion sequences from natural language descriptions. While two-stage frameworks leveraging discrete motion representations have advanced T2M research, they often neglect…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Hongsong Wang , Wenjing Yan , Qiuxia Lai , Xin Geng

We propose a standalone autoregressive (AR) Action Expert that generates actions as a continuous causal sequence while conditioning on refreshable vision-language prefixes. In contrast to existing Vision-Language-Action (VLA) models and…

Audio-Visual Question Answering (AVQA) requires not only question-based multimodal reasoning but also precise temporal grounding to capture subtle dynamics for accurate prediction. However, existing methods mainly use question information…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Hongyeob Kim , Inyoung Jung , Dayoon Suh , Youjia Zhang , Sangmin Lee , Sungeun Hong

TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to…

多媒体 · 计算机科学 2026-03-20 Xinyuan Qian , Xinjia Zhu , Alessio Brutti , Dong Liang

Accurate engagement estimation is essential for adaptive human-computer interaction systems, yet robust deployment is hindered by poor generalizability across diverse domains and challenges in modeling complex interaction dynamics.To tackle…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Yangche Yu , Yin Chen , Jia Li , Peng Jia , Yu Zhang , Li Dai , Zhenzhen Hu , Meng Wang , Richang Hong

We present SentiAvatar, a framework for building expressive interactive 3D digital humans, and use it to create SuSu, a virtual character that speaks, gestures, and emotes in real time. Achieving such a system remains challenging, as it…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Chuhao Jin , Rui Zhang , Qingzhe Gao , Haoyu Shi , Dayu Wu , Yichen Jiang , Yihan Wu , Ruihua Song

We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input audio clip. Unlike prior methods that emphasize lip…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Chenxu Zhang , Zenan Li , Hongyi Xu , You Xie , Xiaochen Zhao , Tianpei Gu , Guoxian Song , Xin Chen , Chao Liang , Jianwen Jiang , Linjie Luo

Dialogue act annotations are important to improve response generation quality in task-oriented dialogue systems. However, it can be challenging to use dialogue acts to control response generation in a generalizable way because different…

计算与语言 · 计算机科学 2023-08-03 Qingyang Wu , James Gung , Raphael Shu , Yi Zhang

Over the past years, significant progress has been made in creating photorealistic and drivable 3D avatars solely from videos of real humans. However, a core remaining challenge is the fine-grained and user-friendly editing of clothing…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Basavaraj Sunagad , Heming Zhu , Mohit Mendiratta , Adam Kortylewski , Christian Theobalt , Marc Habermann

Audio-visual embodied navigation, as a hot research topic, aims training a robot to reach an audio target using egocentric visual (from the sensors mounted on the robot) and audio (emitted from the target) input. The audio-visual…

声音 · 计算机科学 2022-10-06 Yinfeng Yu , Lele Cao , Fuchun Sun , Xiaohong Liu , Liejun Wang

Data replay is a successful incremental learning technique for images. It prevents catastrophic forgetting by keeping a reservoir of previous data, original or synthesized, to ensure the model retains past knowledge while adapting to novel…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Guodong Ding , Hans Golong , Angela Yao

Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present AVoCaDO, a…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Xinlong Chen , Yue Ding , Weihong Lin , Jingyun Hua , Linli Yao , Yang Shi , Bozhou Li , Yuanxing Zhang , Qiang Liu , Pengfei Wan , Liang Wang , Tieniu Tan

In this paper, we introduce Attention Prompt Tuning (APT) - a computationally efficient variant of prompt tuning for video-based applications such as action recognition. Prompt tuning approaches involve injecting a set of learnable prompts…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Wele Gedara Chaminda Bandara , Vishal M. Patel

Accurate and stable field-of-view (FoV) guidance is critical for safe and efficient minimally invasive surgery, yet existing approaches often conflate visual attention estimation with downstream camera control or rely on direct…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Rulin Zhou , Guankun Wang , An Wang , Yujie Ma , Lixin Ouyang , Bolin Cui , Junyan Li , Chaowei Zhu , Mingyang Li , Ming Chen , Xiaopin Zhong , Peng Lu , Jiankun Wang , Xianming Liu , Hongliang Ren

In order to be widely applicable, speech-driven 3D head avatars must articulate their lips in accordance with speech, while also conveying the appropriate emotions with dynamically changing facial expressions. The key problem is that…

图形学 · 计算机科学 2026-01-28 Radek Daněček , Carolin Schmitt , Senya Polikovsky , Michael J. Black

In video-based emotion recognition, audio and visual modalities are often expected to have a complementary relationship, which is widely explored using cross-attention. However, they may also exhibit weak complementary relationships,…

计算机视觉与模式识别 · 计算机科学 2024-03-29 R. Gnana Praveen , Jahangir Alam

Evaluating whether human action is standard or not and providing reasonable feedback to improve action standardization is very crucial but challenging in real-world scenarios. However, current video understanding methods are mainly…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Mengshi Qi , Yeteng Wu , Xianlin Zhang , Huadong Ma

Many vision-language tasks can be reduced to the problem of sequence prediction for natural language output. In particular, recent advances in image captioning use deep reinforcement learning (RL) to alleviate the "exposure bias" during…

计算机视觉与模式识别 · 计算机科学 2018-08-23 Daqing Liu , Zheng-Jun Zha , Hanwang Zhang , Yongdong Zhang , Feng Wu

Vision-Language-Action (VLA) models have demonstrated strong performance in robotic manipulation, yet their closed-loop deployment is hindered by the high latency and compute cost of repeatedly running large vision-language backbones at…

机器人学 · 计算机科学 2026-01-28 Wenda Yu , Tianshi Wang , Fengling Li , Jingjing Li , Lei Zhu