English
Related papers

Related papers: ViSTAR: Virtual Skill Training with Augmented Real…

200 papers

Reinforcement Learning (RL) is a rapidly growing area of machine learning that finds its application in a broad range of domains, from finance and healthcare to robotics and gaming. Compared to other machine learning techniques, RL agents…

Artificial Intelligence · Computer Science 2024-11-14 Geetansh Kalra , Divye Singh , Justin Jose

Vision-language-action (VLA) models demonstrate strong generalization in robotic manipulation but face challenges in complex, real-world tasks. While supervised fine-tuning with demonstrations is constrained by data quality, reinforcement…

Robotics · Computer Science 2025-09-18 Piaopiao Jin , Qi Wang , Guokang Sun , Ziwen Cai , Pinjia He , Yangwei You

Developing robust and general-purpose manipulation policies represents a fundamental objective in robotics research. While Vision-Language-Action (VLA) models have demonstrated promising capabilities for end-to-end robot control, existing…

Multi-object tracking (MOT) is crucial for various multi-agent analyses such as evaluating team sports tactics and player movements and performance. While pedestrian tracking has advanced with Tracking-by-Detection MOT, team sports like…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Li Yin , Calvin Yeung , Qingrui Hu , Jun Ichikawa , Hirotsugu Azechi , Susumu Takahashi , Keisuke Fujii

Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human behavior as a translation task co-speech gesture or text-to-motion that maps a fixed…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Juze Zhang , Changan Chen , Xin Chen , Heng Yu , Tiange Xiang , Ali Sartaz Khan , Shrinidhi K. Lakshmikanth , Ehsan Adeli

Post-training Large Vision-and-Language Models (LVLMs) typically involves Supervised Fine-Tuning (SFT) for knowledge injection or Reinforcement Learning with Verifiable Rewards (RLVR) for performance enhancement. However, SFT often leads to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Yuqi Liu , Liangyu Chen , Jiazhen Liu , Mingkang Zhu , Zhisheng Zhong , Bei Yu , Jiaya Jia

Pre-training visual and textual representations from large-scale image-text pairs is becoming a standard approach for many downstream vision-language tasks. The transformer-based models learn inter and intra-modal attention through a list…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Mohammad Abuzar Hashemi , Zhanghexuan Li , Mihir Chauhan , Yan Shen , Abhishek Satbhai , Mir Basheer Ali , Mingchen Gao , Sargur Srihari

Recent advancements in physics-based character animation leverage deep learning to generate agile and natural motion, enabling characters to execute movements such as backflips, boxing, and tennis. However, reproducing the selection and use…

Graphics · Computer Science 2024-07-24 Jiashun Wang , Jessica Hodgins , Jungdam Won

As deep reinforcement learning driven by visual perception becomes more widely used there is a growing need to better understand and probe the learned agents. Understanding the decision making process and its relationship to visual inputs…

Computer Vision and Pattern Recognition · Computer Science 2019-04-03 Christian Rupprecht , Cyril Ibrahim , Christopher J. Pal

Hybrid tutoring, where a human tutor supports multiple students in learning with educational technology, is an increasingly common application to deliver high-impact tutoring at scale. However, past hybrid tutoring applications are limited…

Human-Computer Interaction · Computer Science 2025-05-14 Eason Chen , Xinyi Tang , Aprille Xi , Chenyu Lin , Conrad Borchers , Shivang Gupta , Jionghao Lin , Kenneth R Koedinger

Large Language Models (LLMs) are trained and aligned to follow natural language instructions with only a handful of examples, and they are prompted as task-driven autonomous agents to adapt to various sources of execution environments.…

Computation and Language · Computer Science 2023-10-03 Yang Su

We propose AvatarVTON, the first 4D virtual try-on framework that generates realistic try-on results from a single in-shop garment image, enabling free pose control, novel-view rendering, and diverse garment choices. Unlike existing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Zicheng Jiang , Jixin Gao , Shengfeng He , Xinzhe Li , Yulong Zheng , Zhaotong Yang , Junyu Dong , Yong Du

The growing adoption of augmented and virtual reality (AR and VR) technologies in industrial training and on-the-job assistance has created new opportunities for intelligent, context-aware support systems. As workers perform complex tasks…

Human-Computer Interaction · Computer Science 2025-11-18 Mahya Qorbani , Kamran Paynabar , Mohsen Moghaddam

Recently, rapid advancements have been made in multimodal large language models (MLLMs), especially in video understanding tasks. However, current research focuses on simple video scenarios, failing to reflect the complex and diverse nature…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Lu Zhu , Tiantian Geng , Yangye Chen , Teng Wang , Ping Lu , Feng Zheng

Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for generalist robotic control. Built upon vision-language model (VLM) architectures, VLAs predict actions conditioned on visual observations and language…

Robotics · Computer Science 2026-05-26 Weikang Qiu , Huashuo Lei , Tinglin Huang , Rex Ying

Video-based spatial cognition is vital for robotics and embodied AI but challenges current Vision-Language Models (VLMs). This paper makes two key contributions. First, we introduce ViCA (Visuospatial Cognitive Assistant)-322K, a diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Qi Feng

We study reward models for long-horizon manipulation tasks by learning from action-free videos and language instructions, which we term the visual-instruction correlation (VIC) problem. Recent advancements in cross-modality modeling have…

Robotics · Computer Science 2025-02-21 Kuo-Han Hung , Pang-Chi Lo , Jia-Fong Yeh , Han-Yuan Hsu , Yi-Ting Chen , Winston H. Hsu

VLA models have achieved remarkable progress in embodied intelligence; however, their evaluation remains largely confined to simulations or highly constrained real-world settings. This mismatch creates a substantial reality gap, where…

A hallmark of advanced artificial intelligence is the capacity to progress from passive visual perception to the strategic modification of visual information to facilitate complex reasoning. This advanced capability, however, remains…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Jingkun Ma , Runzhe Zhan , Yang Li , Di Sun , Hou Pong Chan , Lidia S. Chao , Derek F. Wong

Existing for audio- and pose-driven human animation methods often struggle with stiff head movements and blurry hands, primarily due to the weak correlation between audio and head movements and the structural complexity of hands. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Donglin Huang , Yongyuan Li , Tianhang Liu , Junming Huang , Xiaoda Yang , Chi Wang , Weiwei Xu