中文
相关论文

相关论文: Stargazer: An Interactive Camera Robot for Capturi…

200 篇论文

Aerial filming is constantly gaining importance due to the recent advances in drone technology. It invites many intriguing, unsolved problems at the intersection of aesthetical and scientific challenges. In this work, we propose a deep…

机器人学 · 计算机科学 2019-10-16 Mirko Gschwindt , Efe Camci , Rogerio Bonatti , Wenshan Wang , Erdal Kayacan , Sebastian Scherer

Flipped learning is a method that flips in/out class activities to make lectures learner-centered. In flipped learning, comments from learners on preparation material are useful information for instructors to consider before deciding…

计算机与社会 · 计算机科学 2020-07-30 Shintaro Uchiyama , Hayato Okumoto , Mitsuo Yoshida , Yuko Ichikawa , Kyoji Umemura

We present a new learning approach, Soft Conditional Prompt Learning (SCP), which leverages the strengths of prompt learning for aerial video action recognition. Our approach is designed to predict the action of each agent by helping the…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Xijun Wang , Ruiqi Xian , Tianrui Guan , Fuxiao Liu , Dinesh Manocha

Robust and efficient learning remains a challenging problem in robotics, in particular with complex visual inputs. Inspired by human attention mechanism, with which we quickly process complex visual scenes and react to changes in the…

机器人学 · 计算机科学 2023-08-30 Daniel Scheuchenstuhl , Stefan Ulmer , Felix Resch , Luigi Berducci , Radu Grosu

Recent years have seen substantial progress in diffusion-based controllable video generation. However, achieving precise control in complex scenarios, including fine-grained object parts, sophisticated motion trajectories, and coherent…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Haitao Zhou , Chuang Wang , Rui Nie , Jinlin Liu , Dongdong Yu , Qian Yu , Changhu Wang

Interactive imitation learning makes an agent's control policy robust by stepwise supervisions from an expert. The recent algorithms mostly employ expert-agent switching systems to reduce the expert's burden by limitedly selecting the…

机器人学 · 计算机科学 2026-04-23 Taisuke Kobayashi

Nowadays, modeling exercises on software development objects are conducted in higher education institutions for information technology. Not only are there many defects such as missing elements in the models created by learners during the…

软件工程 · 计算机科学 2025-08-07 Yuta Saito , Takehiro Kokubu , Takafumi Tanaka , Atsuo Hazeyama , Hiroaki Hashiura

Videos provide a rich source of information, but it is generally hard to extract dynamical parameters of interest. Inferring those parameters from a video stream would be beneficial for physical reasoning. Robots performing tasks in dynamic…

机器人学 · 计算机科学 2020-08-31 Martin Asenov , Michael Burke , Daniel Angelov , Todor Davchev , Kartic Subr , Subramanian Ramamoorthy

We introduce $\textit{InteractiveVideo}$, a user-centric framework for video generation. Different from traditional generative approaches that operate based on user-provided images or text, our framework is designed for dynamic interaction,…

计算机视觉与模式识别 · 计算机科学 2024-02-06 Yiyuan Zhang , Yuhao Kang , Zhixin Zhang , Xiaohan Ding , Sanyuan Zhao , Xiangyu Yue

Motor skills, especially fine motor skills like handwriting, play an essential role in academic pursuits and everyday life. Traditional methods to teach these skills, although effective, can be time-consuming and inconsistent. With the rise…

机器学习 · 计算机科学 2024-08-13 Hadar Mulian , Segev Shlomov , Lior Limonad , Alessia Noccaro , Silvia Buscaglione

The endoscopic camera of a surgical robot provides surgeons with a magnified 3D view of the surgical field, but repositioning it increases mental workload and operation time. Poor camera placement contributes to safety-critical events when…

机器人学 · 计算机科学 2023-01-20 Kay Hutchinson , Mohammad Samin Yasar , Harshneet Bhatia , Homa Alemzadeh

Despite the recent progress in text-to-video generation, existing studies usually overlook the issue that only spatial contents but not temporal motions in synthesized videos are under the control of text. Towards such a challenge, this…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Xi Chen , Zhiheng Liu , Mengting Chen , Yutong Feng , Yu Liu , Yujun Shen , Hengshuang Zhao

Automatic video captioning aims to train models to generate text descriptions for all segments in a video, however, the most effective approaches require large amounts of manual annotation which is slow and expensive. Active learning is a…

计算机视觉与模式识别 · 计算机科学 2020-12-04 David M. Chan , Sudheendra Vijayanarasimhan , David A. Ross , John Canny

Imitation Learning from monocular video demonstrations provides a scalable approach for teaching complex skills to humanoid robots. However, translating human motion to humanoids requires overcoming significant morphological mismatches.…

Robotic manipulation tasks often rely on static cameras for perception, which can limit flexibility, particularly in scenarios like robotic surgery and cluttered environments where mounting static cameras is impractical. Ideally, robots…

机器人学 · 计算机科学 2025-09-18 Xiatao Sun , Francis Fan , Yinxing Chen , Daniel Rakita

We present HelpViz, a tool for generating contextual visual mobile tutorials from text-based instructions that are abundant on the web. HelpViz transforms text instructions to graphical tutorials in batch, by extracting a sequence of…

人机交互 · 计算机科学 2025-04-04 Mingyuan Zhong , Gang Li , Peggy Chi , Yang Li

Observing a human demonstrator manipulate objects provides a rich, scalable and inexpensive source of data for learning robotic policies. However, transferring skills from human videos to a robotic manipulator poses several challenges, not…

机器人学 · 计算机科学 2023-03-08 Minttu Alakuijala , Gabriel Dulac-Arnold , Julien Mairal , Jean Ponce , Cordelia Schmid

Capturing an event from multiple camera angles can give a viewer the most complete and interesting picture of that event. To be suitable for broadcasting, a human director needs to decide what to show at each point in time. This can become…

计算机视觉与模式识别 · 计算机科学 2022-08-11 Bram Vanherle , Tim Vervoort , Nick Michiels , Philippe Bekaert

Realistic simulators are critical for training and verifying robotics systems. While most of the contemporary simulators are hand-crafted, a scaleable way to build simulators is to use machine learning to learn how the environment behaves…

计算机视觉与模式识别 · 计算机科学 2021-05-03 Seung Wook Kim , Jonah Philion , Antonio Torralba , Sanja Fidler

In video prediction tasks, one major challenge is to capture the multi-modal nature of future contents and dynamics. In this work, we propose a simple yet effective framework that can efficiently predict plausible future states. The key…

计算机视觉与模式识别 · 计算机科学 2020-07-06 Jingwei Xu , Huazhe Xu , Bingbing Ni , Xiaokang Yang , Trevor Darrell