中文
相关论文

相关论文: Stargazer: An Interactive Camera Robot for Capturi…

200 篇论文

Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to the challenge of…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Joonghyuk Shin , Daehyeon Choi , Jaesik Park

On-orbit servicing represents a critical frontier in future aerospace engineering, with the manipulation of dynamic non-cooperative targets serving as a key technology. In microgravity environments, objects are typically free-floating,…

机器人学 · 计算机科学 2026-03-31 Siyi Lang , Hongyi Gao , Yingxin Zhang , Zihao Liu , Hanlin Dong , Zhaoke Ning , Zhiqiang Ma , Panfeng Huang

Recent advancements in video generation have been remarkable, yet many existing methods struggle with issues of consistency and poor text-video alignment. Moreover, the field lacks effective techniques for text-guided video inpainting, a…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Bojia Zi , Shihao Zhao , Xianbiao Qi , Jianan Wang , Yukai Shi , Qianyu Chen , Bin Liang , Kam-Fai Wong , Lei Zhang

This paper investigates the challenge of extracting highlight moments from videos. To perform this task, we need to understand what constitutes a highlight for arbitrary video domains while at the same time being able to scale across…

计算机视觉与模式识别 · 计算机科学 2024-06-26 David Chuan-En Lin , Fabian Caba Heilbron , Joon-Young Lee , Oliver Wang , Nikolas Martelaro

The job of a camera operator is challenging, and potentially dangerous, when filming long moving camera shots. Broadly, the operator must keep the actors in-frame while safely navigating around obstacles, and while fulfilling an artistic…

图形学 · 计算机科学 2022-05-03 Mohamed Sayed , Robert Cinca , Enrico Costanza , Gabriel Brostow

Active perception, the ability of a robot to proactively adjust its viewpoint to acquire task-relevant information, is essential for robust operation in unstructured real-world environments. While critical for downstream tasks such as…

机器人学 · 计算机科学 2026-03-03 Yongxi Huang , Zhuohang Wang , Wenjing Tang , Cewu Lu , Panpan Cai

The goal of this work is to generate step-by-step visual instructions in the form of a sequence of images, given an input image that provides the scene context and the sequence of textual instructions. This is a challenging problem as it…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Tomáš Souček , Prajwal Gatti , Michael Wray , Ivan Laptev , Dima Damen , Josef Sivic

Classical policy search algorithms for robotics typically require performing extensive explorations, which are time-consuming and expensive to implement with real physical platforms. To facilitate the efficient learning of robot…

机器人学 · 计算机科学 2023-04-25 Shengzeng Huo , Anqing Duan , Lijun Han , Luyin Hu , Hesheng Wang , David Navarro-Alarcon

Video surveillance cameras generate most of recorded video, and there is far more recorded video than operators can watch. Much progress has recently been made using summarization of recorded video, but such techniques do not have much…

计算机视觉与模式识别 · 计算机科学 2017-01-05 Yedid Hoshen , Shmuel Peleg

We present a novel method for aligning a sequence of instructions to a video of someone carrying out a task. In particular, we focus on the cooking domain, where the instructions correspond to the recipe. Our technique relies on an HMM to…

计算与语言 · 计算机科学 2015-03-16 Jonathan Malmaud , Jonathan Huang , Vivek Rathod , Nick Johnston , Andrew Rabinovich , Kevin Murphy

Despite recent progress, video generative models still struggle to animate static images into videos that portray delicate human actions, particularly when handling uncommon or novel actions whose training data are limited. In this paper,…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Haoxin Li , Yingchen Yu , Qilong Wu , Hanwang Zhang , Song Bai , Boyang Li

Learning from demonstration allows for rapid deployment of robot manipulators to a great many tasks, by relying on a person showing the robot what to do rather than programming it. While this approach provides many opportunities, measuring,…

机器人学 · 计算机科学 2019-05-13 Aran Sena , Matthew J Howard

Recent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Zeqi Xiao , Wenqi Ouyang , Yifan Zhou , Shuai Yang , Lei Yang , Jianlou Si , Xingang Pan

Imitation learning is an effective approach for autonomous systems to acquire control policies when an explicit reward function is unavailable, using supervision provided as demonstrations from an expert, typically a human operator.…

机器学习 · 计算机科学 2018-06-20 YuXuan Liu , Abhishek Gupta , Pieter Abbeel , Sergey Levine

In cinema, large camera lenses create beautiful shallow depth of field (DOF), but make focusing difficult and expensive. Accurate cinema focus usually relies on a script and a person to control focus in realtime. Casual videographers often…

计算机视觉与模式识别 · 计算机科学 2019-05-22 Xuaner Zhang , Kevin Matzen , Vivien Nguyen , Dillon Yao , You Zhang , Ren Ng

Harnessing human movements to command an Unmanned Aerial Vehicle (UAV) holds the potential to revolutionize their deployment, rendering it more intuitive and user-centric. In this research, we introduce a novel methodology adept at…

机器人学 · 计算机科学 2024-08-20 Akash Chaudhary , Tiago Nascimento , Martin Saska

Purpose - Most industrial robots are still programmed using the typical teaching process, through the use of the robot teach pendant. This is a tedious and time-consuming task that requires some technical expertise, and hence new approaches…

机器人学 · 计算机科学 2013-09-10 Pedro Neto , Norberto Pires , Paulo Moreira

Augmented and mixed-reality techniques harbor a great potential for improving human-robot collaboration. Visual signals and cues may be projected to a human partner in order to explicitly communicate robot intentions and goals. However, it…

机器人学 · 计算机科学 2023-08-22 Shubham Sonawani , Yifan Zhou , Heni Ben Amor

Robots that must operate in novel environments and collaborate with humans must be capable of acquiring new knowledge from human experts during operation. We propose teaching a robot novel objects it has not encountered before by pointing a…

机器人学 · 计算机科学 2020-12-29 Sagar Gubbi Venkatesh , Raviteja Upadrashta , Shishir Kolathaya , Bharadwaj Amrutur

This paper presents an unsupervised approach towards automatically extracting video-based guidance on object usage, from egocentric video and wearable gaze tracking, collected from multiple users while performing tasks. The approach i)…

计算机视觉与模式识别 · 计算机科学 2016-03-22 Dima Damen , Teesid Leelasawassuk , Walterio Mayol-Cuevas