中文
相关论文

相关论文: Stargazer: An Interactive Camera Robot for Capturi…

200 篇论文

User-generated cinematic creations are gaining popularity as our daily entertainment, yet it is a challenge to master cinematography for producing immersive contents. Many existing automatic methods focus on roughly controlling predefined…

多媒体 · 计算机科学 2024-05-24 Xinyi Wu , Haohong Wang , Aggelos K. Katsaggelos

Large-scale video generative models have recently demonstrated strong visual capabilities, enabling the prediction of future frames that adhere to the logical and physical cues in the current observation. In this work, we investigate…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Gongfan Fang , Xinyin Ma , Xinchao Wang

This paper explores the integration of visual communication and musical interaction by implementing a robotic camera within a "Guided Harmony" musical game. We aim to examine co-creative behaviors between human musicians and robotic…

人机交互 · 计算机科学 2024-10-29 Ross Greer , Laura Fleig , Shlomo Dubnov

Always-on egocentric cameras are increasingly used as demonstrations for embodied robotics, imitation learning, and assistive AR, but the resulting video streams are dominated by redundant and low-quality frames. Under the storage and…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Ajan Subramanian , Sumukh Bettadapura , Rohan Sathish

Digital creators, from indie filmmakers to animation studios, face a persistent bottleneck: translating their creative vision into precise camera movements. Despite significant progress in computer vision and artificial intelligence,…

Video captioning is the task of automatically generating a textual description of the actions in a video. Although previous work (e.g. sequence-to-sequence model) has shown promising results in abstracting a coarse description of a short…

计算机视觉与模式识别 · 计算机科学 2018-03-30 Xin Wang , Wenhu Chen , Jiawei Wu , Yuan-Fang Wang , William Yang Wang

Temporal action segmentation in videos has drawn much attention recently. Timestamp supervision is a cost-effective way for this task. To obtain more information to optimize the model, the existing method generated pseudo frame-wise labels…

计算机视觉与模式识别 · 计算机科学 2022-12-14 Yang Zhao , Yan Song

In this work, we propose an interactive general instruction framework SketchMeHow to guidance the common users to complete the daily tasks in real-time. In contrast to the conventional augmented reality-based instruction systems, the…

人机交互 · 计算机科学 2021-09-08 Haoran Xie , Yichen Peng , Hange Wang , Kazunori Miyata

This work introduces a robot navigation controller that combines event cameras and other sensors with reinforcement learning to enable real-time human-centered navigation and obstacle avoidance. Unlike conventional image-based controllers,…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Ignacio Bugueno-Cordova , Javier Ruiz-del-Solar , Rodrigo Verschae

Robotic technology can support the creation of new tools that improve the creative process of cinematography. It is crucial to consider the specific requirements and perspectives of industry professionals when designing and developing these…

机器人学 · 计算机科学 2023-04-18 Pragathi Praveena , Bengisu Cagiltay , Michael Gleicher , Bilge Mutlu

Vision-based perception systems are typically exposed to large orientation changes in different robot applications. In such conditions, their performance might be compromised due to the inherent complexity of processing data captured under…

Creating a vivid video from the event or scenario in our imagination is a truly fascinating experience. Recent advancements in text-to-video synthesis have unveiled the potential to achieve this with prompts only. While text is convenient…

计算机视觉与模式识别 · 计算机科学 2023-06-02 Jinbo Xing , Menghan Xia , Yuxin Liu , Yuechen Zhang , Yong Zhang , Yingqing He , Hanyuan Liu , Haoxin Chen , Xiaodong Cun , Xintao Wang , Ying Shan , Tien-Tsin Wong

We introduce Vocal Sandbox, a framework for enabling seamless human-robot collaboration in situated environments. Systems in our framework are characterized by their ability to adapt and continually learn at multiple levels of abstraction…

机器人学 · 计算机科学 2024-11-06 Jennifer Grannen , Siddharth Karamcheti , Suvir Mirchandani , Percy Liang , Dorsa Sadigh

We introduce a simple new method for visual imitation learning, which allows a novel robot manipulation task to be learned from a single human demonstration, without requiring any prior knowledge of the object being interacted with. Our…

机器人学 · 计算机科学 2021-06-11 Edward Johns

Traditionally, learning from human demonstrations via direct behavior cloning can lead to high-performance policies given that the algorithm has access to large amounts of high-quality data covering the most likely scenarios to be…

Can we learn robot manipulation for everyday tasks, only by watching videos of humans doing arbitrary tasks in different unstructured settings? Unlike widely adopted strategies of learning task-specific behaviors or direct imitation of a…

机器人学 · 计算机科学 2023-02-07 Homanga Bharadhwaj , Abhinav Gupta , Shubham Tulsiani , Vikash Kumar

The rise of autonomous AI agents suggests that dynamic benchmark environments with built-in feedback on scientifically grounded tasks are needed to evaluate the capabilities of these agents in research work. We introduce Stargazer, a…

机器学习 · 计算机科学 2026-05-13 Xinge Liu , Terry Jingchen Zhang , Bernhard Schölkopf , Zhijing Jin , Kristen Menou

With advancements in video generative AI models (e.g., SORA), creators are increasingly using these techniques to enhance video previsualization. However, they face challenges with incomplete and mismatched AI workflows. Existing methods…

计算机视觉与模式识别 · 计算机科学 2024-09-02 Yiran Chen , Anyi Rao , Xuekun Jiang , Shishi Xiao , Ruiqing Ma , Zeyu Wang , Hui Xiong , Bo Dai

This paper presents a new video question answering task on screencast tutorials. We introduce a dataset including question, answer and context triples from the tutorial videos for a software. Unlike other video question answering works, all…

计算与语言 · 计算机科学 2020-08-04 Wentian Zhao , Seokhwan Kim , Ning Xu , Hailin Jin

While the human eye can perceive an impressive twenty stops of dynamic range, smartphone camera sensors remain limited to about twelve stops despite decades of research. A variety of high dynamic range (HDR) image capture and processing…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Baiang Li , Ruyu Yan , Ethan Tseng , Zhoutong Zhang , Adam Finkelstein , Jiawen Chen , Felix Heide