中文
相关论文

相关论文: Batteries, camera, action! Learning a semantic con…

200 篇论文

Camera sensor simulation serves as a critical role for autonomous driving (AD), e.g. evaluating vision-based AD algorithms. While existing approaches have leveraged generative models for controllable image/video generation, they remain…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Wenchao Sun , Xuewu Lin , Keyu Chen , Zixiang Pei , Yining Shi , Chuang Zhang , Sifa Zheng

We consider the problem of estimating the parameters of a vehicle dynamics model for predictive control in driving applications. Instead of solely using the instantaneous parameters estimated from the vehicle signals, we combine this with…

系统与控制 · 电气工程与系统科学 2025-11-17 Marcus Greiff , Ray Zhang , Takeru Shirasawa , John Subosits

Camera-controllable video generation aims to synthesize videos with flexible and physically plausible camera movements. However, existing methods either provide imprecise camera control from text prompts or rely on labor-intensive manual…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Haoyu Zhao , Zihao Zhang , Jiaxi Gu , Haoran Chen , Qingping Zheng , Pin Tang , Yeyin Jin , Yuang Zhang , Junqi Cheng , Zenghui Lu , Peng Shu , Zuxuan Wu , Yu-Gang Jiang

Machine-learning excels in many areas with well-defined goals. However, a clear goal is usually not available in art forms, such as photography. The success of a photograph is measured by its aesthetic value, a very subjective concept. This…

计算机视觉与模式识别 · 计算机科学 2017-07-13 Hui Fang , Meng Zhang

Planning with world models offers a powerful paradigm for robotic control. Conventional approaches train a model to predict future frames conditioned on current frames and actions, which can then be used for planning. However, the objective…

机器学习 · 计算机科学 2025-10-23 Jacob Berg , Chuning Zhu , Yanda Bao , Ishan Durugkar , Abhishek Gupta

Aerial cinematography is significantly expanding the capabilities of film-makers. Recent progress in autonomous unmanned aerial vehicles (UAVs) has further increased the potential impact of aerial cameras, with systems that can safely track…

机器人学 · 计算机科学 2021-04-02 Arthur Bucker , Rogerio Bonatti , Sebastian Scherer

This paper introduces CameraCtrl II, a framework that enables large-scale dynamic scene exploration through a camera-controlled video diffusion model. Previous camera-conditioned video generative models suffer from diminished video dynamics…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Hao He , Ceyuan Yang , Shanchuan Lin , Yinghao Xu , Meng Wei , Liangke Gui , Qi Zhao , Gordon Wetzstein , Lu Jiang , Hongsheng Li

Video is a rich and scalable source of 3D/4D visual observations, and camera control is a key capability for video generation models to produce geometrically meaningful content. Existing approaches typically learn a mapping from camera…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Chen Hou , Christian Rupprecht

Generative AI has made image creation more accessible, yet aligning outputs with nuanced creative intent remains challenging, particularly for non-experts. Existing tools often require users to externalize ideas through prompts or…

人机交互 · 计算机科学 2025-08-11 Daniel Lee , Nikhil Sharma , Donghoon Shin , DaEun Choi , Harsh Sharma , Jeonghwan Kim , Heng Ji

Change detection is an important problem in vision field, especially for aerial images. However, most works focus on traditional change detection, i.e., where changes happen, without considering the change type information, i.e., what…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Wensheng Cheng , Yan Zhang , Xu Lei , Wen Yang , Guisong Xia

The rise of Unmanned Aerial Vehicles and their increasing use in the cinema industry calls for the creation of dedicated tools. Though there is a range of techniques to automatically control drones for a variety of applications, none have…

机器人学 · 计算机科学 2017-12-13 Quentin Galvane , Julien Fleureau , Francois-Louis Tariolle , Philippe Guillotel

Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Xuli Shen , Hua Cai , Dingding Yu , Weilin Shen , Qing Xu , Xiangyang Xue

Generalization to unseen real-world scenarios for robot manipulation requires exposure to diverse datasets during training. However, collecting large real-world datasets is intractable due to high operational costs. For robot learning to…

机器人学 · 计算机科学 2024-09-04 Zoey Chen , Zhao Mandi , Homanga Bharadhwaj , Mohit Sharma , Shuran Song , Abhishek Gupta , Vikash Kumar

Understating and controlling generative models' latent space is a complex task. In this paper, we propose a novel method for learning to control any desired attribute in a pre-trained GAN's latent space, for the purpose of editing…

计算机视觉与模式识别 · 计算机科学 2021-11-18 Nir Diamant , Nitsan Sandor , Alex M Bronstein

This study addresses the challenge that generative models struggle to balance flexibility, stability, and controllability in complex interactive scenarios. It proposes a controllable generation framework for dynamic interactive content…

人机交互 · 计算机科学 2026-02-27 Rui Liu

Controllability plays a crucial role in video generation, as it allows users to create and edit content more precisely. Existing models, however, lack control of camera pose that serves as a cinematic language to express deeper narrative…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Hao He , Yinghao Xu , Yuwei Guo , Gordon Wetzstein , Bo Dai , Hongsheng Li , Ceyuan Yang

In social robotics, endowing humanoid robots with the ability to generate bodily expressions of affect can improve human-robot interaction and collaboration, since humans attribute, and perhaps subconsciously anticipate, such traces to…

机器人学 · 计算机科学 2022-05-03 Mina Marmpena , Fernando Garcia , Angelica Lim , Nikolas Hemion , Thomas Wennekers

We propose a shared semantic map architecture to construct and configure Model Predictive Controllers (MPC) dynamically, that solve navigation problems for multiple robotic agents sharing parts of the same environment. The navigation task…

机器人学 · 计算机科学 2024-10-24 K. de Vos , E. Torta , H. Bruyninckx , C. A. Lopez Martinez , M. J. G. van de Molengraft

Semantic segmentation is a crucial task for robot navigation and safety. However, it requires huge amounts of pixelwise annotations to yield accurate results. While recent progress in computer vision algorithms has been heavily boosted by…

计算机视觉与模式识别 · 计算机科学 2019-10-23 Alina Marcu , Dragos Costea , Vlad Licaret , Marius Leordeanu

The field of Text-to-Speech has experienced huge improvements last years benefiting from deep learning techniques. Producing realistic speech becomes possible now. As a consequence, the research on the control of the expressiveness,…

计算与语言 · 计算机科学 2019-03-28 Noé Tits , Fengna Wang , Kevin El Haddad , Vincent Pagel , Thierry Dutoit