中文
相关论文

相关论文: Viewpoint-Agnostic Manipulation Policies with Stra…

200 篇论文

Vision-based policies are widely applied in robotics for tasks such as manipulation and locomotion. On lightweight mobile robots, however, they face a trilemma of limited scene transferability, restricted onboard computation resources, and…

机器人学 · 计算机科学 2026-03-24 Kai Li , Shiyu Zhao

Many computer vision tasks rely on labeled data. Rapid progress in generative modeling has led to the ability to synthesize photorealistic images. However, controlling specific aspects of the generation process such that the data can be…

计算机视觉与模式识别 · 计算机科学 2020-10-26 Yufeng Zheng , Seonwook Park , Xucong Zhang , Shalini De Mello , Otmar Hilliges

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlusions and during cross-task transfer. To address these…

Diffusion-based image synthesis has attracted extensive attention recently. In particular, ControlNet that uses image-based prompts exhibits powerful capability in image tasks such as canny edge detection and generates images well aligned…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Junjie Yang , Jinze Zhao , Peihao Wang , Zhangyang Wang , Yingbin Liang

Real-world robotic manipulation demands visuomotor policies capable of robust spatial scene understanding and strong generalization across diverse camera viewpoints. While recent advances in 3D-aware visual representations have shown…

机器人学 · 计算机科学 2026-02-02 Di Zhang , Weicheng Duan , Dasen Gu , Hongye Lu , Hai Zhang , Hang Yu , Junqiao Zhao , Guang Chen

Novel-view synthesis through diffusion models has demonstrated remarkable potential for generating diverse and high-quality images. Yet, the independent process of image generation in these prevailing methods leads to challenges in…

计算机视觉与模式识别 · 计算机科学 2024-03-01 Xianghui Yang , Yan Zuo , Sameera Ramasinghe , Loris Bazzani , Gil Avraham , Anton van den Hengel

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT) tend to treat multi-view features equally and directly concatenate them for policy learning. However, it…

机器人学 · 计算机科学 2025-07-01 Zihan Lan , Weixin Mao , Haosheng Li , Le Wang , Tiancai Wang , Haoqiang Fan , Osamu Yoshie

Learning robust visuomotor policies for robotic manipulation remains a challenge in real-world settings, where visual distractors can significantly degrade performance and safety. In this work, we propose an effective and scalable…

机器人学 · 计算机科学 2025-12-01 Sajjad Pakdamansavoji , Mozhgan Pourkeshavarz , Adam Sigal , Zhiyuan Li , Rui Heng Yang , Amir Rasouli

Since Transformers are introduced into vision architectures, their quadratic complexity has always been a significant issue that many research efforts aim to address. A representative approach involves grouping tokens, performing…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Qihang Fan , Yuang Ai , Huaibo Huang , Ran He

Decentralized machine learning often relies on outsourcing computations, such as gradient evaluations, to untrusted worker nodes. Existing robust aggregation methods can mitigate malicious behavior under honest-majority assumptions, but may…

机器学习 · 计算机科学 2026-05-11 Hanzaleh Akbari Nodehi , Parsa Moradi , Soheil Mohajer , Mohammad Ali Maddah-Ali

In agricultural automation, inherent occlusion presents a major challenge for robotic harvesting. We propose a novel imitation learning-based viewpoint planning approach to actively adjust camera viewpoint and capture unobstructed images of…

机器人学 · 计算机科学 2025-03-14 Lun Li , Hamidreza Kasaei

Vision-Language-Action (VLA) models have demonstrated strong performance across a wide range of robotic manipulation tasks. Despite the success, extending large pretrained Vision-Language Models (VLMs) to the action space can induce…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Yiye Chen , Yanan Jian , Xiaoyi Dong , Shuxin Cao , Jing Wu , Patricio Vela , Benjamin E. Lundell , Dongdong Chen

Prevailing 2D-centric visuomotor policies exhibit a pronounced deficiency in novel view generalization, as their reliance on static observations hinders consistent action mapping across unseen views. In response, we introduce GenSplat, a…

机器人学 · 计算机科学 2026-04-01 Sen Wang , Huaiyi Dong , Jingyi Tian , Jiayi Li , Zhuo Yang , Tongtong Cao , Anlin Chen , Shuang Wu , Le Wang , Sanping Zhou

Estimating camera pose in dynamic environments is a critical challenge, as most visual SLAM and SfM methods assume static scenes. While recent dynamic-aware methods exist, they are often not unified: semantic-based approaches are brittle,…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Jianhao Zheng , Liyuan Zhu , Zihan Zhu , Iro Armeni

We propose Camera Splatting, a novel view optimization framework for novel view synthesis. Each camera is modeled as a 3D Gaussian, referred to as a camera splat, and virtual cameras, termed point cameras, are placed at 3D points sampled…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Gahye Lee , Hyomin Kim , Gwangjin Ju , Jooeun Son , Hyejeong Yoon , Seungyong Lee

Optimal viewpoint prediction is an essential task in many computer graphics applications. Unfortunately, common viewpoint qualities suffer from two major drawbacks: dependency on clean surface meshes, which are not always available, and the…

图形学 · 计算机科学 2021-02-10 Michael Schelling , Pedro Hermosilla , Pere-Pau Vazquez , Timo Ropinski

Traditional point-based image editing methods rely on iterative latent optimization or geometric transformations, which are either inefficient in their processing or fail to capture the semantic relationships within the image. These methods…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Biao Yang , Muqi Huang , Yuhui Zhang , Yun Xiong , Kun Zhou , Xi Chen , Shiyang Zhou , Huishuai Bao , Chuan Li , Feng Shi , Hualei Liu

Planning contact interactions is one of the core challenges of many robotic tasks. Optimizing contact locations while taking dynamics into account is computationally costly and, in environments that are only partially observable, executing…

机器人学 · 计算机科学 2020-04-20 Alina Kloss , Maria Bauza , Jiajun Wu , Joshua B. Tenenbaum , Alberto Rodriguez , Jeannette Bohg

Class-incremental learning requires a learning system to continually learn knowledge of new classes and meanwhile try to preserve previously learned knowledge of old classes. As current state-of-the-art methods based on Vision-Language…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Jiantao Tan , Peixian Ma , Tong Yu , Wentao Zhang , Ruixuan Wang

Training vision-language models for image-text alignment typically requires large datasets to achieve robust performance. In low-data scenarios, standard contrastive learning can struggle to align modalities effectively due to overfitting…

计算机视觉与模式识别 · 计算机科学 2025-03-06 Sneh Pillai