中文
相关论文

相关论文: DINOv3-Diffusion Policy: Self-Supervised Large Vis…

200 篇论文

We are witnessing a modeling shift from CNN to Transformers in computer vision. In this work, we present a self-supervised learning approach called MoBY, with Vision Transformers as its backbone architecture. The approach basically has no…

计算机视觉与模式识别 · 计算机科学 2021-05-12 Zhenda Xie , Yutong Lin , Zhuliang Yao , Zheng Zhang , Qi Dai , Yue Cao , Han Hu

Many recent Vision-Language-Action models employ diffusion or flow-matching backbones with hundreds of millions of parameters for action generation. However, unlike image synthesis where the output spans millions of diverse pixels, a…

机器人学 · 计算机科学 2026-03-17 Jian Zhou , Sihao Lin , Shuai Fu , Zerui Li , Gengze Zhou , Qi WU

In recent years, large-scale pre-trained diffusion models have demonstrated their outstanding capabilities in image and video generation tasks. However, existing models tend to produce visual objects commonly found in the training dataset,…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Changgu Chen , Libing Yang , Xiaoyan Yang , Lianggangxu Chen , Gaoqi He , CHangbo Wang , Yang Li

If a robot masters folding a kitchen towel, we would expect it to master folding a large beach towel. However, existing policy learning methods that rely on data augmentation still don't guarantee such generalization. Our insight is to add…

机器人学 · 计算机科学 2024-07-03 Jingyun Yang , Congyue Deng , Jimmy Wu , Rika Antonova , Leonidas Guibas , Jeannette Bohg

Mental rotation is a key test of spatial reasoning in humans and has been central to understanding how perception supports cognition. Despite the success of modern vision transformers, it is still unclear how well these models develop…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Sebastian Ray Mason , Anders Gjølbye , Phillip Chavarria Højbjerg , Lenka Tětková , Lars Kai Hansen

The remote sensing (RS) domain suffers from a lack of densely labeled datasets, which are costly to obtain. Thus, models that can segment RS imagery well without supervised fine-tuning are valuable, but existing solutions fall behind…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Ryan Faulkenberry , Saurabh Prasad

Face recognition systems are increasingly used in biometric security for convenience and effectiveness. However, they remain vulnerable to spoofing attacks, where attackers use photos, videos, or masks to impersonate legitimate users. This…

计算机视觉与模式识别 · 计算机科学 2024-10-23 Arman Keresh , Pakizar Shamoi

Recently, 3D vision-based diffusion policies have shown strong capability in learning complex robotic manipulation skills. However, a common architectural mismatch exists in these models: a tiny yet efficient point-cloud encoder is often…

机器人学 · 计算机科学 2026-02-02 Jinhao Zhang , Zhexuan Zhou , Huizhe Li , Yichen Lai , Wenlong Xia , Haoming Song , Youmin Gong , Jie Mei

Recent large vision-language-action models pretrained on diverse robot datasets have demonstrated the potential for generalizing to new environments with a few in-domain data. However, those approaches usually predict individual discretized…

机器人学 · 计算机科学 2025-03-25 Zhi Hou , Tianyi Zhang , Yuwen Xiong , Hengjun Pu , Chengyang Zhao , Ronglei Tong , Yu Qiao , Jifeng Dai , Yuntao Chen

Frozen pretrained image representations are widely used for transfer learning: a backbone is kept fixed, feature vectors are extracted, and a lightweight classifier is trained on top. This pipeline usually feeds the full feature vector to…

机器学习 · 计算机科学 2026-05-12 Indar Kumar , Girish Karhana , Sai Krishna Jasti , Ankit Hemant Lade

Video models have recently been applied with success to problems in content generation, novel view synthesis, and, more broadly, world simulation. Many applications in generation and transfer rely on conditioning these models, typically…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Edoardo A. Dominici , Thomas Deixelberger , Konstantinos Vardis , Markus Steinberger

In this paper, we introduce a self-supervised deep SLAM method that robustly operates in dynamic scenes while accurately identifying dynamic components. Our method leverages a dual-flow representation for static flow and dynamic flow,…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Xingyuan Yu , Weicai Ye , Xiyue Guo , Yuhang Ming , Jinyu Li , Hujun Bao , Zhaopeng Cui , Guofeng Zhang

Diffusion-based visuomotor policies excel at modeling action distributions but are inference-inefficient, since recursively denoising from noise to policy requires many steps and heavy UNet backbones, which hinders deployment on…

机器人学 · 计算机科学 2026-02-16 Zhihao Chen , Yiyuan Ge , Ziyang Wang

The DINO family of self-supervised vision models has shown remarkable transferability, yet effectively adapting their representations for segmentation remains challenging. Existing approaches often rely on heavy decoders with multi-scale…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Sicheng Yang , Hongqiu Wang , Zhaohu Xing , Sixiang Chen , Lei Zhu

Visual imitation learning is effective for robots to learn versatile tasks. However, many existing methods rely on behavior cloning with supervised historical trajectories, limiting their 3D spatial and 4D spatiotemporal awareness.…

机器人学 · 计算机科学 2025-07-15 Zhenyang Liu , Yikai Wang , Kuanning Wang , Longfei Liang , Xiangyang Xue , Yanwei Fu

Recently, the diffusion model has emerged as a powerful generative technique for robotic policy learning, capable of modeling multi-mode action distributions. Leveraging its capability for end-to-end autonomous driving is a promising…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Bencheng Liao , Shaoyu Chen , Haoran Yin , Bo Jiang , Cheng Wang , Sixu Yan , Xinbang Zhang , Xiangyu Li , Ying Zhang , Qian Zhang , Xinggang Wang

Many robotic systems, such as mobile manipulators or quadrotors, cannot be equipped with high-end GPUs due to space, weight, and power constraints. These constraints prevent these systems from leveraging recent developments in visuomotor…

机器人学 · 计算机科学 2024-07-02 Aaditya Prasad , Kevin Lin , Jimmy Wu , Linqi Zhou , Jeannette Bohg

Needle picking is a challenging manipulation task in robot-assisted surgery due to the characteristics of small slender shapes of needles, needles' variations in shapes and sizes, and demands for millimeter-level control. Prior works,…

机器人学 · 计算机科学 2023-07-27 Hongbin Lin , Bin Li , Xiangyu Chu , Qi Dou , Yunhui Liu , Kwok Wai Samuel Au

Building a robust perception module is crucial for visuomotor policy learning. While recent methods incorporate pre-trained 2D foundation models into robotic perception modules to leverage their strong semantic understanding, they struggle…

机器人学 · 计算机科学 2025-07-14 Wenbo Cui , Chengyang Zhao , Yuhui Chen , Haoran Li , Zhizheng Zhang , Dongbin Zhao , He Wang

Training AI models to understand images without costly labeled data remains a challenge. We combine two techniques--DINO (teacher-student learning) and Barlow Twins (redundancy reduction)--to create a model that learns better with fewer…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Michael Podsiadly , Brendon K Lay