English
Related papers

Related papers: SWEET: Sparse World Modeling with Image Editing fo…

200 papers

Text-driven video generation witnesses rapid progress. However, merely using text prompts is not enough to depict the desired subject appearance that accurately aligns with users' intents, especially for customized content creation. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Yuming Jiang , Tianxing Wu , Shuai Yang , Chenyang Si , Dahua Lin , Yu Qiao , Chen Change Loy , Ziwei Liu

We propose a method to train deep networks to decompose videos into 3D geometry (camera and depth), moving objects, and their motions, with no supervision. We build on the idea of view synthesis, which uses classical camera geometry to…

Computer Vision and Pattern Recognition · Computer Science 2021-06-02 Dan Xu , Andrea Vedaldi , Joao F. Henriques

ControlNet has enabled detailed spatial control in text-to-image diffusion models by incorporating additional visual conditions such as depth or edge maps. However, its effectiveness heavily depends on the availability of visual conditions…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Woosung Joung , Daewon Chae , Jinkyu Kim

Recent advances in image editing models have demonstrated remarkable capabilities in executing explicit instructions, such as attribute manipulation, style transfer, and pose synthesis. However, these models often face challenges when…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Wang Lin , Feng Wang , Majun Zhang , Wentao Hu , Tao Jin , Zhou Zhao , Fei Wu , Jingyuan Chen , Alan Yuille , Sucheng Ren

Building generic robotic manipulation systems often requires large amounts of real-world data, which can be dificult to collect. Synthetic data generation offers a promising alternative, but limiting the sim-to-real gap requires significant…

Robotics · Computer Science 2024-11-18 Thomas Lips , Francis wyffels

Diffusion models have demonstrated remarkable generative capabilities in image processing tasks. We propose a Sparse condition Temporal Rewighted Integrated Distribution Estimation guided diffusion model (STRIDE) for sparse-view CT…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Zekun Zhou , Yanru Gong , Liu Shi , Qiegen Liu

Imitation learning has emerged as a promising approach towards building generalist robots. However, scaling imitation learning for large robot foundation models remains challenging due to its reliance on high-quality expert demonstrations.…

Robotics · Computer Science 2025-05-26 Chuning Zhu , Raymond Yu , Siyuan Feng , Benjamin Burchfiel , Paarth Shah , Abhishek Gupta

The use of synthetic (or simulated) data for training machine learning models has grown rapidly in recent years. Synthetic data can often be generated much faster and more cheaply than its real-world counterpart. One challenge of using…

Computer Vision and Pattern Recognition · Computer Science 2022-10-28 Handi Yu , Simiao Ren , Leslie M. Collins , Jordan M. Malof

Most image segmentation algorithms are trained on binary masks formulated as a classification task per pixel. However, in applications such as medical imaging, this "black-and-white" approach is too constraining because the contrast between…

Image and Video Processing · Electrical Eng. & Systems 2020-11-19 Charley Gros , Andreanne Lemay , Julien Cohen-Adad

The effectiveness of scaling up training data in robotic manipulation is still limited. A primary challenge in manipulation is the tasks are diverse, and the trained policy would be confused if the task targets are not specified clearly.…

Robotics · Computer Science 2025-02-12 Zhuoling Li , Liangliang Ren , Jinrong Yang , Yong Zhao , Xiaoyang Wu , Zhenhua Xu , Xiang Bai , Hengshuang Zhao

As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually…

Dynamic scenes that contain both object motion and egomotion are a challenge for monocular visual odometry (VO). Another issue with monocular VO is the scale ambiguity, i.e. these methods cannot estimate scene depth and camera motion in…

Computer Vision and Pattern Recognition · Computer Science 2020-08-31 Hirak J Kashyap , Charless Fowlkes , Jeffrey L Krichmar

Recent works in hand-object reconstruction mainly focus on the single-view and dense multi-view settings. On the one hand, single-view methods can leverage learned shape priors to generalise to unseen objects but are prone to inaccuracies…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Yik Lung Pang , Changjae Oh , Andrea Cavallaro

Sparse autoencoders (SAEs) improve interpretability in multimodal models, but it remains unclear whether SAE features form modular, composable units for reasoning-an assumption underlying many intervention-based steering methods. We test…

Artificial Intelligence · Computer Science 2026-03-27 Yunpeng Zhou

Soft object manipulation tasks in domestic scenes pose a significant challenge for existing robotic skill learning techniques due to their complex dynamics and variable shape characteristics. Since learning new manipulation skills from…

Robotics · Computer Science 2023-09-06 Junjia Liu , Zhihao Li , Wanyu Lin , Sylvain Calinon , Kay Chen Tan , Fei Chen

What if a video generation model could not only imagine a plausible future, but the correct one, accurately reflecting how the world changes with each action? We address this question by presenting the Egocentric World Model (EgoWM), a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Anurag Bagchi , Zhipeng Bao , Homanga Bharadhwaj , Yu-Xiong Wang , Pavel Tokmakov , Martial Hebert

Embodied AI models often employ off the shelf vision backbones like CLIP to encode their visual observations. Although such general purpose representations encode rich syntactic and semantic information about the scene, much of this…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Ainaz Eftekhar , Kuo-Hao Zeng , Jiafei Duan , Ali Farhadi , Ani Kembhavi , Ranjay Krishna

Training robots in simulation requires diverse 3D scenes that reflect the specific challenges of downstream tasks. However, scenes that satisfy strict task requirements, such as high-clutter environments with plausible spatial arrangement,…

Robotics · Computer Science 2025-08-27 Nicholas Pfaff , Hongkai Dai , Sergey Zakharov , Shun Iwase , Russ Tedrake

Text-driven image manipulation remains challenging in training or inference flexibility. Conditional generative models depend heavily on expensive annotated training data. Meanwhile, recent frameworks, which leverage pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Yueming Lyu , Tianwei Lin , Fu Li , Dongliang He , Jing Dong , Tieniu Tan

Do video-text transformers learn to model temporal relationships across frames? Despite their immense capacity and the abundance of multimodal training data, recent work has revealed the strong tendency of video-text models towards…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Yi Li , Kyle Min , Subarna Tripathi , Nuno Vasconcelos
‹ Prev 1 4 5 6 7 8 10 Next ›