English
Related papers

Related papers: COMPASS: Contrastive Multimodal Pretraining for Au…

200 papers

In clinical applications, the utility of segmentation models is often based on the accuracy of derived downstream metrics such as organ size, rather than by the pixel-level accuracy of the segmentation masks themselves. Thus, uncertainty…

Image and Video Processing · Electrical Eng. & Systems 2026-03-03 Matt Y. Cheung , Ashok Veeraraghavan , Guha Balakrishnan

In-context learning allows adapting a model to new tasks given a task description at test time. In this paper, we present IMProv - a generative model that is able to in-context learn visual tasks from multimodal prompts. Given a textual…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Jiarui Xu , Yossi Gandelsman , Amir Bar , Jianwei Yang , Jianfeng Gao , Trevor Darrell , Xiaolong Wang

Developing agents that can perform complex control tasks from high dimensional observations such as pixels is challenging due to difficulties in learning dynamics efficiently. In this work, we propose to learn forward and inverse dynamics…

Robotics · Computer Science 2020-10-26 Jianren Wang , Yujie Lu , Hang Zhao

Self-supervised learning has made substantial strides in image processing, while visual pre-training for autonomous driving is still in its infancy. Existing methods often focus on learning geometric scene information while neglecting…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Shaoqing Xu , Fang Li , Shengyin Jiang , Ziying Song , Li Liu , Zhi-xin Yang

The technology for autonomous vehicles is close to replacing human drivers by artificial systems endowed with high-level decision-making capabilities. In this regard, systems must learn about the usual vehicle's behavior to predict imminent…

Image and Video Processing · Electrical Eng. & Systems 2020-04-22 Mahdyar Ravanbakhsh , Mohamad Baydoun , Damian Campo , Pablo Marin , David Martin , Lucio Marcenaro , andCarlo Regazzoni

The objective of constrained motion planning is to connect start and goal configurations while satisfying task-specific constraints. Motion planning becomes inefficient or infeasible when the configurations lie in disconnected regions,…

Robotics · Computer Science 2026-03-27 Suhyun Jeon , Yumin Lim , Woo-Jeong Baek , Hyeonseo Kim , Suhan Park , Jaeheung Park

The goal of system identification is to learn about underlying physics dynamics behind the time-series data. To model the probabilistic and nonparametric dynamics model, Gaussian process (GP) have been widely used; GP can estimate the…

Machine Learning · Statistics 2018-11-22 Young-Jin Park , Han-Lim Choi

Due to the limited scale and quality of video-text training corpus, most vision-language foundation models employ image-text datasets for pretraining and primarily focus on modeling visually semantic representations while disregarding…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Sihan Chen , Xingjian He , Handong Li , Xiaojie Jin , Jiashi Feng , Jing Liu

Multimodal stock trading volume movement prediction with stock-related news is one of the fundamental problems in the financial area. Existing multimodal works that train models from scratch face the problem of lacking universal knowledge…

Computation and Language · Computer Science 2023-09-12 Ruibo Chen , Zhiyuan Zhang , Yi Liu , Ruihan Bao , Keiko Harimoto , Xu Sun

Many healthcare applications are inherently multimodal, involving several physiological signals. As sensors for these signals become more common, improving machine learning methods for multimodal healthcare data is crucial. Pretraining…

Machine Learning · Computer Science 2024-10-23 Ching Fang , Christopher Sandino , Behrooz Mahasseni , Juri Minxha , Hadi Pouransari , Erdrin Azemi , Ali Moin , Ellen Zippi

While the automotive industry is currently facing a contest among different communication technologies and paradigms about predominance in the connected vehicles sector, the diversity of the various application requirements makes it…

Networking and Internet Architecture · Computer Science 2019-07-09 Benjamin Sliwa , Johannes Pillmann , Maximilian Klaß , Christian Wietfeld

Goal-oriented navigation presents a fundamental challenge for autonomous systems, requiring agents to navigate complex environments to reach designated targets. This survey offers a comprehensive analysis of multimodal navigation approaches…

Robotics · Computer Science 2025-04-23 I-Tak Ieong , Hao Tang

Large-scale multimodal representation learning successfully optimizes for zero-shot transfer at test time. Yet the standard pretraining paradigm (contrastive learning on large amounts of image-text data) does not explicitly encourage…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Karsten Roth , Zeynep Akata , Dima Damen , Ivana Balažević , Olivier J. Hénaff

We propose a unified framework for adaptive routing in multitask, multimodal prediction settings where data heterogeneity and task interactions vary across samples. Motivated by applications in psychotherapy where structured assessments and…

Multimodal pretraining is an effective strategy for the trinity of goals of representation learning in autonomous robots: 1) extracting both local and global task progressions; 2) enforcing temporal consistency of visual representation; 3)…

Traffic scene recognition, which requires various visual classification tasks, is a critical ingredient in autonomous vehicles. However, most existing approaches treat each relevant task independently from one another, never considering the…

Computer Vision and Pattern Recognition · Computer Science 2020-04-06 Younkwan Lee , Jihyo Jeon , Jongmin Yu , Moongu Jeon

Text-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this common failure: 1) the ambiguous nature of data concerning…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Gaoyang Zhang , Bingtao Fu , Qingnan Fan , Qi Zhang , Runxing Liu , Hong Gu , Huaqi Zhang , Xinguo Liu

Learning contextual and spatial environmental representations enhances autonomous vehicle's hazard anticipation and decision-making in complex scenarios. Recent perception systems enhance spatial understanding with sensor fusion but often…

Robotics · Computer Science 2024-01-18 Shoaib Azam , Farzeen Munir , Ville Kyrki , Moongu Jeon , Witold Pedrycz

Critical for the coexistence of humans and robots in dynamic environments is the capability for agents to understand each other's actions, and anticipate their movements. This paper presents Stochastic Process Anticipatory Navigation…

Robotics · Computer Science 2020-11-13 Weiming Zhi , Tin Lai , Lionel Ott , Fabio Ramos

To equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts visual data into nodes to represent entities and edges to capture temporal relations.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Thong Thanh Nguyen , Xiaobao Wu , Yi Bin , Cong-Duy T Nguyen , See-Kiong Ng , Anh Tuan Luu
‹ Prev 1 4 5 6 7 8 10 Next ›