English
Related papers

Related papers: Training Video Foundation Models with NVIDIA NeMo

200 papers

Text-to-video diffusion models are notoriously limited in their ability to model temporal aspects such as motion, physics, and dynamic interactions. Existing approaches address this limitation by retraining the model or introducing external…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Ariel Shaulov , Itay Hazan , Lior Wolf , Hila Chefer

Video inpainting is the task of filling a region in a video in a visually convincing manner. It is very challenging due to the high dimensionality of the data and the temporal consistency required for obtaining convincing results. Recently,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Nicolas Cherel , Andrés Almansa , Yann Gousseau , Alasdair Newson

Large Language Models (LLMs) have been widely used in various tasks, motivating us to develop an LLM-based assistant for videos. Instead of training from scratch, we propose a module to transform arbitrary well-trained image-based LLMs into…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Lishuai Gao , Yujie Zhong , Yingsen Zeng , Haoxian Tan , Dengjie Li , Zheng Zhao

We introduce Xmodel-VLM, a cutting-edge multimodal vision language model. It is designed for efficient deployment on consumer GPU servers. Our work directly confronts a pivotal industry issue by grappling with the prohibitive service costs…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Wanting Xu , Yang Liu , Langping He , Xucheng Huang , Ling Jiang

This paper presents an implementation of multilayer feed forward neural networks (NN) to optimize CMOS analog circuits. For modeling and design recently neural network computational modules have got acceptance as an unorthodox and useful…

Neural and Evolutionary Computing · Computer Science 2012-12-13 Mriganka Chakraborty

This paper studies deep network architectures to address the problem of video classification. A multi-stream framework is proposed to fully utilize the rich multimodal information in videos. Specifically, we first train three Convolutional…

Computer Vision and Pattern Recognition · Computer Science 2015-11-12 Zuxuan Wu , Yu-Gang Jiang , Xi Wang , Hao Ye , Xiangyang Xue , Jun Wang

Neural radiance fields (NeRFs) have emerged as an effective method for novel-view synthesis and 3D scene reconstruction. However, conventional training methods require access to all training views during scene optimization. This assumption…

Computer Vision and Pattern Recognition · Computer Science 2023-09-07 Ryan Po , Zhengyang Dong , Alexander W. Bergman , Gordon Wetzstein

Predicting future dynamics is crucial for applications like autonomous driving and robotics, where understanding the environment is key. Existing pixel-level methods are computationally expensive and often focus on irrelevant details. To…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Efstathios Karypidis , Ioannis Kakogeorgiou , Spyros Gidaris , Nikos Komodakis

Continuous normalizing flows (CNFs) learn an ordinary differential equation to transform prior samples into data. Flow matching (FM) has recently emerged as a simulation-free approach for training CNFs by regressing a velocity model towards…

Machine Learning · Statistics 2024-05-28 Tianyu Xie , Yu Zhu , Longlin Yu , Tong Yang , Ziheng Cheng , Shiyue Zhang , Xiangyu Zhang , Cheng Zhang

Graph Foundation Models (GFMs) are emerging as a significant research topic in the graph domain, aiming to develop graph models trained on extensive and diverse data to enhance their applicability across various tasks and domains.…

Machine Learning · Computer Science 2024-06-03 Haitao Mao , Zhikai Chen , Wenzhuo Tang , Jianan Zhao , Yao Ma , Tong Zhao , Neil Shah , Mikhail Galkin , Jiliang Tang

Diffusion models have emerged as a powerful generative method for synthesizing high-quality and diverse set of images. In this paper, we propose a video generation method based on diffusion models, where the effects of motion are modeled in…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Kangfu Mei , Vishal M. Patel

This study introduces an efficient and effective method, MeDM, that utilizes pre-trained image Diffusion Models for video-to-video translation with consistent temporal flow. The proposed framework can render videos from scene position…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Ernie Chu , Tzuhsuan Huang , Shuo-Yen Lin , Jun-Cheng Chen

Convolutional networks are one of the most widely employed architectures in computer vision and machine learning. In order to leverage their ability to learn complex functions, large amounts of data are required for training. Training a…

Computer Vision and Pattern Recognition · Computer Science 2015-06-09 Michael Mathieu , Mikael Henaff , Yann LeCun

Although vision foundation models (VFMs) are increasingly reused for biomedical image analysis, it remains unclear whether the latent representations they provide are general enough to support effective transfer and reuse across…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Caterina Fuster-Barceló , Virginie Uhlmann

Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in foundation models, particularly Vision Language Models (VLMs), have demonstrated remarkable…

Robotics · Computer Science 2025-07-29 Guangyan Chen , Meiling Wang , Te Cui , Yao Mu , Haoyang Lu , Zicai Peng , Mengxiao Hu , Tianxing Zhou , Mengyin Fu , Yi Yang , Yufeng Yue

We present Fashion-VDM, a video diffusion model (VDM) for generating virtual try-on videos. Given an input garment image and person video, our method aims to generate a high-quality try-on video of the person wearing the given garment,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Johanna Karras , Yingwei Li , Nan Liu , Luyang Zhu , Innfarn Yoo , Andreas Lugmayr , Chris Lee , Ira Kemelmacher-Shlizerman

We introduce Nemotron Nano V2 VL, the latest model of the Nemotron vision-language series designed for strong real-world document understanding, long video comprehension, and reasoning tasks. Nemotron Nano V2 VL delivers significant…

Machine Learning · Computer Science 2025-11-10 NVIDIA , : , Amala Sanjay Deshmukh , Kateryna Chumachenko , Tuomas Rintamaki , Matthieu Le , Tyler Poon , Danial Mohseni Taheri , Ilia Karmanov , Guilin Liu , Jarno Seppanen , Guo Chen , Karan Sapra , Zhiding Yu , Adi Renduchintala , Charles Wang , Peter Jin , Arushi Goel , Mike Ranzinger , Lukas Voegtle , Philipp Fischer , Timo Roman , Wei Ping , Boxin Wang , Zhuolin Yang , Nayeon Lee , Shaokun Zhang , Fuxiao Liu , Zhiqi Li , Di Zhang , Greg Heinrich , Hongxu Yin , Song Han , Pavlo Molchanov , Parth Mannan , Yao Xu , Jane Polak Scowcroft , Tom Balough , Subhashree Radhakrishnan , Paris Zhang , Sean Cha , Ratnesh Kumar , Zaid Pervaiz Bhat , Jian Zhang , Darragh Hanley , Pritam Biswas , Jesse Oliver , Kevin Vasques , Roger Waleffe , Duncan Riach , Oluwatobi Olabiyi , Ameya Sunil Mahabaleshwarkar , Bilal Kartal , Pritam Gundecha , Khanh Nguyen , Alexandre Milesi , Eugene Khvedchenia , Ran Zilberstein , Ofri Masad , Natan Bagrov , Nave Assaf , Tomer Asida , Daniel Afrimi , Amit Zuker , Netanel Haber , Zhiyu Cheng , Jingyu Xin , Di Wu , Nik Spirin , Maryam Moosaei , Roman Ageev , Vanshil Atul Shah , Yuting Wu , Daniel Korzekwa , Unnikrishnan Kizhakkemadam Sreekumar , Wanli Jiang , Padmavathy Subramanian , Alejandra Rico , Sandip Bhaskar , Saeid Motiian , Kedi Wu , Annie Surla , Chia-Chih Chen , Hayden Wolff , Matthew Feinberg , Melissa Corpuz , Marek Wawrzos , Eileen Long , Aastha Jhunjhunwala , Paul Hendricks , Farzan Memarian , Benika Hall , Xin-Yu Wang , David Mosallanezhad , Soumye Singhal , Luis Vega , Katherine Cheung , Krzysztof Pawelec , Michael Evans , Katherine Luna , Jie Lou , Erick Galinkin , Akshay Hazare , Kaustubh Purandare , Ann Guan , Anna Warno , Chen Cui , Yoshi Suhara , Shibani Likhite , Seph Mard , Meredith Price , Laya Sleiman , Saori Kaji , Udi Karpas , Kari Briski , Joey Conway , Michael Lightstone , Jan Kautz , Mohammad Shoeybi , Mostofa Patwary , Jonathen Cohen , Oleksii Kuchaiev , Andrew Tao , Bryan Catanzaro

Motion understanding is fundamental to physical reasoning, enabling models to infer dynamics and predict future states. However, state-of-the-art models still struggle on recent motion benchmarks, primarily due to the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Yulu Gan , Ligeng Zhu , Dandan Shan , Baifeng Shi , Hongxu Yin , Boris Ivanovic , Song Han , Trevor Darrell , Jitendra Malik , Marco Pavone , Boyi Li

Recent developments in foundation models, like Large Language Models (LLMs) and Vision-Language Models (VLMs), trained on extensive data, facilitate flexible application across different tasks and modalities. Their impact spans various…

Training LLMs in distributed environments presents significant challenges due to the complexity of model execution, deployment systems, and the vast space of configurable strategies. Although various optimization techniques exist, achieving…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-15 Mingyu Liang , Hiwot Tadese Kassa , Wenyin Fu , Brian Coutinho , Louis Feng , Christina Delimitrou