English
Related papers

Related papers: Nemotron 3 Nano Omni: Efficient and Open Multimoda…

200 papers

Audio and video are two most common modalities in the mainstream media platforms, e.g., YouTube. To learn from multimodal videos effectively, in this work, we propose a novel audio-video recognition approach termed audio video Transformer,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Wentao Zhu

Any general artificial intelligence system must be able to interpret, operate on, and produce data in a multi-modal latent space that can represent audio, imagery, text, and more. In the last decade, deep neural networks have seen…

Machine Learning · Computer Science 2021-10-12 Sarah Di , Robin Yu , Amol Kapoor

Utilizing the sensor characteristics of the audio, visible camera, and thermal camera, the robustness of person recognition can be enhanced. Existing multimodal person recognition frameworks are primarily formulated assuming that multimodal…

Multimedia · Computer Science 2022-10-25 Vijay John , Yasutomo Kawanishi

Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonomous driving. However, existing vision-language models are…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Minghui Hou , Wei-Hsing Huang , Shaofeng Liang , Daizong Liu , Tai-Hao Wen , Gang Wang , Runwei Guan , Weiping Ding

Autonomous vehicles (AVs) are poised to redefine transportation by enhancing road safety, minimizing human error, and optimizing traffic efficiency. The success of AVs depends on their ability to interpret complex, dynamic environments…

Multimedia · Computer Science 2025-07-11 Abolfazl Zarghani , Amirhossein Ebrahimi , Amir Malekesfandiari

This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-translation.…

Computation and Language · Computer Science 2024-01-17 Eliya Nachmani , Alon Levkovitch , Yifan Ding , Chulayuth Asawaroengchai , Heiga Zen , Michelle Tadmor Ramanovich

Deploying high-quality automatic speech recognition (ASR) on edge devices requires models that jointly optimize accuracy, latency, and memory footprint while operating entirely on CPU without GPU acceleration. We conduct a systematic…

Artificial Intelligence · Computer Science 2026-04-21 Nenad Banfic , David Fan , Kunal Vaishnavi , Sam Kemp , Sunghoon Choi , Rui Ren , Sayan Shaw , Meng Tang

Modern information systems often involve different types of items, e.g., a text query, an image, a video clip, or an audio segment. This motivates omni-modal embedding models that map heterogeneous modalities into a shared space for direct…

Computation and Language · Computer Science 2026-01-12 Haonan Chen , Sicheng Gao , Radu Timofte , Tetsuya Sakai , Zhicheng Dou

Training a family of large language models targeting multiple scales and deployment objectives is prohibitively expensive, requiring separate training runs for each different size. Recent work on model compression through pruning and…

We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art performance across a diverse range of audio benchmarks, while…

Sound · Computer Science 2025-07-18 Alexander H. Liu , Andy Ehrenberg , Andy Lo , Clément Denoix , Corentin Barreau , Guillaume Lample , Jean-Malo Delignon , Khyathi Raghavi Chandu , Patrick von Platen , Pavankumar Reddy Muddireddy , Sanchit Gandhi , Soham Ghosh , Srijan Mishra , Thomas Foubert , Abhinav Rastogi , Adam Yang , Albert Q. Jiang , Alexandre Sablayrolles , Amélie Héliou , Amélie Martin , Anmol Agarwal , Antoine Roux , Arthur Darcet , Arthur Mensch , Baptiste Bout , Baptiste Rozière , Baudouin De Monicault , Chris Bamford , Christian Wallenwein , Christophe Renaudin , Clémence Lanfranchi , Darius Dabert , Devendra Singh Chaplot , Devon Mizelle , Diego de las Casas , Elliot Chane-Sane , Emilien Fugier , Emma Bou Hanna , Gabrielle Berrada , Gauthier Delerce , Gauthier Guinet , Georgii Novikov , Guillaume Martin , Himanshu Jaju , Jan Ludziejewski , Jason Rute , Jean-Hadrien Chabran , Jessica Chudnovsky , Joachim Studnia , Joep Barmentlo , Jonas Amar , Josselin Somerville Roberts , Julien Denize , Karan Saxena , Karmesh Yadav , Kartik Khandelwal , Kush Jain , Lélio Renard Lavaud , Léonard Blier , Lingxiao Zhao , Louis Martin , Lucile Saulnier , Luyu Gao , Marie Pellat , Mathilde Guillaumin , Mathis Felardos , Matthieu Dinot , Maxime Darrin , Maximilian Augustin , Mickaël Seznec , Neha Gupta , Nikhil Raghuraman , Olivier Duchenne , Patricia Wang , Patryk Saffer , Paul Jacob , Paul Wambergue , Paula Kurylowicz , Philomène Chagniot , Pierre Stock , Pravesh Agrawal , Rémi Delacourt , Romain Sauvestre , Roman Soletskyi , Sagar Vaze , Sandeep Subramanian , Saurabh Garg , Shashwat Dalal , Siddharth Gandhi , Sumukh Aithal , Szymon Antoniak , Teven Le Scao , Thibault Schueller , Thibaut Lavril , Thomas Robert , Thomas Wang , Timothée Lacroix , Tom Bewley , Valeriia Nemychnikova , Victor Paltz , Virgile Richard , Wen-Ding Li , William Marshall , Xuanyu Zhang , Yihan Wan , Yunhao Tang

We present a new pre-training strategy called M$^{3}$3D ($\underline{M}$ulti-$\underline{M}$odal $\underline{M}$asked $\underline{3D}$) built based on Multi-modal masked autoencoders that can leverage 3D priors and learned cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Muhammad Abdullah Jamal , Omid Mohareri

Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However, their ability to jointly reason over omni-modal (visual, audio, and textual) signals in long and…

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Given the foundational role of multimodal joint representation in understanding and generation pipelines, high-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Zehan Wang , Ziang Zhang , Hang Zhang , Luping Liu , Rongjie Huang , Xize Cheng , Hengshuang Zhao , Zhou Zhao

In this paper, we introduce a new embedding model called M3-Embedding, which is distinguished for its versatility in \textit{Multi-Linguality}, \textit{Multi-Functionality}, and \textit{Multi-Granularity}. It provides a uniform support for…

Computation and Language · Computer Science 2025-12-15 Jianlv Chen , Shitao Xiao , Peitian Zhang , Kun Luo , Defu Lian , Zheng Liu

We propose a self-supervised shared encoder model that achieves strong results on several visual, language and multimodal benchmarks while being data, memory and run-time efficient. We make three key contributions. First, in contrast to…

Computer Vision and Pattern Recognition · Computer Science 2023-04-13 Rakesh Chada , Zhaoheng Zheng , Pradeep Natarajan

Multimodal models excel in English, supported by abundant image-text and audio-text data, but performance drops sharply for other languages due to limited multilingual multimodal resources. Existing solutions rely on machine translation,…

Machine Learning · Computer Science 2026-01-22 Piyush Singh Pasi

We present BlueLM-2.5-3B, a compact and unified dense Multimodal Large Language Model (MLLM) designed for efficient edge-device deployment, offering strong general-purpose and reasoning capabilities. To the best of our knowledge, this is…

As inference-time scaling becomes critical for enhanced reasoning capabilities, it is increasingly becoming important to build models that are efficient to infer. We introduce Nemotron-H, a family of 8B and 56B/47B hybrid Mamba-Transformer…

Computation and Language · Computer Science 2025-09-09 NVIDIA , : , Aaron Blakeman , Aarti Basant , Abhinav Khattar , Adithya Renduchintala , Akhiad Bercovich , Aleksander Ficek , Alexis Bjorlin , Ali Taghibakhshi , Amala Sanjay Deshmukh , Ameya Sunil Mahabaleshwarkar , Andrew Tao , Anna Shors , Ashwath Aithal , Ashwin Poojary , Ayush Dattagupta , Balaram Buddharaju , Bobby Chen , Boris Ginsburg , Boxin Wang , Brandon Norick , Brian Butterfield , Bryan Catanzaro , Carlo del Mundo , Chengyu Dong , Christine Harvey , Christopher Parisien , Dan Su , Daniel Korzekwa , Danny Yin , Daria Gitman , David Mosallanezhad , Deepak Narayanan , Denys Fridman , Dima Rekesh , Ding Ma , Dmytro Pykhtar , Dong Ahn , Duncan Riach , Dusan Stosic , Eileen Long , Elad Segal , Ellie Evans , Eric Chung , Erick Galinkin , Evelina Bakhturina , Ewa Dobrowolska , Fei Jia , Fuxiao Liu , Gargi Prasad , Gerald Shen , Guilin Liu , Guo Chen , Haifeng Qian , Helen Ngo , Hongbin Liu , Hui Li , Igor Gitman , Ilia Karmanov , Ivan Moshkov , Izik Golan , Jan Kautz , Jane Polak Scowcroft , Jared Casper , Jarno Seppanen , Jason Lu , Jason Sewall , Jiaqi Zeng , Jiaxuan You , Jimmy Zhang , Jing Zhang , Jining Huang , Jinze Xue , Jocelyn Huang , Joey Conway , John Kamalu , Jon Barker , Jonathan Cohen , Joseph Jennings , Jupinder Parmar , Karan Sapra , Kari Briski , Kateryna Chumachenko , Katherine Luna , Keshav Santhanam , Kezhi Kong , Kirthi Sivamani , Krzysztof Pawelec , Kumar Anik , Kunlun Li , Lawrence McAfee , Leon Derczynski , Lindsey Pavao , Luis Vega , Lukas Voegtle , Maciej Bala , Maer Rodrigues de Melo , Makesh Narsimhan Sreedhar , Marcin Chochowski , Markus Kliegl , Marta Stepniewska-Dziubinska , Matthieu Le , Matvei Novikov , Mehrzad Samadi , Michael Andersch , Michael Evans , Miguel Martinez , Mike Chrzanowski , Mike Ranzinger , Mikolaj Blaz , Misha Smelyanskiy , Mohamed Fawzy , Mohammad Shoeybi , Mostofa Patwary , Nayeon Lee , Nima Tajbakhsh , Ning Xu , Oleg Rybakov , Oleksii Kuchaiev , Olivier Delalleau , Osvald Nitski , Parth Chadha , Pasha Shamis , Paulius Micikevicius , Pavlo Molchanov , Peter Dykas , Philipp Fischer , Pierre-Yves Aquilanti , Piotr Bialecki , Prasoon Varshney , Pritam Gundecha , Przemek Tredak , Rabeeh Karimi , Rahul Kandu , Ran El-Yaniv , Raviraj Joshi , Roger Waleffe , Ruoxi Zhang , Sabrina Kavanaugh , Sahil Jain , Samuel Kriman , Sangkug Lym , Sanjeev Satheesh , Saurav Muralidharan , Sean Narenthiran , Selvaraj Anandaraj , Seonmyeong Bak , Sergey Kashirsky , Seungju Han , Shantanu Acharya , Shaona Ghosh , Sharath Turuvekere Sreenivas , Sharon Clay , Shelby Thomas , Shrimai Prabhumoye , Shubham Pachori , Shubham Toshniwal , Shyamala Prayaga , Siddhartha Jain , Sirshak Das , Slawek Kierat , Somshubra Majumdar , Song Han , Soumye Singhal , Sriharsha Niverty , Stefania Alborghetti , Suseella Panguluri , Swetha Bhendigeri , Syeda Nahida Akter , Szymon Migacz , Tal Shiri , Terry Kong , Timo Roman , Tomer Ronen , Trisha Saar , Tugrul Konuk , Tuomas Rintamaki , Tyler Poon , Ushnish De , Vahid Noroozi , Varun Singh , Vijay Korthikanti , Vitaly Kurin , Wasi Uddin Ahmad , Wei Du , Wei Ping , Wenliang Dai , Wonmin Byeon , Xiaowei Ren , Yao Xu , Yejin Choi , Yian Zhang , Ying Lin , Yoshi Suhara , Zhiding Yu , Zhiqi Li , Zhiyu Li , Zhongbo Zhu , Zhuolin Yang , Zijia Chen

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Gaojie Lin , Jianwen Jiang , Jiaqi Yang , Zerong Zheng , Chao Liang

Recent advances in Large Multimodal Models (LMM) have made it possible for various applications in human-machine interactions. However, developing LMMs that can comprehend, reason, and plan in complex and diverse 3D environments remains a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Sijin Chen , Xin Chen , Chi Zhang , Mingsheng Li , Gang Yu , Hao Fei , Hongyuan Zhu , Jiayuan Fan , Tao Chen
‹ Prev 1 4 5 6 7 8 10 Next ›