English
Related papers

Related papers: Nemotron 3 Nano Omni: Efficient and Open Multimoda…

200 papers

We introduce NVLM 1.0, a family of frontier-class multimodal large language models (LLMs) that achieve state-of-the-art results on vision-language tasks, rivaling the leading proprietary models (e.g., GPT-4o) and open-access models (e.g.,…

Computation and Language · Computer Science 2024-10-24 Wenliang Dai , Nayeon Lee , Boxin Wang , Zhuolin Yang , Zihan Liu , Jon Barker , Tuomas Rintamaki , Mohammad Shoeybi , Bryan Catanzaro , Wei Ping

Recent advances in GPT-4o like multi-modality models have demonstrated remarkable progress for direct speech-to-speech conversation, with real-time speech interaction experience and strong speech understanding ability. However, current…

Sound · Computer Science 2024-12-09 Ze Yuan , Yanqing Liu , Shujie Liu , Sheng Zhao

Last year, multimodal architectures served up a revolution in AI-based approaches and solutions, extending the capabilities of large language models (LLM). We propose an \textit{OmniFusion} model based on a pretrained LLM and adapters for…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Elizaveta Goncharova , Anton Razzhigaev , Matvey Mikhalchuk , Maxim Kurkin , Irina Abdullaeva , Matvey Skripkin , Ivan Oseledets , Denis Dimitrov , Andrey Kuznetsov

Many promising applications of multimodal wearables require continuous sensing and heavy computation, yet users reject such devices due to privacy concerns. This paper shares our experiences building an ear-mounted voice-and-vision wearable…

Human-Computer Interaction · Computer Science 2025-11-26 Yonatan Tussa , Andy Heredia , Nirupam Roy

In this work, we present a conceptually simple yet powerful baseline for the multimodal dialog task, an S3 model, that achieves near state-of-the-art results on two compelling leaderboards: MMMU and AI Journey Contest 2023. The system is…

Computation and Language · Computer Science 2024-06-27 Elisei Rykov , Egor Malkershin , Alexander Panchenko

In an era defined by the explosive growth of data and rapid technological advancements, Multimodal Large Language Models (MLLMs) stand at the forefront of artificial intelligence (AI) systems. Designed to seamlessly integrate diverse data…

Large language models (LLMs) have recently achieved impressive results in speech recognition across multiple modalities, including Auditory Speech Recognition (ASR), Visual Speech Recognition (VSR), and Audio-Visual Speech Recognition…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-28 Umberto Cappellazzo , Xubo Liu , Pingchuan Ma , Stavros Petridis , Maja Pantic

We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audio encoder trained…

Recent advancements in sensor technology and deep learning have led to significant progress in 3D human body reconstruction. However, most existing approaches rely on data from a specific sensor, which can be unreliable due to the inherent…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Anjun Chen , Xiangyu Wang , Zhi Xu , Kun Shi , Yan Qin , Yuchi Huo , Jiming Chen , Qi Ye

We present Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we tokenize inputs and outputs -- images, text, audio, action,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Jiasen Lu , Christopher Clark , Sangho Lee , Zichen Zhang , Savya Khosla , Ryan Marten , Derek Hoiem , Aniruddha Kembhavi

Medical image analysis is essential to clinical diagnosis and treatment, which is increasingly supported by multi-modal large language models (MLLMs). However, previous research has primarily focused on 2D medical images, leaving 3D images…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Fan Bai , Yuxin Du , Tiejun Huang , Max Q. -H. Meng , Bo Zhao

Information comes in diverse modalities. Multimodal native AI models are essential to integrate real-world information and deliver comprehensive understanding. While proprietary multimodal native models exist, their lack of openness imposes…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Dongxu Li , Yudong Liu , Haoning Wu , Yue Wang , Zhiqi Shen , Bowen Qu , Xinyao Niu , Fan Zhou , Chengen Huang , Yanpeng Li , Chongyan Zhu , Xiaoyi Ren , Chao Li , Yifan Ye , Peng Liu , Lihuan Zhang , Hanshu Yan , Guoyin Wang , Bei Chen , Junnan Li

Recently, Multi-modal Large Language Models (MLLMs) have shown remarkable effectiveness for multi-modal tasks due to their abilities to generate and understand cross-modal data. However, processing long sequences of visual tokens extracted…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Haicheng Wang , Zhemeng Yu , Gabriele Spadaro , Chen Ju , Victor Quétu , Shuai Xiao , Enzo Tartaglione

This paper presents OmniDataComposer, an innovative approach for multimodal data fusion and unlimited data generation with an intent to refine and uncomplicate interplay among diverse data modalities. Coming to the core breakthrough, it…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Dongyang Yu , Shihao Wang , Yuan Fang , Wangpeng An

Multimodal emotion recognition identifies human emotions from various data modalities like video, text, and audio. However, we found that this task can be easily affected by noisy information that does not contain useful semantics. To this…

Multimedia · Computer Science 2023-05-05 Yuanyuan Liu , Haoyu Zhang , Yibing Zhan , Zijing Chen , Guanghao Yin , Lin Wei , Zhe Chen

Non-autoregressive (NAR) text-to-speech synthesis relies on length alignment between text sequences and audio representations, constraining naturalness and expressiveness. Existing methods depend on duration modeling or pseudo-alignment…

Efficient audio feature extraction is critical for low-latency, resource-constrained speech recognition. Conventional preprocessing techniques, such as Mel Spectrogram, Perceptual Linear Prediction (PLP), and Learnable Spectrogram, achieve…

Sound · Computer Science 2025-10-28 Akshaya Rajesh , Pavithra Ananthasubramanian , Nagarajan Raghavan , Ankush Kumar

Sensory earables have evolved from basic audio enhancement devices into sophisticated platforms for clinical-grade health monitoring and wellbeing management. This paper introduces OmniBuds, an advanced sensory earable platform integrating…

Large Multimodal Models (LMMs) have demonstrated exceptional comprehension and interpretation capabilities in Autonomous Driving (AD) by incorporating large language models. Despite the advancements, current data-driven AD approaches tend…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Zhijian Huang , Chengjian Feng , Feng Yan , Baihui Xiao , Zequn Jie , Yujie Zhong , Xiaodan Liang , Lin Ma

The advancement of sophisticated artificial intelligence (AI) algorithms has led to a notable increase in energy usage and carbon dioxide emissions, intensifying concerns about climate change. This growing problem has brought the…

Machine Learning · Computer Science 2024-05-22 Hasib-Al Rashid , Tinoosh Mohsenin