English
Related papers

Related papers: Jet-Nemotron: Efficient Language Model with Post N…

200 papers

We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while achieving substantially higher inference efficiency on…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Dongyun Zou , Zhuoyang Zhang , Junyu Chen , Wenkun He , Qinhe Peng , Hanrong Ye , Yao Lu , Hongxu Yin , Yu Wang , Song Han , Han Cai

Hybrid attention architectures are becoming an increasingly important paradigm for improving LLM inference efficiency while preserving model quality, making hybrid architecture design a central problem. Existing designs often rely on manual…

Machine Learning · Computer Science 2026-05-21 Weizhe Chen , Miao Zhang , Junpeng Jiang , Yaping Li , Weili Guan , Liqiang Nie

Deep learning methods have become very successful at solving many complex tasks such as image classification and segmentation, speech recognition and machine translation. Nevertheless, manually designing a neural network for a specific…

Computer Vision and Pattern Recognition · Computer Science 2020-11-10 Maria Baldeon Calisto , Susana Lai-Yuen

Parameter-efficient tuning (PET) methods fit pre-trained language models (PLMs) to downstream tasks by either computing a small compressed update for a subset of model parameters, or appending and fine-tuning a small number of new model…

Computation and Language · Computer Science 2023-05-29 Neal Lawton , Anoop Kumar , Govind Thattai , Aram Galstyan , Greg Ver Steeg

We introduce the Llama-Nemotron series of models, an open family of heterogeneous reasoning models that deliver exceptional reasoning capabilities, inference efficiency, and an open license for enterprise use. The family comes in three…

Computation and Language · Computer Science 2025-09-10 Akhiad Bercovich , Itay Levy , Izik Golan , Mohammad Dabbah , Ran El-Yaniv , Omri Puny , Ido Galil , Zach Moshe , Tomer Ronen , Najeeb Nabwani , Ido Shahaf , Oren Tropp , Ehud Karpas , Ran Zilberstein , Jiaqi Zeng , Soumye Singhal , Alexander Bukharin , Yian Zhang , Tugrul Konuk , Gerald Shen , Ameya Sunil Mahabaleshwarkar , Bilal Kartal , Yoshi Suhara , Olivier Delalleau , Zijia Chen , Zhilin Wang , David Mosallanezhad , Adi Renduchintala , Haifeng Qian , Dima Rekesh , Fei Jia , Somshubra Majumdar , Vahid Noroozi , Wasi Uddin Ahmad , Sean Narenthiran , Aleksander Ficek , Mehrzad Samadi , Jocelyn Huang , Siddhartha Jain , Igor Gitman , Ivan Moshkov , Wei Du , Shubham Toshniwal , George Armstrong , Branislav Kisacanin , Matvei Novikov , Daria Gitman , Evelina Bakhturina , Prasoon Varshney , Makesh Narsimhan , Jane Polak Scowcroft , John Kamalu , Dan Su , Kezhi Kong , Markus Kliegl , Rabeeh Karimi Mahabadi , Ying Lin , Sanjeev Satheesh , Jupinder Parmar , Pritam Gundecha , Brandon Norick , Joseph Jennings , Shrimai Prabhumoye , Syeda Nahida Akter , Mostofa Patwary , Abhinav Khattar , Deepak Narayanan , Roger Waleffe , Jimmy Zhang , Bor-Yiing Su , Guyue Huang , Terry Kong , Parth Chadha , Sahil Jain , Christine Harvey , Elad Segal , Jining Huang , Sergey Kashirsky , Robert McQueen , Izzy Putterman , George Lam , Arun Venkatesan , Sherry Wu , Vinh Nguyen , Manoj Kilaru , Andrew Wang , Anna Warno , Abhilash Somasamudramath , Sandip Bhaskar , Maka Dong , Nave Assaf , Shahar Mor , Omer Ullman Argov , Scot Junkin , Oleksandr Romanenko , Pedro Larroy , Monika Katariya , Marco Rovinelli , Viji Balas , Nicholas Edelman , Anahita Bhiwandiwalla , Muthu Subramaniam , Smita Ithape , Karthik Ramamoorthy , Yuting Wu , Suguna Varshini Velury , Omri Almog , Joyjit Daw , Denys Fridman , Erick Galinkin , Michael Evans , Shaona Ghosh , Katherine Luna , Leon Derczynski , Nikki Pope , Eileen Long , Seth Schneider , Guillermo Siman , Tomasz Grzegorzek , Pablo Ribalta , Monika Katariya , Chris Alexiuk , Joey Conway , Trisha Saar , Ann Guan , Krzysztof Pawelec , Shyamala Prayaga , Oleksii Kuchaiev , Boris Ginsburg , Oluwatobi Olabiyi , Kari Briski , Jonathan Cohen , Bryan Catanzaro , Jonah Alben , Yonatan Geifman , Eric Chung

Transformer-based pre-trained language models like BERT and its variants have recently achieved promising performance in various natural language processing (NLP) tasks. However, the conventional paradigm constructs the backbone by purely…

Computation and Language · Computer Science 2022-02-08 Jiahui Gao , Hang Xu , Han Shi , Xiaozhe Ren , Philip L. H. Yu , Xiaodan Liang , Xin Jiang , Zhenguo Li

Self-attention architectures have emerged as a recent advancement for improving the performance of vision tasks. Manual determination of the architecture for self-attention networks relies on the experience of experts and cannot…

Computer Vision and Pattern Recognition · Computer Science 2021-11-16 Yuan Zhou , Haiyang Wang , Shuwei Huo , Boyu Wang

Point cloud architecture design has become a crucial problem for 3D deep learning. Several efforts exist to manually design architectures with high accuracy in point cloud tasks such as classification, segmentation, and detection. Recent…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Guohao Li , Mengmeng Xu , Silvio Giancola , Ali Thabet , Bernard Ghanem

Vision Transformers have enabled recent attention-based Deep Learning (DL) architectures to achieve remarkable results in Computer Vision (CV) tasks. However, due to the extensive computational resources required, these architectures are…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Lotfi Abdelkrim Mecharbat , Hadjer Benmeziane , Hamza Ouarnoughi , Smail Niar

Neural architecture search (NAS) has shown great promise in designing state-of-the-art (SOTA) models that are both accurate and efficient. Recently, two-stage NAS, e.g. BigNAS, decouples the model training and searching process and achieves…

Computer Vision and Pattern Recognition · Computer Science 2021-04-15 Dilin Wang , Meng Li , Chengyue Gong , Vikas Chandra

In this work, we explore multiple neural architectures adapted for the task of automatic post-editing of machine translation output. We focus on neural end-to-end models that combine both inputs $mt$ (raw MT output) and $src$ (source…

Computation and Language · Computer Science 2017-10-03 Marcin Junczys-Dowmunt , Roman Grundkiewicz

Neural Architecture Search (NAS) has proven effective in discovering new Convolutional Neural Network (CNN) architectures, particularly for scenarios with well-defined accuracy optimization goals. However, previous approaches often involve…

Machine Learning · Computer Science 2024-08-28 Ye Qiao , Haocheng Xu , Yifan Zhang , Sitao Huang

We introduce FFN Fusion, an architectural optimization technique that reduces sequential computation in large language models by identifying and exploiting natural opportunities for parallelization. Our key insight is that sequences of…

Mixture of Experts (MoEs) have become a central component of many state-of-the-art open-source and proprietary large language models. Despite their widespread adoption, it remains unclear how close existing MoE architectures are to optimal…

We present Hebatron, a Hebrew-specialized open-weight large language model built on the NVIDIA Nemotron-3 sparse Mixture-of-Experts architecture. Training employs a three-phase easy-to-hard curriculum with continuous anti-forgetting…

Deep neural networks (DNNs) have shown superior performances on various multimodal learning problems. However, it often requires huge efforts to adapt DNNs to individual multimodal tasks by manually engineering unimodal features and…

Computer Vision and Pattern Recognition · Computer Science 2022-01-04 Yihang Yin , Siyu Huang , Xiang Zhang

The traditional design approaches for high-degree-of-freedom metamaterials have been computationally intensive and, in many cases, even intractable due to the vast design space. In this work, we introduce a novel fixed-attention mechanism…

Intent detection and slot filling are two main tasks in natural language understanding and play an essential role in task-oriented dialogue systems. The joint learning of both tasks can improve inference accuracy and is popular in recent…

Computation and Language · Computer Science 2022-05-17 Liang Huang , Senjie Liang , Feiyang Ye , Nan Gao

While modern Transformer-based language models (LMs) have achieved major success in multi-task generalization, they often struggle to capture long-range dependencies within their context window. This work introduces a novel approach using…

Computation and Language · Computer Science 2025-09-23 Alok N. Shah , Khush Gupta , Keshav Ramji , Pratik Chaudhari

Pre-trained language models (PLM), for example BERT or RoBERTa, mark the state-of-the-art for natural language understanding task when fine-tuned on labeled data. However, their large size poses challenges in deploying them for inference in…

Machine Learning · Computer Science 2024-08-27 Aaron Klein , Jacek Golebiowski , Xingchen Ma , Valerio Perrone , Cedric Archambeau
‹ Prev 1 2 3 10 Next ›