English
Related papers

Related papers: Jupiter-N Technical Report

200 papers

In this paper, we introduced our joint team SJTU-NICT 's participation in the WMT 2020 machine translation shared task. In this shared task, we participated in four translation directions of three language pairs: English-Chinese,…

Computation and Language · Computer Science 2020-10-13 Zuchao Li , Hai Zhao , Rui Wang , Kehai Chen , Masao Utiyama , Eiichiro Sumita

Large language models have recently evolved from fluent text generation to advanced reasoning across diverse domains, giving rise to reasoning language models. Among these domains, mathematical reasoning serves as a representative benchmark…

Open-weight LLMs have been released by frontier labs; however, sovereign Large Language Models (for languages other than English) remain low in supply yet high in demand. Training large language models (LLMs) for low-resource languages such…

Computation and Language · Computer Science 2026-02-03 Shaltiel Shmidman , Avi Shmidman , Amir DN Cohen , Moshe Koppel

This paper proposes a hybrid neural network (HNN) model for commonsense reasoning. An HNN consists of two component models, a masked language model and a semantic similarity model, which share a BERT-based contextual encoder but use…

Computation and Language · Computer Science 2019-07-30 Pengcheng He , Xiaodong Liu , Weizhu Chen , Jianfeng Gao

This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limited tri-modal datasets, high computational costs, and complex…

High-quality mathematical reasoning supervision requires diverse reasoning styles, long-form traces, and effective tool integration, capabilities that existing datasets provide only in limited form. Leveraging the multi-mode generation…

We introduce Nemotron-Nano-9B-v2, a hybrid Mamba-Transformer language model designed to increase throughput for reasoning workloads while achieving state-of-the-art accuracy compared to similarly-sized models. Nemotron-Nano-9B-v2 builds on…

Computation and Language · Computer Science 2025-09-03 NVIDIA , : , Aarti Basant , Abhijit Khairnar , Abhijit Paithankar , Abhinav Khattar , Adithya Renduchintala , Aditya Malte , Akhiad Bercovich , Akshay Hazare , Alejandra Rico , Aleksander Ficek , Alex Kondratenko , Alex Shaposhnikov , Alexander Bukharin , Ali Taghibakhshi , Amelia Barton , Ameya Sunil Mahabaleshwarkar , Amy Shen , Andrew Tao , Ann Guan , Anna Shors , Anubhav Mandarwal , Arham Mehta , Arun Venkatesan , Ashton Sharabiani , Ashwath Aithal , Ashwin Poojary , Ayush Dattagupta , Balaram Buddharaju , Banghua Zhu , Barnaby Simkin , Bilal Kartal , Bita Darvish Rouhani , Bobby Chen , Boris Ginsburg , Brandon Norick , Brian Yu , Bryan Catanzaro , Charles Wang , Charlie Truong , Chetan Mungekar , Chintan Patel , Chris Alexiuk , Christian Munley , Christopher Parisien , Dan Su , Daniel Afrimi , Daniel Korzekwa , Daniel Rohrer , Daria Gitman , David Mosallanezhad , Deepak Narayanan , Dima Rekesh , Dina Yared , Dmytro Pykhtar , Dong Ahn , Duncan Riach , Eileen Long , Elliott Ning , Eric Chung , Erick Galinkin , Evelina Bakhturina , Gargi Prasad , Gerald Shen , Haifeng Qian , Haim Elisha , Harsh Sharma , Hayley Ross , Helen Ngo , Herman Sahota , Hexin Wang , Hoo Chang Shin , Hua Huang , Iain Cunningham , Igor Gitman , Ivan Moshkov , Jaehun Jung , Jan Kautz , Jane Polak Scowcroft , Jared Casper , Jian Zhang , Jiaqi Zeng , Jimmy Zhang , Jinze Xue , Jocelyn Huang , Joey Conway , John Kamalu , Jonathan Cohen , Joseph Jennings , Julien Veron Vialard , Junkeun Yi , Jupinder Parmar , Kari Briski , Katherine Cheung , Katherine Luna , Keith Wyss , Keshav Santhanam , Kezhi Kong , Krzysztof Pawelec , Kumar Anik , Kunlun Li , Kushan Ahmadian , Lawrence McAfee , Laya Sleiman , Leon Derczynski , Luis Vega , Maer Rodrigues de Melo , Makesh Narsimhan Sreedhar , Marcin Chochowski , Mark Cai , Markus Kliegl , Marta Stepniewska-Dziubinska , Matvei Novikov , Mehrzad Samadi , Meredith Price , Meriem Boubdir , Michael Boone , Michael Evans , Michal Bien , Michal Zawalski , Miguel Martinez , Mike Chrzanowski , Mohammad Shoeybi , Mostofa Patwary , Namit Dhameja , Nave Assaf , Negar Habibi , Nidhi Bhatia , Nikki Pope , Nima Tajbakhsh , Nirmal Kumar Juluru , Oleg Rybakov , Oleksii Hrinchuk , Oleksii Kuchaiev , Oluwatobi Olabiyi , Pablo Ribalta , Padmavathy Subramanian , Parth Chadha , Pavlo Molchanov , Peter Dykas , Peter Jin , Piotr Bialecki , Piotr Januszewski , Pradeep Thalasta , Prashant Gaikwad , Prasoon Varshney , Pritam Gundecha , Przemek Tredak , Rabeeh Karimi Mahabadi , Rajen Patel , Ran El-Yaniv , Ranjit Rajan , Ria Cheruvu , Rima Shahbazyan , Ritika Borkar , Ritu Gala , Roger Waleffe , Ruoxi Zhang , Russell J. Hewett , Ryan Prenger , Sahil Jain , Samuel Kriman , Sanjeev Satheesh , Saori Kaji , Sarah Yurick , Saurav Muralidharan , Sean Narenthiran , Seonmyeong Bak , Sepehr Sameni , Seungju Han , Shanmugam Ramasamy , Shaona Ghosh , Sharath Turuvekere Sreenivas , Shelby Thomas , Shizhe Diao , Shreya Gopal , Shrimai Prabhumoye , Shubham Toshniwal , Shuoyang Ding , Siddharth Singh , Siddhartha Jain , Somshubra Majumdar , Soumye Singhal , Stefania Alborghetti , Syeda Nahida Akter , Terry Kong , Tim Moon , Tomasz Hliwiak , Tomer Asida , Tony Wang , Tugrul Konuk , Twinkle Vashishth , Tyler Poon , Udi Karpas , Vahid Noroozi , Venkat Srinivasan , Vijay Korthikanti , Vikram Fugro , Vineeth Kalluru , Vitaly Kurin , Vitaly Lavrukhin , Wasi Uddin Ahmad , Wei Du , Wonmin Byeon , Ximing Lu , Xin Dong , Yashaswi Karnati , Yejin Choi , Yian Zhang , Ying Lin , Yonggan Fu , Yoshi Suhara , Zhen Dong , Zhiyu Li , Zhongbo Zhu , Zijia Chen

Transfer learning in reinforcement learning (RL) seeks to accelerate learning in new tasks by leveraging knowledge from related sources. Existing neurosymbolic transfer methods, however, typically rely on manually specified task automata,…

Artificial Intelligence · Computer Science 2026-05-08 Mahyar Alinejad , Yue Wang , Amrit Singh Bedi , George Atia

We introduce QwenLong-L1.5, a model that achieves superior long-context reasoning capabilities through systematic post-training innovations. The key technical breakthroughs of QwenLong-L1.5 are as follows: (1) Long-Context Data Synthesis…

Popular Neural Machine Translation model training uses strategies like backtranslation to improve BLEU scores, requiring large amounts of additional data and training. We introduce a class of conditional generative-discriminative hybrid…

Computation and Language · Computer Science 2020-10-16 Prathyusha Jwalapuram , Shafiq Joty , Youlin Shen

Although multimodal large language models (MLLMs) have achieved impressive performance, the multimodal instruction tuning stage often causes catastrophic forgetting of the base LLM's language ability, even in strong models like Llama3. To…

Computation and Language · Computer Science 2025-05-23 Zeping Yu , Sophia Ananiadou

Despite the remarkable success of large language models (LLMs) in English, a significant performance gap remains in non-English languages. To address this, we introduce a novel approach for strategically constructing a multilingual…

We present INTELLECT-3, a 106B-parameter Mixture-of-Experts model (12B active) trained with large-scale reinforcement learning on our end-to-end RL infrastructure stack. INTELLECT-3 achieves state of the art performance for its size across…

Large language models (LLMs) can potentially democratize access to medical knowledge. While many efforts have been made to harness and improve LLMs' medical knowledge and reasoning capacities, the resulting models are either closed-source…

We introduce Nemotron-Cascade 2, an open 30B MoE model with 3B activated parameters that delivers best-in-class reasoning and strong agentic capabilities. Despite its compact size, its mathematical and coding reasoning performance…

Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across modalities such as images and text. However, tabular data, despite being a critical real-world modality, remains relatively underexplored in…

Computation and Language · Computer Science 2026-03-26 Kun-Yang Yu , Zhi Zhou , Shi-Yu Tian , Xiao-Wen Yang , Zi-Yi Jia , Ming Yang , Zi-Jian Cheng , Lan-Zhe Guo , Yu-Feng Li

Multimodal Large Language Models (MLLMs) have made remarkable progress in multimodal perception and reasoning by bridging vision and language. However, most existing MLLMs perform reasoning primarily with textual CoT, which limits their…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Jintao Tong , Shilin Yan , Hongwei Xue , Xiaojun Tang , Kunyu Shi , Guannan Zhang , Ruixuan Li , Yixiong Zou

Pre-training datasets are typically collected from web content and lack inherent domain divisions. For instance, widely used datasets like Common Crawl do not include explicit domain labels, while manually curating labeled datasets such as…

Computation and Language · Computer Science 2025-12-02 Shizhe Diao , Yu Yang , Yonggan Fu , Xin Dong , Dan Su , Markus Kliegl , Zijia Chen , Peter Belcak , Yoshi Suhara , Hongxu Yin , Mostofa Patwary , Yingyan , Lin , Jan Kautz , Pavlo Molchanov

Reasoning capabilities are crucial for Large Language Models (LLMs), yet a notable gap exists between English and non-English languages. To bridge this disparity, some works fine-tune LLMs to relearn reasoning capabilities in non-English…

Computation and Language · Computer Science 2024-05-28 Zixian Huang , Wenhao Zhu , Gong Cheng , Lei Li , Fei Yuan

Pre-trained language models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks. GPT-3 has shown that scaling up pre-trained language models can further exploit their enormous potential. A unified…