English
Related papers

Related papers: Phi-3 Technical Report: A Highly Capable Language …

200 papers

Large Language Models (LLMs) are widely adopted for assisting in software development tasks, yet their performance evaluations have narrowly focused on the functional correctness of generated code. Human programmers, however, require…

Software Engineering · Computer Science 2024-12-06 Yun Peng , Akhilesh Deepak Gotmare , Michael Lyu , Caiming Xiong , Silvio Savarese , Doyen Sahoo

We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus…

In this paper, we explore FP8 low-bit data formats for efficient training of large language models (LLMs). Our key insight is that most variables, such as gradients and optimizer states, in LLM training can employ low-precision data formats…

General-purpose large language models (LLMs), despite their broad capabilities accrued from open-world data, frequently exhibit suboptimal performance when confronted with the nuanced and specialized demands inherent in real-time…

Networking and Internet Architecture · Computer Science 2025-05-14 Vignesh Ethiraj , Divya Vijay , Sidhanth Menon , Heblin Berscilla

A multimodal AI agent is characterized by its ability to process and learn from various types of data, including natural language, visual, and audio inputs, to inform its actions. Despite advancements in large language models that…

Computation and Language · Computer Science 2024-04-19 Wei Chen , Zhiyuan Li

The resurgence and rapid advancement of Generative Artificial Intelligence (GenAI) in 2023 has catalyzed transformative shifts across numerous industry sectors, including urban transportation and logistics. This study investigates the…

Artificial Intelligence · Computer Science 2024-09-24 Shaowei Ying , Zhenlong Li , Manzhu Yu

Large language models are powerful but often limited by high computational cost, privacy concerns, and English-centric training. Recent progress demonstrates that small, efficient models with around one billion parameters can deliver strong…

Computation and Language · Computer Science 2025-12-16 Anna Aksenova , Boris Zverkov , Nicola Dainese , Alexander Nikitin , Pekka Marttinen

There has been a surge in LLM evaluation research to understand LLM capabilities and limitations. However, much of this research has been confined to English, leaving LLM building and evaluation for non-English languages relatively…

Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE models on HPC platforms is hindered by large memory footprints, frequent large-scale…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-07 Sajal Dash , Feiyi Wang

Motivational interviewing (MI) promotes behavioural change in substance use disorders. Its fidelity is measured using the Motivational Interviewing Treatment Integrity (MITI) framework. While large language models (LLMs) can potentially…

Computation and Language · Computer Science 2026-03-05 Aishwariya Jha , Prakrithi Shivaprakash , Lekhansh Shukla , Animesh Mukherjee , Prabhat Chand , Pratima Murthy

In this work, we introduce LokiLM, a 1.4B parameter large language model trained on 500B tokens. Our model performs strongly in natural language reasoning tasks and achieves state-of-the-art performance among models with 1.5B parameters or…

Computation and Language · Computer Science 2024-07-11 Justin Kiefel , Shrey Shah

We present foundation language models developed to power Apple Intelligence features, including a ~3 billion parameter model designed to run efficiently on devices and a large server-based language model designed for Private Cloud Compute.…

Artificial Intelligence · Computer Science 2026-05-28 Tom Gunter , Zirui Wang , Chong Wang , Ruoming Pang , Andy Narayanan , Aonan Zhang , Bowen Zhang , Chen Chen , Chung-Cheng Chiu , David Qiu , Deepak Gopinath , Dian Ang Yap , Dong Yin , Feng Nan , Floris Weers , Guoli Yin , Haoshuo Huang , Jianyu Wang , Jiarui Lu , John Peebles , Ke Ye , Mark Lee , Nan Du , Qibin Chen , Quentin Keunebroek , Sam Wiseman , Syd Evans , Tao Lei , Vivek Rathod , Xiang Kong , Xianzhi Du , Yanghao Li , Yongqiang Wang , Yuan Gao , Zaid Ahmed , Zhaoyang Xu , Zhiyun Lu , Al Rashid , Albin Madappally Jose , Alec Doane , Alfredo Bencomo , Allison Vanderby , Andrew Hansen , Ankur Jain , Anupama Mann Anupama , Areeba Kamal , Bugu Wu , Carolina Brum , Charlie Maalouf , Chinguun Erdenebileg , Chris Dulhanty , Daniel Parilla , Dominik Moritz , Doug Kang , Eduardo Jimenez , Evan Ladd , Fangping Shi , Felix Bai , Frank Chu , Fred Hohman , Hadas Kotek , Hannah Gillis Coleman , Jane Li , Jeffrey Bigham , Jeffery Cao , Jeff Lai , Jessica Cheung , Jiulong Shan , Joe Zhou , John Li , Jun Qin , Karanjeet Singh , Karla Vega , Kelvin Zou , Laura Heckman , Lauren Gardiner , Margit Bowler , Maria Cordell , Meng Cao , Nicole Hay , Nilesh Shahdadpuri , Otto Godwin , Pranay Dighe , Pushyami Rachapudi , Ramsey Tantawi , Roman Frigg , Sam Davarnia , Sanskruti Shah , Saptarshi Guha , Sasha Sirovica , Shen Ma , Shuang Ma , Simon Wang , Sulgi Kim , Suma Jayaram , Vaishaal Shankar , Varsha Paidi , Vivek Kumar , Xin Wang , Xin Zheng , Walker Cheng , Yael Shrager , Yang Ye , Yasu Tanaka , Yihao Guo , Yunsong Meng , Zhao Tang Luo , Zhi Ouyang , Alp Aygar , Alvin Wan , Andrew Walkingshaw , Andy Narayanan , Antonie Lin , Arsalan Farooq , Brent Ramerth , Colorado Reed , Chris Bartels , Chris Chaney , David Riazati , Eric Liang Yang , Erin Feldman , Gabriel Hochstrasser , Guillaume Seguin , Irina Belousova , Joris Pelemans , Karen Yang , Keivan Alizadeh Vahid , Liangliang Cao , Mahyar Najibi , Marco Zuliani , Max Horton , Minsik Cho , Nikhil Bhendawade , Patrick Dong , Piotr Maj , Pulkit Agrawal , Qi Shan , Qichen Fu , Regan Poston , Sam Xu , Shuangning Liu , Sushma Rao , Tashweena Heeramun , Thomas Merth , Uday Rayala , Victor Cui , Vivek Rangarajan Sridhar , Wencong Zhang , Wenqi Zhang , Wentao Wu , Xingyu Zhou , Xinwen Liu , Yang Zhao , Yin Xia , Zhile Ren , Zhongzheng Ren

The interest in developing small language models (SLM) for on-device deployment is fast growing. However, the existing SLM design hardly considers the device hardware characteristics. Instead, this work presents a simple yet effective…

Computation and Language · Computer Science 2024-11-11 Rongjie Yi , Xiang Li , Weikai Xie , Zhenyan Lu , Chenghua Wang , Ao Zhou , Shangguang Wang , Xiwen Zhang , Mengwei Xu

We present Mify-Coder, a 2.5B-parameter code model trained on 4.2T tokens using a compute-optimal strategy built on the Mify-2.5B foundation model. Mify-Coder achieves comparable accuracy and safety while significantly outperforming much…

Software Engineering · Computer Science 2026-01-01 Abhinav Parmar , Abhisek Panigrahi , Abhishek Kumar Dwivedi , Abhishek Bhattacharya , Adarsh Ramachandra , Aditya Choudhary , Aditya Garg , Aditya Raj , Alankrit Bhatt , Alpesh Yadav , Anant Vishnu , Ananthu Pillai , Ankush Kumar , Aryan Patnaik , Aswatha Narayanan S , Avanish Raj Singh , Bhavya Shree Gadda , Brijesh Pankajbhai Kachhadiya , Buggala Jahnavi , Chidurala Nithin Krishna , Chintan Shah , Chunduru Akshaya , Debarshi Banerjee , Debrup Dey , Deepa R. , Deepika B G , Faiz ur Rahman , Gagan Gayari , Gudhi Jagadeesh Kumar Naidu , Gursimar Singh , Harshal Tyagi , Harshini K , James Mani Vathalloor , Jayarama Nettar , Jayashree Gajjam , Joe Walter Sugil George , Kamalakara Sri Krishna Tadepalli , Kamalkumar Rathinasamy , Karan Chaurasia , Karthikeyan S , Kashish Arora , Kaushal Desai , Khushboo Buwade , Kiran Manjrekar , Malikireddy Venkata Sai Likhitha , Manjunath A , Mitali Mahavir Bedmutha , Mohammed Rafee Tarafdar , Nikhil Tiwari , Nikitha K Gigi , Pavan Ravikumar , Pendyala Swarnanjali , Piyush Anand , Prakash Chandrasekar , Prasanna Bhalchandra Gawade , Prasanth Sivan , Preeti Khurana , Priyanshi Babbar , Rajab Ali Mondal , Rajesh Kumar Vissapragada , Rajeshwari Ganesan , Rajeswari Koppisetti , Ramjee R. , Ramkumar Thiruppathisamy , Rani G. S. , S Reka , Samarth Gupta , Sandeep Reddy Kothakota , Sarathy K , Sathyanarayana Sampath Kumar , Saurabh Kumar , Shashank Khasare , Shenbaga Devi Venkatesh Kumar , Shiva Rama Krishna Parvatham , Shoeb Shaikh , Shrishanmathi A , Shubham Pathak , Sree Samhita Koppaka , Sreenivasa Raghavan K S , Sreeram Venkatasubramanian , Suprabha Desai Bojja , Swetha R , Syed Ahmed , Chinmai Harshitha Thota , Tushar Yadav , Veeravelly Kusumitha , V V S S Prasanth Patnaik , Vidya Sri Sesetti , Vijayakeerthi K , Vikram Raj Bakshi , Vinay K K , Vinoth Kumar Loganathan , Vipin Tiwari , Vivek Kumar Shrivastav , V Venkata Sri Datta Charan , Wasim Akhtar Khan

Multimodal large language models (MLLMs) have demonstrated impressive performance in vision-language tasks across a broad spectrum of domains. However, the large model scale and associated high computational costs pose significant…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Zhangwei Gao , Zhe Chen , Erfei Cui , Yiming Ren , Weiyun Wang , Jinguo Zhu , Hao Tian , Shenglong Ye , Junjun He , Xizhou Zhu , Lewei Lu , Tong Lu , Yu Qiao , Jifeng Dai , Wenhai Wang

We introduce F2LLM - Foundation to Feature Large Language Models, a suite of state-of-the-art embedding models in three sizes: 0.6B, 1.7B, and 4B. Unlike previous top-ranking embedding models that require massive contrastive pretraining,…

Computation and Language · Computer Science 2025-10-03 Ziyin Zhang , Zihan Liao , Hang Yu , Peng Di , Rui Wang

Although recent advances in scaling large language models (LLMs) have resulted in improvements on many NLP tasks, it remains unclear whether these models trained primarily with general web text are the right tool in highly specialized,…

Large language models (LLMs) have showcased profound capabilities in language understanding and generation, facilitating a wide array of applications. However, there is a notable paucity of detailed, open-sourced methodologies on…

Smaller vision-language models (VLMs) are becoming increasingly important for privacy-focused, on-device applications due to their ability to run efficiently on consumer hardware for processing enterprise commercial documents and images.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Shaikat Galib , Shanshan Wang , Guanshuo Xu , Pascal Pfeiffer , Ryan Chesler , Mark Landry , Sri Satish Ambati

Recent work in language modeling demonstrates that training large transformer models advances the state of the art in Natural Language Processing applications. However, very large models can be quite difficult to train due to memory…

Computation and Language · Computer Science 2020-03-17 Mohammad Shoeybi , Mostofa Patwary , Raul Puri , Patrick LeGresley , Jared Casper , Bryan Catanzaro
‹ Prev 1 4 5 6 7 8 10 Next ›