English
Related papers

Related papers: Nanbeige4-3B Technical Report: Exploring the Front…

200 papers

We introduce two multilingual, multimodal foundation language models that power Apple Intelligence features across Apple devices and services: i a 3B-parameter on-device model optimized for Apple silicon through architectural innovations…

Machine Learning · Computer Science 2025-08-28 Ethan Li , Anders Boesen Lindbo Larsen , Chen Zhang , Xiyou Zhou , Jun Qin , Dian Ang Yap , Narendran Raghavan , Xuankai Chang , Margit Bowler , Eray Yildiz , John Peebles , Hannah Gillis Coleman , Matteo Ronchi , Peter Gray , Keen You , Anthony Spalvieri-Kruse , Ruoming Pang , Reed Li , Yuli Yang , Emad Soroush , Zhiyun Lu , Crystal Xiao , Rong Situ , Jordan Huffaker , David Griffiths , Zaid Ahmed , Peng Zhang , Daniel Parilla , Asaf Liberman , Jennifer Mallalieu , Parsa Mazaheri , Qibin Chen , Manjot Bilkhu , Aonan Zhang , Eric Wang , Dave Nelson , Michael FitzMaurice , Thomas Voice , Jeremy Liu , Josh Shaffer , Shiwen Zhao , Prasanth Yadla , Farzin Rasteh , Pengsheng Guo , Arsalan Farooq , Jeremy Snow , Stephen Murphy , Tao Lei , Minsik Cho , George Horrell , Sam Dodge , Lindsay Hislop , Sumeet Singh , Alex Dombrowski , Aiswarya Raghavan , Sasha Sirovica , Mandana Saebi , Faye Lao , Max Lam , TJ Lu , Zhaoyang Xu , Karanjeet Singh , Marc Kirchner , David Mizrahi , Rajat Arora , Haotian Zhang , Henry Mason , Lawrence Zhou , Yi Hua , Ankur Jain , Felix Bai , Joseph Astrauskas , Floris Weers , Josh Gardner , Mira Chiang , Yi Zhang , Pulkit Agrawal , Tony Sun , Quentin Keunebroek , Matthew Hopkins , Bugu Wu , Tao Jia , Chen Chen , Xingyu Zhou , Nanzhu Wang , Peng Liu , Ruixuan Hou , Rene Rauch , Yuan Gao , Afshin Dehghan , Jonathan Janke , Zirui Wang , Cha Chen , Xiaoyi Ren , Feng Nan , Josh Elman , Dong Yin , Yusuf Goren , Jeff Lai , Yiran Fei , Syd Evans , Muyang Yu , Guoli Yin , Yi Qin , Erin Feldman , Isha Garg , Aparna Rajamani , Karla Vega , Walker Cheng , TJ Collins , Hans Han , Raul Rea Menacho , Simon Yeung , Sophy Lee , Phani Mutyala , Ying-Chang Cheng , Zhe Gan , Sprite Chu , Justin Lazarow , Alessandro Pappalardo , Federico Scozzafava , Jing Lu , Erik Daxberger , Laurent Duchesne , Jen Liu , David Güera , Stefano Ligas , Mary Beth Kery , Brent Ramerth , Ciro Sannino , Marcin Eichner , Haoshuo Huang , Rui Qian , Moritz Schwarzer-Becker , David Riazati , Mingfei Gao , Bailin Wang , Jack Cackler , Yang Lu , Ransen Niu , John Dennison , Guillaume Klein , Jeffrey Bigham , Deepak Gopinath , Navid Shiee , Darren Botten , Guillaume Tartavel , Alex Guillen Garcia , Sam Xu , Victoria MönchJuan Haladjian , Zi-Yi Dou , Matthias Paulik , Adolfo Lopez Mendez , Zhen Li , Hong-You Chen , Chao Jia , Dhaval Doshi , Zhengdong Zhang , Raunak Manjani , Aaron Franklin , Zhile Ren , David Chen , Artsiom Peshko , Nandhitha Raghuram , Hans Hao , Jiulong Shan , Kavya Nerella , Ramsey Tantawi , Vivek Kumar , Saiwen Wang , Brycen Wershing , Bhuwan Dhingra , Dhruti Shah , Ob Adaranijo , Xin Zheng , Tait Madsen , Hadas Kotek , Chang Liu , Yin Xia , Hanli Li , Suma Jayaram , Yanchao Sun , Ahmed Fakhry , Vasileios Saveris , Dustin Withers , Yanghao Li , Alp Aygar , Andres Romero Mier Y Teran , Kaiwei Huang , Mark Lee , Xiujun Li , Yuhong Li , Tyler Johnson , Jay Tang , Joseph Yitan Cheng , Futang Peng , Andrew Walkingshaw , Lucas Guibert , Abhishek Sharma , Cheng Shen , Piotr Maj , Yasutaka Tanaka , You-Cyuan Jhang , Vivian Ma , Tommi Vehvilainen , Kelvin Zou , Jeff Nichols , Matthew Lei , David Qiu , Yihao Qian , Gokul Santhanam , Wentao Wu , Yena Han , Dominik Moritz , Haijing Fu , Mingze Xu , Vivek Rathod , Jian Liu , Louis D'hauwe , Qin Ba , Haitian Sun , Haoran Yan , Philipp Dufter , Anh Nguyen , Yihao Feng , Emma Wang , Keyu He , Rahul Nair , Sanskruti Shah , Jiarui Lu , Patrick Sonnenberg , Jeremy Warner , Yuanzhi Li , Bowen Pan , Ziyi Zhong , Joe Zhou , Sam Davarnia , Olli Saarikivi , Irina Belousova , Rachel Burger , Shang-Chen Wu , Di Feng , Bas Straathof , James Chou , Yuanyang Zhang , Marco Zuliani , Eduardo Jimenez , Abhishek Sundararajan , Xianzhi Du , Chang Lan , Nilesh Shahdadpuri , Peter Grasch , Sergiu Sima , Josh Newnham , Varsha Paidi , Jianyu Wang , Kaelen Haag , Alex Braunstein , Daniele Molinari , Richard Wei , Brenda Yang , Nicholas Lusskin , Joanna Arreaza-Taylor , Meng Cao , Nicholas Seidl , Simon Wang , Jiaming Hu , Yiping Ma , Mengyu Li , Kieran Liu , Hang Su , Sachin Ravi , Chong Wang , Xin Wang , Kevin Smith , Haoxuan You , Binazir Karimzadeh , Rui Li , Jinhao Lei , Wei Fang , Alec Doane , Sam Wiseman , Ismael Fernandez , Jane Li , Andrew Hansen , Javier Movellan , Christopher Neubauer , Hanzhi Zhou , Chris Chaney , Nazir Kamaldin , Valentin Wolf , Fernando Bermúdez-Medina , Joris Pelemans , Peter Fu , Howard Xing , Xiang Kong , Wayne Shan , Gabriel Jacoby-Cooper , Dongcai Shen , Tom Gunter , Guillaume Seguin , Fangping Shi , Shiyu Li , Yang Xu , Areeba Kamal , Dan Masi , Saptarshi Guha , Qi Zhu , Jenna Thibodeau , Changyuan Zhang , Rebecca Callahan , Charles Maalouf , Wilson Tsao , Boyue Li , Qingqing Cao , Naomy Sabo , Cheng Leong , Yi Wang , Anupama Mann Anupama , Colorado Reed , Kenneth Jung , Zhifeng Chen , Mohana Prasad Sathya Moorthy , Yifei He , Erik Hornberger , Devi Krishna , Senyu Tong , Michael , Lee , David Haldimann , Yang Zhao , Bowen Zhang , Chang Gao , Chris Bartels , Sushma Rao , Nathalie Tran , Simon Lehnerer , Co Giang , Patrick Dong , Junting Pan , Biyao Wang , Dongxu Li , Mehrdad Farajtabar , Dongseong Hwang , Grace Duanmu , Eshan Verma , Sujeeth Reddy , Qi Shan , Hongbin Gao , Nan Du , Pragnya Sridhar , Forrest Huang , Yingbo Wang , Nikhil Bhendawade , Diane Zhu , Sai Aitharaju , Fred Hohman , Lauren Gardiner , Chung-Cheng Chiu , Yinfei Yang , Alper Kokmen , Frank Chu , Ke Ye , Kaan Elgin , Oron Levy , John Park , Donald Zhang , Eldon Schoop , Nina Wenzel , Michael Booker , Hyunjik Kim , Chinguun Erdenebileg , Nan Dun , Eric Liang Yang , Priyal Chhatrapati , Vishaal Mahtani , Haiming Gang , Kohen Chia , Deepa Seshadri , Donghan Yu , Yan Meng , Kelsey Peterson , Zhen Yang , Yongqiang Wang , Carina Peng , Doug Kang , Anuva Agarwal , Albert Antony , Juan Lao Tebar , Albin Madappally Jose , Regan Poston , Andy De Wang , Gerard Casamayor , Elmira Amirloo , Violet Yao , Wojciech Kryscinski , Kun Duan , Lezhi L

Large language models (LLMs) with billions of parameters have demonstrated outstanding performance on various natural language processing tasks. This report presents OpenBA, an open-sourced 15B bilingual asymmetric seq2seq model, to…

Computation and Language · Computer Science 2024-11-26 Juntao Li , Zecheng Tang , Yuyang Ding , Pinzheng Wang , Pei Guo , Wangjie You , Dan Qiao , Wenliang Chen , Guohong Fu , Qiaoming Zhu , Guodong Zhou , Min Zhang

We introduce CRPE (Code Reasoning Process Enhancer), an innovative three-stage framework for data synthesis and model training that advances the development of sophisticated code reasoning capabilities in large language models (LLMs).…

Software Engineering · Computer Science 2025-05-19 Ningxin Gui , Qianghuai Jia , Feijun Jiang , Yuling Jiao , dechun wang , Jerry Zhijian Yang

Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually computationally expensive, so it is difficult to efficiently…

Computation and Language · Computer Science 2020-10-19 Xiaoqi Jiao , Yichun Yin , Lifeng Shang , Xin Jiang , Xiao Chen , Linlin Li , Fang Wang , Qun Liu

While large language models (LLMs) have demonstrated exceptional performance in recent natural language processing (NLP) tasks, their deployment poses substantial challenges due to high computational and memory demands in real-world…

Computation and Language · Computer Science 2024-02-27 Chenglin Li , Qianglong Chen , Liangyue Li , Caiyu Wang , Yicheng Li , Zulong Chen , Yin Zhang

Despite the effectiveness of data selection for large language models (LLMs) during pretraining and instruction fine-tuning phases, improving data efficiency in supervised fine-tuning (SFT) for specialized domains poses significant…

Computation and Language · Computer Science 2024-12-06 Yu Yang , Siddhartha Mishra , Jeffrey N Chiang , Baharan Mirzasoleiman

We ask whether a pure spiking backbone can learn large-scale language modeling from random initialization, without Transformer distillation. We introduce NeuronSpark, a 0.9B-parameter SNN language model trained with next-token prediction…

Artificial Intelligence · Computer Science 2026-03-18 Zhengzheng Tang

Logical reasoning remains a challenge for natural language processing, but it can be improved by training language models to mimic theorem provers on procedurally generated problems. Previous work used domain-specific proof generation…

Computation and Language · Computer Science 2024-06-18 Damien Sileo

Continued pre-training of small language models offers a promising path for domain adaptation with limited computational resources. I've investigated this approach within educational domains, evaluating it as a resource-efficient…

Computation and Language · Computer Science 2025-04-15 Salman Faroz

The challenge of reducing the size of Large Language Models (LLMs) while maintaining their performance has gained significant attention. However, existing methods, such as model distillation and transfer learning, often fail to achieve high…

Large language models (LLMs) with instruction fine-tuning demonstrate superior generative capabilities. However, these models are resource-intensive. To alleviate this issue, we explore distilling knowledge from instruction-tuned LLMs into…

Computation and Language · Computer Science 2024-01-30 Minghao Wu , Abdul Waheed , Chiyu Zhang , Muhammad Abdul-Mageed , Alham Fikri Aji

Instruction fine-tuning pretrained LLMs for diverse downstream tasks has demonstrated remarkable success and has captured the interest of both academics and practitioners. To ensure such fine-tuned LLMs align with human preferences,…

Computation and Language · Computer Science 2024-04-19 Chandeepa Dissanayake , Lahiru Lowe , Sachith Gunasekara , Yasiru Ratnayake

Small Language Models (SLMs) are attractive for cost-sensitive and resource-limited settings due to their efficient, low-latency inference. However, they often struggle with complex, knowledge-intensive tasks that require structured…

Artificial Intelligence · Computer Science 2026-02-03 Shaoxiong Yang , Junting Li , Mengyuan Zhang , Chao Li , Wei Liu , Jian Luan

Large language models show promise for legal applications, but deploying frontier models raises concerns about cost, latency, and data privacy. We evaluate whether sub-10B parameter models can serve as practical alternatives by testing nine…

Computation and Language · Computer Science 2026-03-30 Snehit Vaddi

We present NanoFlux, a novel adversarial framework for generating targeted training data to improve LLM reasoning, where adversarially-generated datasets containing fewer than 200 examples outperform conventional fine-tuning approaches. The…

Machine Learning · Computer Science 2026-03-18 Raviteja Anantha , Soheil Hor , Teodor Nicola Antoniu , Layne C. Price

We present LFM2, a family of Liquid Foundation Models designed for efficient on-device deployment and strong task capabilities. Using hardware-in-the-loop architecture search under edge latency and memory constraints, we obtain a compact…

Large language model fine-tuning techniques typically depend on extensive labeled data, external guidance, and feedback, such as human alignment, scalar rewards, and demonstration. However, in practical application, the scarcity of specific…

Computation and Language · Computer Science 2024-12-31 Jia Liu , Yue Wang , Zhiqi Lin , Min Chen , Yixue Hao , Long Hu

Increasing test-time compute for LLMs shows promise across domains but remains underexplored in code generation, despite extensive study in math. In this paper, we propose S*, the first hybrid test-time scaling framework that substantially…

Machine Learning · Computer Science 2025-02-21 Dacheng Li , Shiyi Cao , Chengkun Cao , Xiuyu Li , Shangyin Tan , Kurt Keutzer , Jiarong Xing , Joseph E. Gonzalez , Ion Stoica

Preference optimization techniques have become a standard final stage for training state-of-art large language models (LLMs). However, despite widespread adoption, the vast majority of work to-date has focused on first-class citizen…

Computation and Language · Computer Science 2024-07-04 John Dang , Arash Ahmadian , Kelly Marchisio , Julia Kreutzer , Ahmet Üstün , Sara Hooker

We present STEP3-VL-10B, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. STEP3-VL-10B is realized through two strategic shifts: first, a…

‹ Prev 1 3 4 5 6 7 10 Next ›