English
Related papers

Related papers: Tiny Model, Big Logic: Diversity-Driven Optimizati…

200 papers

Deploying large language models (LLMs) is challenging because they are memory inefficient and compute-intensive for practical applications. In reaction, researchers train smaller task-specific models by either finetuning with human labels…

Computation and Language · Computer Science 2023-07-06 Cheng-Yu Hsieh , Chun-Liang Li , Chih-Kuan Yeh , Hootan Nakhost , Yasuhisa Fujii , Alexander Ratner , Ranjay Krishna , Chen-Yu Lee , Tomas Pfister

In recent years, general-purpose large language models (LLMs) such as GPT, Gemini, Claude, and DeepSeek have advanced at an unprecedented pace. Despite these achievements, their application to finance remains challenging, due to fragmented…

Large reasoning models (LRMs) such as Claude 3.7 Sonnet and OpenAI o1 achieve strong performance on mathematical benchmarks using lengthy chain-of-thought (CoT) reasoning, but the resulting traces are often unnecessarily verbose. This…

Computation and Language · Computer Science 2025-06-13 Ye Yu , Yaoning Yu , Haohan Wang

Large reasoning models (LRMs) have exhibited remarkable reasoning capabilities through inference-time scaling, but this progress has also introduced considerable redundancy and inefficiency into their reasoning processes, resulting in…

Artificial Intelligence · Computer Science 2025-07-18 Xingyang He , Xiao Ling , Jie Liu

Recent advancements, such as DeepSeek-Prover-V2-671B and Kimina-Prover-Preview-72B, demonstrate a prevailing trend in leveraging reinforcement learning (RL)-based large-scale training for automated theorem proving. Surprisingly, we discover…

Artificial Intelligence · Computer Science 2025-06-16 Chenrui Cao , Liangcheng Song , Zenan Li , Xinyi Le , Xian Zhang , Hui Xue , Fan Yang

Large language models (LLMs) have demonstrated impressive reasoning capabilities, but scaling their performance often relies on massive reasoning datasets that are computationally expensive to train on. Existing data selection methods aim…

Artificial Intelligence · Computer Science 2025-10-24 Shaobo Wang , Yongliang Miao , Yuancheng Liu , Qianli Ma , Ning Liao , Linfeng Zhang

The power of large language models (LLMs) has been demonstrated through numerous data and computing resources. However, the application of language models on mobile devices is facing huge challenge on the computation and memory costs, that…

Computation and Language · Computer Science 2025-04-04 Yehui Tang , Kai Han , Fangcheng Liu , Yunsheng Ni , Yuchuan Tian , Zheyuan Bai , Yi-Qi Hu , Sichao Liu , Shangling Jui , Yunhe Wang

Large language models (LLMs) open up new horizons for sequential recommendations, owing to their remarkable language comprehension and generation capabilities. However, there are still numerous challenges that should be addressed to…

Information Retrieval · Computer Science 2024-03-29 Yuling Wang , Changxin Tian , Binbin Hu , Yanhua Yu , Ziqi Liu , Zhiqiang Zhang , Jun Zhou , Liang Pang , Xiao Wang

Language model pretraining with next token prediction has proved effective for scaling compute but is limited to the amount of available training data. Scaling reinforcement learning (RL) unlocks a new axis for the continued improvement of…

Recent advancements in Large Language Models (LLMs) have revealed a significant performance gap between closed-source and open-source models, particularly in tasks requiring complex reasoning and precise instruction following. This paper…

Artificial Intelligence · Computer Science 2025-07-01 Ziqi Zhong , Xunzhu Tang

Recently developed large language models (LLMs) such as ChatGPT, Claude, and Llama have demonstrated impressive abilities, and even surpass human-level performance in several tasks. Despite their success, the resource-intensive demands of…

Computation and Language · Computer Science 2024-06-17 Jie Wu , Yufeng Zhu , Lei Shen , Xuqing Lu

We present Mify-Coder, a 2.5B-parameter code model trained on 4.2T tokens using a compute-optimal strategy built on the Mify-2.5B foundation model. Mify-Coder achieves comparable accuracy and safety while significantly outperforming much…

Software Engineering · Computer Science 2026-01-01 Abhinav Parmar , Abhisek Panigrahi , Abhishek Kumar Dwivedi , Abhishek Bhattacharya , Adarsh Ramachandra , Aditya Choudhary , Aditya Garg , Aditya Raj , Alankrit Bhatt , Alpesh Yadav , Anant Vishnu , Ananthu Pillai , Ankush Kumar , Aryan Patnaik , Aswatha Narayanan S , Avanish Raj Singh , Bhavya Shree Gadda , Brijesh Pankajbhai Kachhadiya , Buggala Jahnavi , Chidurala Nithin Krishna , Chintan Shah , Chunduru Akshaya , Debarshi Banerjee , Debrup Dey , Deepa R. , Deepika B G , Faiz ur Rahman , Gagan Gayari , Gudhi Jagadeesh Kumar Naidu , Gursimar Singh , Harshal Tyagi , Harshini K , James Mani Vathalloor , Jayarama Nettar , Jayashree Gajjam , Joe Walter Sugil George , Kamalakara Sri Krishna Tadepalli , Kamalkumar Rathinasamy , Karan Chaurasia , Karthikeyan S , Kashish Arora , Kaushal Desai , Khushboo Buwade , Kiran Manjrekar , Malikireddy Venkata Sai Likhitha , Manjunath A , Mitali Mahavir Bedmutha , Mohammed Rafee Tarafdar , Nikhil Tiwari , Nikitha K Gigi , Pavan Ravikumar , Pendyala Swarnanjali , Piyush Anand , Prakash Chandrasekar , Prasanna Bhalchandra Gawade , Prasanth Sivan , Preeti Khurana , Priyanshi Babbar , Rajab Ali Mondal , Rajesh Kumar Vissapragada , Rajeshwari Ganesan , Rajeswari Koppisetti , Ramjee R. , Ramkumar Thiruppathisamy , Rani G. S. , S Reka , Samarth Gupta , Sandeep Reddy Kothakota , Sarathy K , Sathyanarayana Sampath Kumar , Saurabh Kumar , Shashank Khasare , Shenbaga Devi Venkatesh Kumar , Shiva Rama Krishna Parvatham , Shoeb Shaikh , Shrishanmathi A , Shubham Pathak , Sree Samhita Koppaka , Sreenivasa Raghavan K S , Sreeram Venkatasubramanian , Suprabha Desai Bojja , Swetha R , Syed Ahmed , Chinmai Harshitha Thota , Tushar Yadav , Veeravelly Kusumitha , V V S S Prasanth Patnaik , Vidya Sri Sesetti , Vijayakeerthi K , Vikram Raj Bakshi , Vinay K K , Vinoth Kumar Loganathan , Vipin Tiwari , Vivek Kumar Shrivastav , V Venkata Sri Datta Charan , Wasim Akhtar Khan

Recent advances leverage post-training to enhance model reasoning performance, which typically requires costly training pipelines and still suffers from inefficient, overly lengthy outputs. We introduce Speculative Thinking, a training-free…

Computation and Language · Computer Science 2025-04-18 Wang Yang , Xiang Yue , Vipin Chaudhary , Xiaotian Han

Large language models (LLMs) such as ChatGPT o1, ChatGPT o3, and DeepSeek R1 have shown great potential in solving difficult problems. However, current LLM evaluation benchmarks are limited to one-step interactions. Some of the existing…

Machine Learning · Computer Science 2025-12-01 Huanyu Li , Zongyuan Li , Wei Huang , Xian Guo

The reasoning capabilities of large language models (LLMs) have advanced rapidly, particularly following the release of DeepSeek R1, which has inspired a surge of research into data quality and reinforcement learning (RL) algorithms.…

Machine Learning · Computer Science 2025-11-04 Jian Yao , Ran Cheng , Xingyu Wu , Jibin Wu , Kay Chen Tan

The success of DeepSeek-R1 underscores the significant role of reinforcement learning (RL) in enhancing the reasoning capabilities of large language models (LLMs). In this work, we present Skywork-OR1, an effective and scalable RL…

Large language models (LLMs) have demonstrated remarkable progress in reasoning, often through supervised fine-tuning (SFT). However, SFT is resource-intensive, relying on large curated datasets, rejection-sampled demonstrations, and…

Computation and Language · Computer Science 2026-05-22 Jingyuan Wang , Yankai Chen , Zhonghang Li , Chao Huang

This study investigates the in-context learning capabilities of various decoder-only transformer-based language models with different model sizes and training data, including GPT2, SmolLM2, OpenELM, TinyLlama, Stable LM, and Gemma 2. We…

Computation and Language · Computer Science 2025-02-24 Yen-Che Hsiao , Abhishek Dutta

Large language models (LLMs) have demonstrated remarkable capabilities in various natural language processing tasks. However, achieving strong performance in specialized domains like mathematical reasoning and non-English languages often…

Computation and Language · Computer Science 2025-03-19 Huy Hoang Ha

We introduce Bielik v3, a series of parameter-efficient generative text models (1.5B and 4.5B) optimized for Polish language processing. These models demonstrate that smaller, well-optimized architectures can achieve performance comparable…

Machine Learning · Computer Science 2025-05-12 Krzysztof Ociepa , Łukasz Flis , Remigiusz Kinas , Krzysztof Wróbel , Adrian Gwoździej
‹ Prev 1 3 4 5 6 7 10 Next ›