English
Related papers

Related papers: StarCoder 2 and The Stack v2: The Next Generation

200 papers

Training LLMs for code-related tasks typically depends on high-quality code-documentation pairs, which are costly to curate and often scarce for niche programming languages. We introduce BatCoder, a self-supervised reinforcement learning…

Large language models (LLMs) such as Llama 2 perform very well on tasks that involve both natural language and source code, particularly code summarization and code generation. We show that for the task of code summarization, the…

Software Engineering · Computer Science 2024-04-15 Rajarshi Haldar , Julia Hockenmaier

Language models have shown promising performance on the task of translating natural language questions into SQL queries (Text-to-SQL). However, most of the state-of-the-art (SOTA) approaches rely on powerful yet closed-source large language…

Computation and Language · Computer Science 2024-02-27 Haoyang Li , Jing Zhang , Hanbing Liu , Ju Fan , Xiaokang Zhang , Jun Zhu , Renjie Wei , Hongyan Pan , Cuiping Li , Hong Chen

As large language models (LLMs) play an increasingly important role in code generation, enhancing both correctness and efficiency has become crucial. Current methods primarily focus on correctness, often overlooking efficiency. To address…

Computation and Language · Computer Science 2025-06-17 Dong Huang , Guangtao Zeng , Jianbo Dai , Meng Luo , Han Weng , Yuhao Qing , Heming Cui , Zhijiang Guo , Jie M. Zhang

Large Language Models for Code (Code LLM) are flourishing. New and powerful models are released on a weekly basis, demonstrating remarkable performance on the code generation task. Various approaches have been proposed to boost the code…

Computation and Language · Computer Science 2023-07-28 Bo Shen , Jiaxin Zhang , Taihong Chen , Daoguang Zan , Bing Geng , An Fu , Muhan Zeng , Ailun Yu , Jichuan Ji , Jingyang Zhao , Yuenan Guo , Qianxiang Wang

Large language models have shown good potential in supporting software development tasks. This is why more and more developers turn to LLMs (e.g. ChatGPT) to support them in fixing their buggy code. While this can save time and effort, many…

Software Engineering · Computer Science 2024-09-06 Yacine Majdoub , Eya Ben Charrada

Bug reproduction is a critical developer activity that is also challenging to automate, as bug reports are often in natural language and thus can be difficult to transform to test cases consistently. As a result, existing techniques mostly…

Software Engineering · Computer Science 2023-11-10 Sungmin Kang , Juyeon Yoon , Nargiz Askarbekkyzy , Shin Yoo

Large Language Models (LLMs) have recently shown strong reasoning and generalization capabilities, motivating their use as decision-making policies in complex environments. StarCraft II (SC2), with its massive state-action space and partial…

Artificial Intelligence · Computer Science 2026-02-17 Yixin Zhang , Ziyi Wang , Yiming Rong , Haoxi Wang , Jinling Jiang , Shuang Xu , Haoran Wu , Shiyu Zhou , Bo Xu

We present Mify-Coder, a 2.5B-parameter code model trained on 4.2T tokens using a compute-optimal strategy built on the Mify-2.5B foundation model. Mify-Coder achieves comparable accuracy and safety while significantly outperforming much…

Software Engineering · Computer Science 2026-01-01 Abhinav Parmar , Abhisek Panigrahi , Abhishek Kumar Dwivedi , Abhishek Bhattacharya , Adarsh Ramachandra , Aditya Choudhary , Aditya Garg , Aditya Raj , Alankrit Bhatt , Alpesh Yadav , Anant Vishnu , Ananthu Pillai , Ankush Kumar , Aryan Patnaik , Aswatha Narayanan S , Avanish Raj Singh , Bhavya Shree Gadda , Brijesh Pankajbhai Kachhadiya , Buggala Jahnavi , Chidurala Nithin Krishna , Chintan Shah , Chunduru Akshaya , Debarshi Banerjee , Debrup Dey , Deepa R. , Deepika B G , Faiz ur Rahman , Gagan Gayari , Gudhi Jagadeesh Kumar Naidu , Gursimar Singh , Harshal Tyagi , Harshini K , James Mani Vathalloor , Jayarama Nettar , Jayashree Gajjam , Joe Walter Sugil George , Kamalakara Sri Krishna Tadepalli , Kamalkumar Rathinasamy , Karan Chaurasia , Karthikeyan S , Kashish Arora , Kaushal Desai , Khushboo Buwade , Kiran Manjrekar , Malikireddy Venkata Sai Likhitha , Manjunath A , Mitali Mahavir Bedmutha , Mohammed Rafee Tarafdar , Nikhil Tiwari , Nikitha K Gigi , Pavan Ravikumar , Pendyala Swarnanjali , Piyush Anand , Prakash Chandrasekar , Prasanna Bhalchandra Gawade , Prasanth Sivan , Preeti Khurana , Priyanshi Babbar , Rajab Ali Mondal , Rajesh Kumar Vissapragada , Rajeshwari Ganesan , Rajeswari Koppisetti , Ramjee R. , Ramkumar Thiruppathisamy , Rani G. S. , S Reka , Samarth Gupta , Sandeep Reddy Kothakota , Sarathy K , Sathyanarayana Sampath Kumar , Saurabh Kumar , Shashank Khasare , Shenbaga Devi Venkatesh Kumar , Shiva Rama Krishna Parvatham , Shoeb Shaikh , Shrishanmathi A , Shubham Pathak , Sree Samhita Koppaka , Sreenivasa Raghavan K S , Sreeram Venkatasubramanian , Suprabha Desai Bojja , Swetha R , Syed Ahmed , Chinmai Harshitha Thota , Tushar Yadav , Veeravelly Kusumitha , V V S S Prasanth Patnaik , Vidya Sri Sesetti , Vijayakeerthi K , Vikram Raj Bakshi , Vinay K K , Vinoth Kumar Loganathan , Vipin Tiwari , Vivek Kumar Shrivastav , V Venkata Sri Datta Charan , Wasim Akhtar Khan

In this technical report, we present Skywork-13B, a family of large language models (LLMs) trained on a corpus of over 3.2 trillion tokens drawn from both English and Chinese texts. This bilingual foundation model is the most extensively…

Multilingual programming, which involves using multiple programming languages (PLs) in a single project, is increasingly common due to its benefits. However, it introduces cross-language bugs (CLBs), which arise from interactions between…

Software Engineering · Computer Science 2026-04-22 Zengyang Li , Yimeng Li , Binbin Huang , Peng Liang , Ran Mo , Hui Liu , Yutao Ma

Recently, pre-trained large language models (LLMs) have shown impressive abilities in generating codes from natural language descriptions, repairing buggy codes, translating codes between languages, and retrieving relevant code segments.…

Computation and Language · Computer Science 2023-11-07 Mohammad Abdullah Matin Khan , M Saiful Bari , Xuan Long Do , Weishi Wang , Md Rizwan Parvez , Shafiq Joty

Large Language Models (LLMs) exhibit remarkable code generation capabilities but falter when adapting to frequent updates in external library APIs. This critical limitation, stemming from reliance on outdated API knowledge from their…

Computation and Language · Computer Science 2025-11-25 Haoze Wu , Yunzhi Yao , Wenhao Yu , Ningyu Zhang

Large language models (LLMs) have shown remarkable progress in code generation, but their generated code often suffers from inefficiency, resulting in longer execution times and higher memory consumption. To address this issue, we propose…

Software Engineering · Computer Science 2025-05-13 Dong Huang , Jianbo Dai , Han Weng , Puzhen Wu , Yuhao Qing , Heming Cui , Zhijiang Guo , Jie M. Zhang

Diffusion-based language models (DLLMs) offer non-sequential, block-wise generation and richer data reuse compared to autoregressive (AR) models, but existing code DLLMs still lag behind strong AR baselines under comparable budgets. We…

Computation and Language · Computer Science 2026-01-26 Chenghao Fan , Wen Heng , Bo Li , Sichen Liu , Yuxuan Song , Jing Su , Xiaoye Qu , Kai Shen , Wei Wei

Code large language models (Code LLMs) have made significant progress in code generation by translating natural language descriptions into functional code; however, real-world applications often demand stricter adherence to detailed…

Computation and Language · Computer Science 2025-08-04 Jian Yang , Wei Zhang , Shukai Liu , Linzheng Chai , Yingshui Tan , Jiaheng Liu , Ge Zhang , Wangchunshu Zhou , Guanglin Niu , Zhoujun Li , Binyuan Hui , Junyang Lin

Large language models (LLMs) have showcased profound capabilities in language understanding and generation, facilitating a wide array of applications. However, there is a notable paucity of detailed, open-sourced methodologies on…

We explore the novel application of Large Language Models to code optimization. We present a 7B-parameter transformer model trained from scratch to optimize LLVM assembly for code size. The model takes as input unoptimized assembly and…

We introduce the Falcon series: 7B, 40B, and 180B parameters causal decoder-only models trained on a diverse high-quality corpora predominantly assembled from web data. The largest model, Falcon-180B, has been trained on over 3.5 trillion…

In recent years, large language models (LLMs) have emerged as powerful tools with potential applications in various fields, including software engineering. Within the scope of this research, we evaluate five different state-of-the-art LLMs…

Computation and Language · Computer Science 2024-09-09 Luis Mayer , Christian Heumann , Matthias Aßenmacher
‹ Prev 1 3 4 5 6 7 10 Next ›