English
Related papers

Related papers: React-ing to Grace Hopper 200: Five Open-Weights C…

200 papers

Executable software engineering data is valuable for training SWE agents, but scaling it remains difficult for two reasons: only a small fraction of real repository changes yield verifiable, high-signal task instances, and naively building…

Software Engineering · Computer Science 2026-03-24 Jiarong Liang , Zhiheng Lyu , Zijie Liu , Xiangchao Chen , Ping Nie , Kai Zou , Wenhu Chen

Large Language Models (LLMs) face significant deployment challenges due to their substantial resource requirements. While low-bit quantized weights can reduce memory usage and improve inference efficiency, current hardware lacks native…

Machine Learning · Computer Science 2025-06-10 Pengxiang Zhao , Xiaoming Yuan

Modern day Language Models see extensive use in text classification, yet this comes at significant computational cost. Compute-effective classification models are needed for low-resource environments, most notably on edge devices. We…

Machine Learning · Computer Science 2024-11-22 Stan Loosmore , Alexander Titus

Most vulnerability detection studies focus on datasets of vulnerabilities in C/C++ code, offering limited language diversity. Thus, the effectiveness of deep learning methods, including large language models (LLMs), in detecting software…

Software Engineering · Computer Science 2026-02-18 Kohei Dozono , Tiago Espinha Gasiba , Andrea Stocco

Deep learning (DL)-based code completion tools have transformed software development by enabling advanced code generation. These tools leverage models trained on vast amounts of code from numerous repositories, capturing general coding…

Software Engineering · Computer Science 2025-03-19 Alessandro Giagnorio , Alberto Martin-Lopez , Gabriele Bavota

Large Language Models (LLMs) are widely used for code generation. However, commercial models like ChatGPT require significant computing power, which leads to high energy use and carbon emissions. This has raised concerns about their…

Software Engineering · Computer Science 2025-08-13 Humza Ashraf , Syed Muhammad Danish , Aris Leivadeas , Yazan Otoum , Zeeshan Sattar

Inorganic synthesis planning currently relies primarily on heuristic approaches or machine-learning models trained on limited datasets, which constrains its generality. We demonstrate that language models, without task-specific fine-tuning,…

Materials Science · Physics 2025-06-17 Thorben Prein , Elton Pan , Janik Jehkul , Steffen Weinmann , Elsa A. Olivetti , Jennifer L. M. Rupp

We propose SWE-Universe, a scalable and efficient framework for automatically constructing real-world software engineering (SWE) verifiable environments from GitHub pull requests (PRs). To overcome the prevalent challenges of automatic…

This paper illustrates an empirical study of the working efficiency of machine learning techniques in classifying code review text by semantic meaning. The code review comments from the source control repository in GitHub were extracted for…

Software Engineering · Computer Science 2025-08-25 Shadikur Rahman , Umme Ayman Koana , Hasibul Karim Shanto , Mahmuda Akter , Chitra Roy , Aras M. Ismael

What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks? The strongest vision-language models (VLMs) show such broad visual reasoning is within reach, but the recipe behind…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Gabriel Sarch , Linrong Cai , Qunzhong Wang , Haoyang Wu , Danqi Chen , Zhuang Liu

This study systematically evaluated the mathematical reasoning capabilities of Large Language Models (LLMs) using the 2026 Korean College Scholastic Ability Test (CSAT) Mathematics section, ensuring a completely contamination-free…

Computation and Language · Computer Science 2025-12-02 Goun Pyeon , Inbum Heo , Jeesu Jung , Taewook Hwang , Hyuk Namgoong , Hyein Seo , Yerim Han , Eunbin Kim , Hyeonseok Kang , Sangkeun Jung

Speculative decoding and quantization effectively accelerate memory-bound inference of large language models. Speculative decoding mitigates the memory bandwidth bottleneck by verifying multiple tokens within a single forward pass, which…

Computation and Language · Computer Science 2025-05-30 Yudi Zhang , Weilin Zhao , Xu Han , Tiejun Zhao , Wang Xu , Hailong Cao , Conghui Zhu

We present GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that supports both thinking and direct response modes. Through…

Computation and Language · Computer Science 2025-08-11 5 Team , Aohan Zeng , Xin Lv , Qinkai Zheng , Zhenyu Hou , Bin Chen , Chengxing Xie , Cunxiang Wang , Da Yin , Hao Zeng , Jiajie Zhang , Kedong Wang , Lucen Zhong , Mingdao Liu , Rui Lu , Shulin Cao , Xiaohan Zhang , Xuancheng Huang , Yao Wei , Yean Cheng , Yifan An , Yilin Niu , Yuanhao Wen , Yushi Bai , Zhengxiao Du , Zihan Wang , Zilin Zhu , Bohan Zhang , Bosi Wen , Bowen Wu , Bowen Xu , Can Huang , Casey Zhao , Changpeng Cai , Chao Yu , Chen Li , Chendi Ge , Chenghua Huang , Chenhui Zhang , Chenxi Xu , Chenzheng Zhu , Chuang Li , Congfeng Yin , Daoyan Lin , Dayong Yang , Dazhi Jiang , Ding Ai , Erle Zhu , Fei Wang , Gengzheng Pan , Guo Wang , Hailong Sun , Haitao Li , Haiyang Li , Haiyi Hu , Hanyu Zhang , Hao Peng , Hao Tai , Haoke Zhang , Haoran Wang , Haoyu Yang , He Liu , He Zhao , Hongwei Liu , Hongxi Yan , Huan Liu , Huilong Chen , Ji Li , Jiajing Zhao , Jiamin Ren , Jian Jiao , Jiani Zhao , Jianyang Yan , Jiaqi Wang , Jiayi Gui , Jiayue Zhao , Jie Liu , Jijie Li , Jing Li , Jing Lu , Jingsen Wang , Jingwei Yuan , Jingxuan Li , Jingzhao Du , Jinhua Du , Jinxin Liu , Junkai Zhi , Junli Gao , Ke Wang , Lekang Yang , Liang Xu , Lin Fan , Lindong Wu , Lintao Ding , Lu Wang , Man Zhang , Minghao Li , Minghuan Xu , Mingming Zhao , Mingshu Zhai , Pengfan Du , Qian Dong , Shangde Lei , Shangqing Tu , Shangtong Yang , Shaoyou Lu , Shijie Li , Shuang Li , Shuang-Li , Shuxun Yang , Sibo Yi , Tianshu Yu , Wei Tian , Weihan Wang , Wenbo Yu , Weng Lam Tam , Wenjie Liang , Wentao Liu , Xiao Wang , Xiaohan Jia , Xiaotao Gu , Xiaoying Ling , Xin Wang , Xing Fan , Xingru Pan , Xinyuan Zhang , Xinze Zhang , Xiuqing Fu , Xunkai Zhang , Yabo Xu , Yandong Wu , Yida Lu , Yidong Wang , Yilin Zhou , Yiming Pan , Ying Zhang , Yingli Wang , Yingru Li , Yinpei Su , Yipeng Geng , Yitong Zhu , Yongkun Yang , Yuhang Li , Yuhao Wu , Yujiang Li , Yunan Liu , Yunqing Wang , Yuntao Li , Yuxuan Zhang , Zezhen Liu , Zhen Yang , Zhengda Zhou , Zhongpei Qiao , Zhuoer Feng , Zhuorui Liu , Zichen Zhang , Zihan Wang , Zijun Yao , Zikang Wang , Ziqiang Liu , Ziwei Chai , Zixuan Li , Zuodong Zhao , Wenguang Chen , Jidong Zhai , Bin Xu , Minlie Huang , Hongning Wang , Juanzi Li , Yuxiao Dong , Jie Tang

Evaluating language models fairly is increasingly difficult as static benchmarks risk contamination by training data, obscuring whether models truly reason or recall. We introduce BeyondBench, an evaluation framework using algorithmic…

Computation and Language · Computer Science 2026-03-06 Gaurav Srivastava , Aafiya Hussain , Zhenyu Bi , Swastik Roy , Priya Pitre , Meng Lu , Morteza Ziyadi , Xuan Wang

Large language models (LLMs) have demonstrated remarkable capabilities in code generation across various domains. However, their effectiveness in generating simulation scripts for domain-specific environments like ns-3 remains…

Networking and Internet Architecture · Computer Science 2025-07-16 Tasnim Ahmed , Mirza Mohammad Azwad , Salimur Choudhury

We introduce SWE-ZERO to SWE-HERO, a two-stage SFT recipe that achieves state-of-the-art results on SWE-bench by distilling open-weight frontier LLMs. Our pipeline replaces resource-heavy dependencies with an evolutionary refinement…

Software Engineering · Computer Science 2026-05-07 Nikolai Ludwig , Wasi Uddin Ahmad , Somshubra Majumdar , Boris Ginsburg

Recent advances in mathematical problem-solving with language models (LMs) integrate chain-of-thought (CoT) reasoning and code execution to harness their complementary strengths. However, existing hybrid frameworks exhibit a critical…

Artificial Intelligence · Computer Science 2025-07-21 Haozhe Wang , Long Li , Chao Qu , Fengming Zhu , Weidi Xu , Wei Chu , Fangzhen Lin

Developers rely on code comments to document their work, track issues, and understand the source code. As such, comments provide valuable insights into developers' understanding of their code and describe their various intentions in writing…

Software Engineering · Computer Science 2025-07-03 Moritz Mock , Thomas Borsani , Giuseppe Di Fatta , Barbara Russo

Post-Training Quantization (PTQ) is crucial for efficient model deployment, yet its effectiveness on Ascend NPU remains under-explored compared to GPU architectures. This paper presents a case study of representative PTQ baselines applied…

Machine Learning · Computer Science 2026-02-23 Yuchen Luo , Fangyue Zhu , Ruining Zhou , Mingzhe Huang , Jian Zhu , Fanyu Fan , Wei Shao

Existing benchmarks for AI coding agents focus on isolated, single-issue tasks such as fixing a bug or adding a small feature. However, real-world software engineering is a long-horizon endeavor: developers interpret high-level…

Software Engineering · Computer Science 2026-05-25 Tue Le , Minh V. T. Thai , Dung Nguyen Manh , Huy Phan Nhat , Nghi D. Q. Bui