English
Related papers

Related papers: SuperMerge: An Approach For Gradient-Based Model M…

200 papers

The drastic increase in language models' parameters has led to a new trend of deploying models in cloud servers, raising growing concerns about private inference for Transformer-based models. Existing two-party privacy-preserving…

Computation and Language · Computer Science 2023-12-12 Zi Liang , Pinghui Wang , Ruofei Zhang , Nuo Xu , Lifeng Xing , Shuo Zhang

With the emerging trend of GPT models, we have established a framework called AutoML-GPT that integrates a comprehensive set of tools and libraries. This framework grants users access to a wide range of data preprocessing techniques,…

Machine Learning · Computer Science 2023-09-06 Yun-Da Tsai , Yu-Che Tsai , Bo-Wei Huang , Chun-Pai Yang , Shou-De Lin

Fine-tuning large language models for different tasks can be costly and inefficient, and even methods that reduce the number of tuned parameters still require full gradient-based optimization. We propose HyperTuning, a novel approach to…

Computation and Language · Computer Science 2022-11-23 Jason Phang , Yi Mao , Pengcheng He , Weizhu Chen

Large language models (LLMs) have revolutionized natural language processing (NLP) by excelling at understanding and generating human-like text. However, their widespread deployment can be prohibitively expensive. SortedNet is a recent…

Computation and Language · Computer Science 2024-02-12 Parsa Kavehzadeh , Mojtaba Valipour , Marzieh Tahaei , Ali Ghodsi , Boxing Chen , Mehdi Rezagholizadeh

With the surge of ChatGPT,the use of large models has significantly increased,rapidly rising to prominence across the industry and sweeping across the internet. This article is a comprehensive review of fine-tuning methods for large models.…

Machine Learning · Computer Science 2024-04-16 Benjue Weng

In collaborative software development, program merging is the mechanism to integrate changes from multiple programmers. Merge algorithms in modern version control systems report a conflict when changes interfere textually. Merge conflicts…

Software Engineering · Computer Science 2021-09-08 Elizabeth Dinella , Todd Mytkowicz , Alexey Svyatkovskiy , Christian Bird , Mayur Naik , Shuvendu K. Lahiri

Ensembles of generative large language models (LLMs) are a promising way to compensate for individual model limitations, integrating the strengths of different LLMs. Existing LLM ensemble methods, however, face limitations such as…

Computation and Language · Computer Science 2026-03-09 Bo Lv , Nayu Liu , Chen Tang , Xin Liu , Yue Yu , Ping Luo

In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised and multilingual-capable model, with support for 23…

Computation and Language · Computer Science 2025-04-15 Team Cohere , : , Aakanksha , Arash Ahmadian , Marwan Ahmed , Jay Alammar , Milad Alizadeh , Yazeed Alnumay , Sophia Althammer , Arkady Arkhangorodsky , Viraat Aryabumi , Dennis Aumiller , Raphaël Avalos , Zahara Aviv , Sammie Bae , Saurabh Baji , Alexandre Barbet , Max Bartolo , Björn Bebensee , Neeral Beladia , Walter Beller-Morales , Alexandre Bérard , Andrew Berneshawi , Anna Bialas , Phil Blunsom , Matt Bobkin , Adi Bongale , Sam Braun , Maxime Brunet , Samuel Cahyawijaya , David Cairuz , Jon Ander Campos , Cassie Cao , Kris Cao , Roman Castagné , Julián Cendrero , Leila Chan Currie , Yash Chandak , Diane Chang , Giannis Chatziveroglou , Hongyu Chen , Claire Cheng , Alexis Chevalier , Justin T. Chiu , Eugene Cho , Eugene Choi , Eujeong Choi , Tim Chung , Volkan Cirik , Ana Cismaru , Pierre Clavier , Henry Conklin , Lucas Crawhall-Stein , Devon Crouse , Andres Felipe Cruz-Salinas , Ben Cyrus , Daniel D'souza , Hugo Dalla-Torre , John Dang , William Darling , Omar Darwiche Domingues , Saurabh Dash , Antoine Debugne , Théo Dehaze , Shaan Desai , Joan Devassy , Rishit Dholakia , Kyle Duffy , Ali Edalati , Ace Eldeib , Abdullah Elkady , Sarah Elsharkawy , Irem Ergün , Beyza Ermis , Marzieh Fadaee , Boyu Fan , Lucas Fayoux , Yannis Flet-Berliac , Nick Frosst , Matthias Gallé , Wojciech Galuba , Utsav Garg , Matthieu Geist , Mohammad Gheshlaghi Azar , Ellen Gilsenan-McMahon , Seraphina Goldfarb-Tarrant , Tomas Goldsack , Aidan Gomez , Victor Machado Gonzaga , Nithya Govindarajan , Manoj Govindassamy , Nathan Grinsztajn , Nikolas Gritsch , Patrick Gu , Shangmin Guo , Kilian Haefeli , Rod Hajjar , Tim Hawes , Jingyi He , Sebastian Hofstätter , Sungjin Hong , Sara Hooker , Tom Hosking , Stephanie Howe , Eric Hu , Renjie Huang , Hemant Jain , Ritika Jain , Nick Jakobi , Madeline Jenkins , JJ Jordan , Dhruti Joshi , Jason Jung , Trushant Kalyanpur , Siddhartha Rao Kamalakara , Julia Kedrzycki , Gokce Keskin , Edward Kim , Joon Kim , Wei-Yin Ko , Tom Kocmi , Michael Kozakov , Wojciech Kryściński , Arnav Kumar Jain , Komal Kumar Teru , Sander Land , Michael Lasby , Olivia Lasche , Justin Lee , Patrick Lewis , Jeffrey Li , Jonathan Li , Hangyu Lin , Acyr Locatelli , Kevin Luong , Raymond Ma , Lukáš Mach , Marina Machado , Joanne Magbitang , Brenda Malacara Lopez , Aryan Mann , Kelly Marchisio , Olivia Markham , Alexandre Matton , Alex McKinney , Dominic McLoughlin , Jozef Mokry , Adrien Morisot , Autumn Moulder , Harry Moynehan , Maximilian Mozes , Vivek Muppalla , Lidiya Murakhovska , Hemangani Nagarajan , Alekhya Nandula , Hisham Nasir , Shauna Nehra , Josh Netto-Rosen , Daniel Ohashi , James Owers-Bardsley , Jason Ozuzu , Dennis Padilla , Gloria Park , Sam Passaglia , Jeremy Pekmez , Laura Penstone , Aleksandra Piktus , Case Ploeg , Andrew Poulton , Youran Qi , Shubha Raghvendra , Miguel Ramos , Ekagra Ranjan , Pierre Richemond , Cécile Robert-Michon , Aurélien Rodriguez , Sudip Roy , Sebastian Ruder , Laura Ruis , Louise Rust , Anubhav Sachan , Alejandro Salamanca , Kailash Karthik Saravanakumar , Isha Satyakam , Alice Schoenauer Sebag , Priyanka Sen , Sholeh Sepehri , Preethi Seshadri , Ye Shen , Tom Sherborne , Sylvie Shang Shi , Sanal Shivaprasad , Vladyslav Shmyhlo , Anirudh Shrinivason , Inna Shteinbuk , Amir Shukayev , Mathieu Simard , Ella Snyder , Ava Spataru , Victoria Spooner , Trisha Starostina , Florian Strub , Yixuan Su , Jimin Sun , Dwarak Talupuru , Eugene Tarassov , Elena Tommasone , Jennifer Tracey , Billy Trend , Evren Tumer , Ahmet Üstün , Bharat Venkitesh , David Venuto , Pat Verga , Maxime Voisin , Alex Wang , Donglu Wang , Shijian Wang , Edmond Wen , Naomi White , Jesse Willman , Marysia Winkels , Chen Xia , Jessica Xie , Minjie Xu , Bowen Yang , Tan Yi-Chern , Ivan Zhang , Zhenyu Zhao , Zhoujie Zhao

Model merging, a method that combines the parameters and embeddings of multiple fine-tuned large language models (LLMs), offers a promising approach to enhance model performance across various tasks while maintaining computational…

Computation and Language · Computer Science 2025-11-10 Amin Heyrani Nobari , Kaveh Alim , Ali ArjomandBigdeli , Akash Srivastava , Faez Ahmed , Navid Azizan

Model merging enables the combination of multiple specialized expert models into a single model capable of performing multiple tasks. However, the benefits of merging an increasing amount of specialized experts generally lead to diminishing…

Machine Learning · Computer Science 2025-12-23 Ronald Skorobogat , Karsten Roth , Mariana-Iuliana Georgescu

Large pre-trained models (LPMs), such as large language models, have become ubiquitous and are employed in many applications. These models are often adapted to a desired domain or downstream task through a fine-tuning stage. This paper…

Machine Learning · Computer Science 2024-10-08 Juan Pablo Muñoz , Jinjie Yuan , Nilesh Jain

Large language models (LLMs) have significantly advanced various natural language processing (NLP) tasks. Recent research indicates that moderately-sized LLMs often outperform larger ones after task-specific fine-tuning. This study focuses…

Computation and Language · Computer Science 2024-10-14 Minghao Wu , Thuy-Trang Vu , Lizhen Qu , George Foster , Gholamreza Haffari

Expanding Large Language Models~(LLMs) to new languages is a costly endeavor, demanding extensive Continued Pre-Training~(CPT) and data-intensive alignment. While recent data-free merging techniques attempt to bypass alignment by fusing a…

Computation and Language · Computer Science 2026-05-19 Hao Zhou , Tianhao Li , Zhijun Wang , Shuaijie She , Linjuan Wu , Hao-Ran Wei , Baosong Yang , Jiajun Chen , Shujian Huang

As Large Language Models (LLMs) excel across tasks and specialized domains, scaling LLMs based on existing models has garnered significant attention, which faces the challenge of decreasing performance when combining disparate models.…

We propose a novel approach to enhancing the performance and efficiency of large language models (LLMs) by combining domain prompt routing with domain-specialized models. We introduce a system that utilizes a BERT-based router to direct…

Computation and Language · Computer Science 2024-10-11 Toby Simonds , Kemal Kurniawan , Jey Han Lau

Transfer learning - i.e., further fine-tuning a pre-trained model on a downstream task - can confer significant advantages, including improved downstream performance, faster convergence, and better sample efficiency. These advantages have…

Machine Learning · Computer Science 2023-10-30 Prateek Yadav , Derek Tam , Leshem Choshen , Colin Raffel , Mohit Bansal

Fine-tuning of Large Language Models (LLMs) for downstream tasks, performed on domain-specific data has shown significant promise. However, commercial use of such LLMs is limited by the high computational cost required for their deployment…

Computation and Language · Computer Science 2025-03-06 Boris Nazarov , Darya Frolova , Yackov Lubarsky , Alexei Gaissinski , Pavel Kisilev

The transition from System 1 to System 2 reasoning in large language models (LLMs) has marked significant advancements in handling complex tasks through deliberate, iterative thinking. However, this progress often comes at the cost of…

Computation and Language · Computer Science 2025-05-26 Han Wu , Yuxuan Yao , Shuqi Liu , Zehua Liu , Xiaojin Fu , Xiongwei Han , Xing Li , Hui-Ling Zhen , Tao Zhong , Mingxuan Yuan

Pretrained Large Language Models (LLMs) have achieved remarkable success across diverse domains, with education and research emerging as particularly impactful areas. Among current state-of-the-art LLMs, ChatGPT and DeepSeek exhibit strong…

Artificial Intelligence · Computer Science 2025-12-10 Md Mostafizer Rahman , Ariful Islam Shiplu , Md Faizul Ibne Amin , Yutaka Watanobe , Lu Peng

From a multi-model compression perspective, model merging enables memory-efficient serving of multiple models fine-tuned from the same base, but suffers from degraded performance due to interference among their task-specific parameter…

Machine Learning · Computer Science 2025-05-19 Hangyu Zhou , Aaron Gokaslan , Volodymyr Kuleshov , Bharath Hariharan
‹ Prev 1 8 9 10 Next ›