English
Related papers

Related papers: Skill over Scale: The Case for Medium, Domain-Spec…

200 papers

Pretrained transformer models have achieved state-of-the-art results in many tasks and benchmarks recently. Many state-of-the-art Language Models (LMs), however, do not scale well above the threshold of 512 input tokens. In specialized…

Computation and Language · Computer Science 2022-12-01 Joel Niklaus , Daniele Giofré

Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks. However, at some point further model increases become harder due to GPU/TPU memory limitations and longer…

Computation and Language · Computer Science 2020-02-11 Zhenzhong Lan , Mingda Chen , Sebastian Goodman , Kevin Gimpel , Piyush Sharma , Radu Soricut

There has been an influx of biomedical domain-specific language models, showing language models pre-trained on biomedical text perform better on biomedical domain benchmarks than those trained on general domain text corpora such as…

Computation and Language · Computer Science 2020-10-15 Hoo-Chang Shin , Yang Zhang , Evelina Bakhturina , Raul Puri , Mostofa Patwary , Mohammad Shoeybi , Raghav Mani

Recent works have demonstrated the effectiveness of self-alignment in which a large language model is aligned to follow general instructions using instructional data generated from the model itself starting from a handful of human-written…

Computation and Language · Computer Science 2024-06-07 Junmo Kang , Hongyin Luo , Yada Zhu , Jacob Hansen , James Glass , David Cox , Alan Ritter , Rogerio Feris , Leonid Karlinsky

There are many cases where LLMs are used for specific tasks in a single domain. These usually require less general, but more domain-specific knowledge. Highly capable, general-purpose state-of-the-art language models like GPT-4 or…

Machine Learning · Computer Science 2024-07-30 Tobias Kerner

Large Language Model (LLM) pre-training exhausts an ever growing compute budget, yet recent research has demonstrated that careful document selection enables comparable model quality with only a fraction of the FLOPs. Inspired by efforts…

Computation and Language · Computer Science 2024-06-10 Xiang Kong , Tom Gunter , Ruoming Pang

Skill Extraction (SE) is an important and widely-studied task useful to gain insights into labor market dynamics. However, there is a lacuna of datasets and annotation guidelines; available datasets are few and contain crowd-sourced labels…

Computation and Language · Computer Science 2022-04-28 Mike Zhang , Kristian Nørgaard Jensen , Sif Dam Sonniks , Barbara Plank

The burgeoning field of Large Language Models (LLMs), exemplified by sophisticated models like OpenAI's ChatGPT, represents a significant advancement in artificial intelligence. These models, however, bring forth substantial challenges in…

This study aims to guide language model selection by investigating: 1) the necessity of finetuning versus zero-shot usage, 2) the benefits of domain-adjacent versus generic pretrained models, 3) the value of further domain-specific…

Computation and Language · Computer Science 2025-09-25 Lovedeep Gondara , Jonathan Simkin , Graham Sayle , Shebnum Devji , Gregory Arbour , Raymond Ng

The grand goal of AI research, and particularly Self Supervised Learning (SSL), is to produce systems that can successfully solve any possible task. In contrast, current evaluation methods available to AI researchers typically rely on a…

Machine Learning · Computer Science 2025-10-22 Niket Patel , Randall Balestriero

For tasks like code synthesis from natural language, code retrieval, and code summarization, data-driven models have shown great promise. However, creating these models require parallel data between natural language (NL) and code with…

Computation and Language · Computer Science 2018-05-24 Pengcheng Yin , Bowen Deng , Edgar Chen , Bogdan Vasilescu , Graham Neubig

There is growing evidence that pretrained language models improve task-specific fine-tuning not just for the languages seen in pretraining, but also for new languages and even non-linguistic data. What is the nature of this surprising…

Computation and Language · Computer Science 2021-04-20 Zhengxuan Wu , Nelson F. Liu , Christopher Potts

Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform…

Computation and Language · Computer Science 2023-06-13 Aarohi Srivastava , Abhinav Rastogi , Abhishek Rao , Abu Awal Md Shoeb , Abubakar Abid , Adam Fisch , Adam R. Brown , Adam Santoro , Aditya Gupta , Adrià Garriga-Alonso , Agnieszka Kluska , Aitor Lewkowycz , Akshat Agarwal , Alethea Power , Alex Ray , Alex Warstadt , Alexander W. Kocurek , Ali Safaya , Ali Tazarv , Alice Xiang , Alicia Parrish , Allen Nie , Aman Hussain , Amanda Askell , Amanda Dsouza , Ambrose Slone , Ameet Rahane , Anantharaman S. Iyer , Anders Andreassen , Andrea Madotto , Andrea Santilli , Andreas Stuhlmüller , Andrew Dai , Andrew La , Andrew Lampinen , Andy Zou , Angela Jiang , Angelica Chen , Anh Vuong , Animesh Gupta , Anna Gottardi , Antonio Norelli , Anu Venkatesh , Arash Gholamidavoodi , Arfa Tabassum , Arul Menezes , Arun Kirubarajan , Asher Mullokandov , Ashish Sabharwal , Austin Herrick , Avia Efrat , Aykut Erdem , Ayla Karakaş , B. Ryan Roberts , Bao Sheng Loe , Barret Zoph , Bartłomiej Bojanowski , Batuhan Özyurt , Behnam Hedayatnia , Behnam Neyshabur , Benjamin Inden , Benno Stein , Berk Ekmekci , Bill Yuchen Lin , Blake Howald , Bryan Orinion , Cameron Diao , Cameron Dour , Catherine Stinson , Cedrick Argueta , César Ferri Ramírez , Chandan Singh , Charles Rathkopf , Chenlin Meng , Chitta Baral , Chiyu Wu , Chris Callison-Burch , Chris Waites , Christian Voigt , Christopher D. Manning , Christopher Potts , Cindy Ramirez , Clara E. Rivera , Clemencia Siro , Colin Raffel , Courtney Ashcraft , Cristina Garbacea , Damien Sileo , Dan Garrette , Dan Hendrycks , Dan Kilman , Dan Roth , Daniel Freeman , Daniel Khashabi , Daniel Levy , Daniel Moseguí González , Danielle Perszyk , Danny Hernandez , Danqi Chen , Daphne Ippolito , Dar Gilboa , David Dohan , David Drakard , David Jurgens , Debajyoti Datta , Deep Ganguli , Denis Emelin , Denis Kleyko , Deniz Yuret , Derek Chen , Derek Tam , Dieuwke Hupkes , Diganta Misra , Dilyar Buzan , Dimitri Coelho Mollo , Diyi Yang , Dong-Ho Lee , Dylan Schrader , Ekaterina Shutova , Ekin Dogus Cubuk , Elad Segal , Eleanor Hagerman , Elizabeth Barnes , Elizabeth Donoway , Ellie Pavlick , Emanuele Rodola , Emma Lam , Eric Chu , Eric Tang , Erkut Erdem , Ernie Chang , Ethan A. Chi , Ethan Dyer , Ethan Jerzak , Ethan Kim , Eunice Engefu Manyasi , Evgenii Zheltonozhskii , Fanyue Xia , Fatemeh Siar , Fernando Martínez-Plumed , Francesca Happé , Francois Chollet , Frieda Rong , Gaurav Mishra , Genta Indra Winata , Gerard de Melo , Germán Kruszewski , Giambattista Parascandolo , Giorgio Mariani , Gloria Wang , Gonzalo Jaimovitch-López , Gregor Betz , Guy Gur-Ari , Hana Galijasevic , Hannah Kim , Hannah Rashkin , Hannaneh Hajishirzi , Harsh Mehta , Hayden Bogar , Henry Shevlin , Hinrich Schütze , Hiromu Yakura , Hongming Zhang , Hugh Mee Wong , Ian Ng , Isaac Noble , Jaap Jumelet , Jack Geissinger , Jackson Kernion , Jacob Hilton , Jaehoon Lee , Jaime Fernández Fisac , James B. Simon , James Koppel , James Zheng , James Zou , Jan Kocoń , Jana Thompson , Janelle Wingfield , Jared Kaplan , Jarema Radom , Jascha Sohl-Dickstein , Jason Phang , Jason Wei , Jason Yosinski , Jekaterina Novikova , Jelle Bosscher , Jennifer Marsh , Jeremy Kim , Jeroen Taal , Jesse Engel , Jesujoba Alabi , Jiacheng Xu , Jiaming Song , Jillian Tang , Joan Waweru , John Burden , John Miller , John U. Balis , Jonathan Batchelder , Jonathan Berant , Jörg Frohberg , Jos Rozen , Jose Hernandez-Orallo , Joseph Boudeman , Joseph Guerr , Joseph Jones , Joshua B. Tenenbaum , Joshua S. Rule , Joyce Chua , Kamil Kanclerz , Karen Livescu , Karl Krauth , Karthik Gopalakrishnan , Katerina Ignatyeva , Katja Markert , Kaustubh D. Dhole , Kevin Gimpel , Kevin Omondi , Kory Mathewson , Kristen Chiafullo , Ksenia Shkaruta , Kumar Shridhar , Kyle McDonell , Kyle Richardson , Laria Reynolds , Leo Gao , Li Zhang , Liam Dugan , Lianhui Qin , Lidia Contreras-Ochando , Louis-Philippe Morency , Luca Moschella , Lucas Lam , Lucy Noble , Ludwig Schmidt , Luheng He , Luis Oliveros Colón , Luke Metz , Lütfi Kerem Şenel , Maarten Bosma , Maarten Sap , Maartje ter Hoeve , Maheen Farooqi , Manaal Faruqui , Mantas Mazeika , Marco Baturan , Marco Marelli , Marco Maru , Maria Jose Ramírez Quintana , Marie Tolkiehn , Mario Giulianelli , Martha Lewis , Martin Potthast , Matthew L. Leavitt , Matthias Hagen , Mátyás Schubert , Medina Orduna Baitemirova , Melody Arnaud , Melvin McElrath , Michael A. Yee , Michael Cohen , Michael Gu , Michael Ivanitskiy , Michael Starritt , Michael Strube , Michał Swędrowski , Michele Bevilacqua , Michihiro Yasunaga , Mihir Kale , Mike Cain , Mimee Xu , Mirac Suzgun , Mitch Walker , Mo Tiwari , Mohit Bansal , Moin Aminnaseri , Mor Geva , Mozhdeh Gheini , Mukund Varma T , Nanyun Peng , Nathan A. Chi , Nayeon Lee , Neta Gur-Ari Krakover , Nicholas Cameron , Nicholas Roberts , Nick Doiron , Nicole Martinez , Nikita Nangia , Niklas Deckers , Niklas Muennighoff , Nitish Shirish Keskar , Niveditha S. Iyer , Noah Constant , Noah Fiedel , Nuan Wen , Oliver Zhang , Omar Agha , Omar Elbaghdadi , Omer Levy , Owain Evans , Pablo Antonio Moreno Casares , Parth Doshi , Pascale Fung , Paul Pu Liang , Paul Vicol , Pegah Alipoormolabashi , Peiyuan Liao , Percy Liang , Peter Chang , Peter Eckersley , Phu Mon Htut , Pinyu Hwang , Piotr Miłkowski , Piyush Patil , Pouya Pezeshkpour , Priti Oli , Qiaozhu Mei , Qing Lyu , Qinlang Chen , Rabin Banjade , Rachel Etta Rudolph , Raefer Gabriel , Rahel Habacker , Ramon Risco , Raphaël Millière , Rhythm Garg , Richard Barnes , Rif A. Saurous , Riku Arakawa , Robbe Raymaekers , Robert Frank , Rohan Sikand , Roman Novak , Roman Sitelew , Ronan LeBras , Rosanne Liu , Rowan Jacobs , Rui Zhang , Ruslan Salakhutdinov , Ryan Chi , Ryan Lee , Ryan Stovall , Ryan Teehan , Rylan Yang , Sahib Singh , Saif M. Mohammad , Sajant Anand , Sam Dillavou , Sam Shleifer , Sam Wiseman , Samuel Gruetter , Samuel R. Bowman , Samuel S. Schoenholz , Sanghyun Han , Sanjeev Kwatra , Sarah A. Rous , Sarik Ghazarian , Sayan Ghosh , Sean Casey , Sebastian Bischoff , Sebastian Gehrmann , Sebastian Schuster , Sepideh Sadeghi , Shadi Hamdan , Sharon Zhou , Shashank Srivastava , Sherry Shi , Shikhar Singh , Shima Asaadi , Shixiang Shane Gu , Shubh Pachchigar , Shubham Toshniwal , Shyam Upadhyay , Shyamolima , Debnath , Siamak Shakeri , Simon Thormeyer , Simone Melzi , Siva Reddy , Sneha Priscilla Makini , Soo-Hwan Lee , Spencer Torene , Sriharsha Hatwar , Stanislas Dehaene , Stefan Divic , Stefano Ermon , Stella Biderman , Stephanie Lin , Stephen Prasad , Steven T. Piantadosi , Stuart M. Shieber , Summer Misherghi , Svetlana Kiritchenko , Swaroop Mishra , Tal Linzen , Tal Schuster , Tao Li , Tao Yu , Tariq Ali , Tatsu Hashimoto , Te-Lin Wu , Théo Desbordes , Theodore Rothschild , Thomas Phan , Tianle Wang , Tiberius Nkinyili , Timo Schick , Timofei Kornev , Titus Tunduny , Tobias Gerstenberg , Trenton Chang , Trishala Neeraj , Tushar Khot , Tyler Shultz , Uri Shaham , Vedant Misra , Vera Demberg , Victoria Nyamai , Vikas Raunak , Vinay Ramasesh , Vinay Uday Prabhu , Vishakh Padmakumar , Vivek Srikumar , William Fedus , William Saunders , William Zhang , Wout Vossen , Xiang Ren , Xiaoyu Tong , Xinran Zhao , Xinyi Wu , Xudong Shen , Yadollah Yaghoobzadeh , Yair Lakretz , Yangqiu Song , Yasaman Bahri , Yejin Choi , Yichi Yang , Yiding Hao , Yifu Chen , Yonatan Belinkov , Yu Hou , Yufang Hou , Yuntao Bai , Zachary Seid , Zhuoye Zhao , Zijian Wang , Zijie J. Wang , Zirui Wang , Ziyi Wu

With the urgent demand for generalized deep models, many pre-trained big models are proposed, such as BERT, ViT, GPT, etc. Inspired by the success of these models in single domains (like computer vision and natural language processing), the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Xiao Wang , Guangyao Chen , Guangwu Qian , Pengcheng Gao , Xiao-Yong Wei , Yaowei Wang , Yonghong Tian , Wen Gao

Pre-trained machine learning (ML) models have shown great performance for a wide range of applications, in particular in natural language processing (NLP) and computer vision (CV). Here, we study how pre-training could be used for…

Machine Learning · Computer Science 2024-01-05 Shashank Subramanian , Peter Harrington , Kurt Keutzer , Wahid Bhimji , Dmitriy Morozov , Michael Mahoney , Amir Gholami

Large pre-trained language models have recently gained significant traction due to their improved performance on various down-stream tasks like text classification and question answering, requiring only few epochs of fine-tuning. However,…

Computation and Language · Computer Science 2023-09-01 Souvik Kundu , Sharath Nittur Sridhar , Maciej Szankin , Sairam Sundaresan

The exponential growth of online textual content across diverse domains has necessitated advanced methods for automated text classification. Large Language Models (LLMs) based on transformer architectures have shown significant success in…

Computation and Language · Computer Science 2025-09-09 Zhyar Rzgar K Rostam , Gábor Kertész

Following its success for vision and text, the "foundation model" (FM) paradigm -- pretraining large models on massive data, then fine-tuning on target tasks -- has rapidly expanded to domains in the sciences, engineering, healthcare, and…

Machine Learning · Computer Science 2025-03-24 Zongzhe Xu , Ritvik Gupta , Wenduo Cheng , Alexander Shen , Junhong Shen , Ameet Talwalkar , Mikhail Khodak

Language is essentially a complex, intricate system of human expressions governed by grammatical rules. It poses a significant challenge to develop capable AI algorithms for comprehending and grasping a language. As a major approach,…

Domain-adaptive pretraining (DAPT) offers a practical path to specializing large language models for high-value domains without full retraining. We conduct an early-stage scaling-law analysis of continued pretraining on U.S. SEC filings,…

Machine Learning · Computer Science 2025-12-16 Jesse Ponnock