English
Related papers

Related papers: A Brain Cell Type Resource Created by Large Langua…

200 papers

Precision medicine has the potential to revolutionize healthcare, but much of the data for patients is locked away in unstructured free-text, limiting research and delivery of effective personalized treatments. Generating large annotated…

Computation and Language · Computer Science 2020-12-16 Nick Altieri , Briton Park , Mara Olson , John DeNero , Anobel Odisho , Bin Yu

Since FineWeb-Edu, data curation for LLM pretraining has predominantly relied on single scalar quality scores produced by small classifiers. A single score conflates multiple quality dimensions, prevents flexible filtering, and offers no…

Computation and Language · Computer Science 2026-02-20 Maximilian Idahl , Benedikt Droste , Björn Plüster , Jan Philipp Harries

Various machine learning approaches have gained significant popularity for the automated classification of educational text to identify indicators of learning engagement -- i.e. learning engagement classification (LEC). LEC can offer…

Computation and Language · Computer Science 2025-10-24 Shiqi Liu , Sannyuya Liu , Lele Sha , Zijie Zeng , Dragan Gasevic , Zhi Liu

Large language models (LLMs) offer strategy researchers powerful tools for annotating text at scale, but treating LLM-generated labels as deterministic overlooks substantial instability. Grounded in content analysis and generalizability…

Computers and Society · Computer Science 2026-01-21 Arnaldo Camuffo , Alfonso Gambardella , Saeid Kazemi , Jakub Malachowski , Abhinav Pandey

Large language models (LLMs) are increasingly positioned as scalable tools for annotating educational data, including classroom discourse, interaction logs, and qualitative learning artifacts. Their ability to rapidly summarize…

Artificial Intelligence · Computer Science 2026-03-17 Bakhtawar Ahtisham , Kirk Vanacore , Rene F. Kizilcec

Pre-training large language models on genomic sequences is a powerful approach for learning biologically meaningful representations. Masked language modeling (MLM) methods, such as DNABERT and Nucleotide Transformer (NT), achieve strong…

Genomics · Quantitative Biology 2025-08-20 Ke Ding , Brian Parker , Jiayu Wen

End-to-end automatic speech recognition systems often fail to transcribe domain-specific named entities, causing catastrophic failures in downstream tasks. Numerous fast and lightweight named entity correction (NEC) models have been…

Computation and Language · Computer Science 2025-10-27 Yuanchang Luo , Daimeng Wei , Shaojun Li , Hengchao Shang , Jiaxin Guo , Zongyao Li , Zhanglin Wu , Xiaoyu Chen , Zhiqiang Rao , Jinlong Yang , Hao Yang

Linguistic annotation of transcribed speech is essential for research in language acquisition, language disorders, and sociolinguistics, yet remains labor-intensive and time-consuming. While Large Language Models (LLMs) have shown promise…

Computation and Language · Computer Science 2026-05-19 Qingwen Zhao , Hongao Zhu , Yunqi He , Rui Wang , Aijun Huang , Hai Hu

Accurate multi-turn intent classification is essential for advancing conversational AI systems. However, challenges such as the scarcity of comprehensive datasets and the complexity of contextual dependencies across dialogue turns hinder…

Computation and Language · Computer Science 2024-11-20 Junhua Liu , Yong Keat Tan , Bin Fu , Kwan Hui Lim

Longitudinal information in radiology reports refers to the sequential tracking of findings across multiple examinations over time, which is crucial for monitoring disease progression and guiding clinical decisions. Many recent automated…

Computation and Language · Computer Science 2026-01-26 Xinyi Wang , Grazziela Figueredo , Ruizhe Li , Xin Chen

Existing deep-learning approaches to semantic column type annotation (CTA) have important shortcomings: they rely on semantic types which are fixed at training time; require a large number of training samples per type and incur large…

Computation and Language · Computer Science 2024-08-20 Benjamin Feuer , Yurong Liu , Chinmay Hegde , Juliana Freire

\textbf{Background:} Regulatory frameworks for AI in healthcare, including the EU AI Act and FDA guidance on AI/ML-based medical devices, require clinical decision support to demonstrate not only accuracy but auditability. Existing formal…

Artificial Intelligence · Computer Science 2026-04-24 Michael Bouzinier , Sergey Trifonov , Michael Chumack , Eugenia Lvova , Dmitry Etin

Accurately annotating and controlling protein function from sequence data remains a major challenge, particularly within homologous families where annotated sequences are scarce and structural variation is minimal. We present a two-stage…

Quantitative Methods · Quantitative Biology 2025-07-22 Lorenzo Rosset , Martin Weigt , Francesco Zamponi

This research examines the use of Reinforcement Learning from AI Feedback (RLAIF) techniques to improve healthcare dialogue models, with the aim of tackling the challenges of preference-aligned data annotation while reducing the reliance on…

Computation and Language · Computer Science 2024-10-08 Chengfeng Dou , Ying Zhang , Zhi Jin , Wenpin Jiao , Haiyan Zhao , Yongqiang Zhao , Zhengwei Tao

Academic researchers need efficient and reliable methods for collecting high-quality information from trusted sources, but modern tools for AI-assisted research still suffer from the tendency of Large Language Models (LLMs) to produce…

Computation and Language · Computer Science 2026-05-21 Gábor Recski , Szilveszter Tóth , Nadia Verdha , István Boros , Ádám Kovács

Recent research in the field of computer vision strongly focuses on deep learning architectures to tackle image processing problems. Deep neural networks are often considered in complex image processing scenarios since traditional computer…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Marcel P. Schilling , Luca Rettenberger , Friedrich Münke , Haijun Cui , Anna A. Popova , Pavel A. Levkin , Ralf Mikut , Markus Reischl

Text-based person retrieval aims to identify specific individuals within an image database using textual descriptions. Due to the high cost of annotation and privacy protection, researchers resort to synthesized data for the paradigm of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Hang Yu , Jiahao Wen , Zhedong Zheng

Many recent approaches to natural language tasks are built on the remarkable abilities of large language models. Large language models can perform in-context learning, where they learn a new task from a few task demonstrations, without any…

Computation and Language · Computer Science 2022-09-07 Hongjin Su , Jungo Kasai , Chen Henry Wu , Weijia Shi , Tianlu Wang , Jiayi Xin , Rui Zhang , Mari Ostendorf , Luke Zettlemoyer , Noah A. Smith , Tao Yu

The adoption of machine learning (ML) and deep learning methods has revolutionized molecular medicine by driving breakthroughs in genomics, transcriptomics, drug discovery, and biological systems modeling. The increasing quantity,…

With the rapid advancement and strong generalization capabilities of large language models (LLMs), they have been increasingly incorporated into the active learning pipelines as annotators to reduce annotation costs. However, considering…

Machine Learning · Computer Science 2026-01-23 Yuanyuan Qi , Xiaohao Yang , Jueqing Lu , Guoxiang Guo , Joanne Enticott , Gang Liu , Lan Du