中文
相关论文

相关论文: Towards Best Practices for Open Datasets for LLM T…

200 篇论文

Large Language Models (LLMs) have achieved unparalleled success across diverse language modeling tasks in recent years. However, this progress has also intensified ethical concerns, impacting the deployment of LLMs in everyday contexts.…

The reliance of language model training on massive amounts of computation and vast datasets scraped from potentially low-quality, copyrighted, or sensitive data has come into question practically, legally, and ethically. Federated learning…

机器学习 · 计算机科学 2024-05-28 Alex Iacob , Lorenzo Sani , Bill Marino , Preslav Aleksandrov , William F. Shen , Nicholas Donald Lane

Recent work has advocated for training AI models on ever-larger datasets, arguing that as the size of a dataset increases, the performance of a model trained on that dataset will correspondingly increase (referred to as "scaling laws"). In…

机器学习 · 计算机科学 2024-07-30 Fernando Diaz , Michael Madaio

Language models may memorize more than just facts, including entire chunks of texts seen during training. Fair use exemptions to copyright laws typically allow for limited use of copyrighted material without permission from the copyright…

计算与语言 · 计算机科学 2023-10-24 Antonia Karamolegkou , Jiaang Li , Li Zhou , Anders Søgaard

Systemic property dispossession from minority groups has often been carried out in the name of technological progress. In this paper, we identify evidence that the current paradigm of large language models (LLMs) likely continues this long…

计算机与社会 · 计算机科学 2024-03-21 Heila Precel , Allison McDonald , Brent Hecht , Nicholas Vincent

The scarcity of accessible, compliant, and ethically sourced data presents a considerable challenge to the adoption of artificial intelligence (AI) in sensitive fields like healthcare, finance, and biomedical research. Furthermore, access…

机器学习 · 计算机科学 2025-04-02 Kumar Kshitij Patel , Weitong Zhang , Lingxiao Wang

New capabilities in foundation models are owed in large part to massive, widely-sourced, and under-documented training data collections. Existing practices in data collection have led to challenges in tracing authenticity, verifying…

The growing adoption of large language models in legal practice brings both significant promise and serious risk. Legal professionals stand to benefit from AI that can reason over contracts, draft documents, and analyze sources at scale,…

人工智能 · 计算机科学 2026-05-15 Olivia Peiyu Wang , Leilani H. Gilpin

Large Language Models (LLMs) are a powerful technology that augment human skill to create new opportunities, akin to the development of steam engines and the internet. However, LLMs come with a high cost. They require significant computing…

计算机与社会 · 计算机科学 2024-04-16 Vishwas Sathish , Hannah Lin , Aditya K Kamath , Anish Nyayachavadi

The rapid advancement of Large Language Models (LLMs) has transformed knowledge-intensive has led to its widespread usage by knowledge workers to enhance their productivity. As these professionals handle sensitive information, and the…

人机交互 · 计算机科学 2025-07-28 Siying Hu , Piaohong Wang , Ka I Chan , Yaxing Yao , Zhicong Lu

Large Language Models (LLMs) are being adopted across a wide range of tasks, including decision-making processes in industries where bias in AI systems is a significant concern. Recent research indicates that LLMs can harbor implicit biases…

计算与语言 · 计算机科学 2024-10-18 Divyanshu Kumar , Umang Jain , Sahil Agarwal , Prashanth Harshangi

Large Language Models (LLMs) have raised significant concerns regarding the fair use of copyright-protected content. While prior studies have examined the extent to which LLMs reproduce copyrighted materials, they have predominantly focused…

计算机与社会 · 计算机科学 2025-03-11 Yupeng Chen , Xiaoyu Zhang , Yixian Huang , Qian Xie

This paper explores the economic underpinnings of open sourcing advanced large language models (LLMs) by for-profit companies. Empirical analysis reveals that: (1) LLMs are compatible with R&D portfolios of numerous technologically…

综合经济学 · 经济学 2025-01-22 Mahyar Habibi

Ensuring the safety and compliance of large language models (LLMs) is of paramount importance. However, existing LLM safety datasets often rely on ad-hoc taxonomies for data generation and suffer from a significant shortage of…

计算与语言 · 计算机科学 2026-04-17 Wenbin Hu , Huihao Jing , Haochen Shi , Changxuan Fan , Haoran Li , Yangqiu Song

Children are increasingly using technologies powered by Artificial Intelligence (AI). However, there are growing concerns about privacy risks, particularly for children. Although existing privacy regulations require companies and…

人工智能 · 计算机科学 2026-02-20 Diana Addae , Diana Rogachova , Nafiseh Kahani , Masoud Barati , Michael Christensen , Chen Zhou

Large Language Models (LLMs) have shown greatly enhanced performance in recent years, attributed to increased size and extensive training data. This advancement has led to widespread interest and adoption across industries and the public.…

计算与语言 · 计算机科学 2024-06-19 Victoria Smith , Ali Shahin Shamsabadi , Carolyn Ashurst , Adrian Weller

High-quality datasets are typically required for accomplishing data-driven tasks, such as training medical diagnosis models, predicting real-time traffic conditions, or conducting experiments to validate research hypotheses. Consequently,…

信息检索 · 计算机科学 2025-09-03 Pengyue Li , Sheng Wang , Hua Dai , Zhiyu Chen , Zhifeng Bao , Brian D. Davison

Recent advancements in Large Language Models (LLMs), such as ChatGPT and LLaMA, have significantly transformed Natural Language Processing (NLP) with their outstanding abilities in text generation, summarization, and classification.…

计算与语言 · 计算机科学 2024-08-12 Md Nazmus Sakib , Md Athikul Islam , Royal Pathak , Md Mashrur Arifin

As machine learning systems grow in scale, so do their training data requirements, forcing practitioners to automate and outsource the curation of training data in order to achieve state-of-the-art performance. The absence of trustworthy…

The development of Large Language Models (LLMs) in various languages has been advancing, but the combination of non-English languages with domain-specific contexts remains underexplored. This paper presents our findings from training and…

计算与语言 · 计算机科学 2024-11-07 Kosuke Takahashi , Takahiro Omi , Kosuke Arima , Tatsuya Ishigaki