English
Related papers

Related papers: Putting GenAI on Notice: GenAI Exceptionalism and …

200 papers

GenAI companies are strip-mining the web. Their scraping bots harvest content at an unprecedented scale, circumventing technical barriers to fuel billion-dollar models while creators receive nothing. Courts have enabled this exploitation by…

Computers and Society · Computer Science 2026-02-02 David Atkinson

Search engines are a combination of hardware and computer software supplied by a particular company through the website which has been determined. Search engines collect information from the web through bots or web crawlers that crawls the…

Information Retrieval · Computer Science 2014-10-22 Ahmad Josi , Leon Andretti Abdillah , Suryayusra

The success of generative AI relies heavily on training on data scraped through extensive crawling of the Internet, a practice that has raised significant copyright, privacy, and ethical concerns. While few measures are designed to resist a…

Human-Computer Interaction · Computer Science 2025-05-08 Enze Liu , Elisa Luo , Shawn Shan , Geoffrey M. Voelker , Ben Y. Zhao , Stefan Savage

This paper explores the legal implications of violating "robots.txt", a technical standard widely used by webmasters to communicate restrictions on automated access to website content. Although historically regarded as a voluntary…

Computers and Society · Computer Science 2025-09-17 Chien-yi Chang , Xin He

Recent trends reveal the search by companies for a legal hook to prevent the undesired and unauthorized copying of information posted on websites. In the center of this controversy are metasites, websites that display prices for a variety…

Computers and Society · Computer Science 2007-05-23 Jeffrey M. Rosenfeld

Online data scraping has taken on new dimensions in recent years, as traditional scrapers have been joined by new AI-specific bots. To counteract unwanted scraping, many sites use tools like the Robots Exclusion Protocol (REP), which places…

Networking and Internet Architecture · Computer Science 2025-10-24 Taein Kim , Karstan Bock , Claire Luo , Amanda Liswood , Chloe Poroslay , Emily Wenger

Scientists across disciplines often use data from the internet to conduct research, generating valuable insights about human behavior. However, as generative AI relying on massive text corpora becomes increasingly valuable, platforms have…

Computers and Society · Computer Science 2024-12-20 Megan A. Brown , Andrew Gruen , Gabe Maldoff , Solomon Messing , Zeve Sanderson , Michael Zimmer

The research investigates the crucial role of clear and intelligible terms of service in cultivating user trust and facilitating informed decision-making in the context of AI, in specific GenAI. It highlights the obstacles presented by…

Computers and Society · Computer Science 2024-06-21 Sundaraparipurnan Narayanan

Artificial intelligence (AI) model creators commonly attach restrictive terms of use to both their models and their outputs. These terms typically prohibit activities ranging from creating competing AI models to spreading disinformation.…

Computers and Society · Computer Science 2024-12-11 Peter Henderson , Mark A. Lemley

The increasing adoption of econometric and machine-learning approaches by empirical researchers has led to a widespread use of one data collection method: web scraping. Web scraping refers to the use of automated computer programs to access…

General Economics · Economics 2023-08-07 Jens Foerderer

Generative AI services like ChatGPT and Gemini are some of the fastest-growing consumer services. Individuals using such services must accept their terms of use before access, and conform to these terms for continued use of the service.…

Computers and Society · Computer Science 2026-03-23 Harshvardhan J. Pandit , Dick A. H. Blankvoort , Adel Shaaban , Sasha Luccioni , Abeba Birhane

With much of our lives taking place online, researchers are increasingly turning to information from the World Wide Web to gain insights into geographic patterns and processes. Web scraping as an online data acquisition technique allows us…

Information Retrieval · Computer Science 2023-06-01 Alexander Brenning , Sebastian Henn

This paper challenges the argument that generative artificial intelligence (GenAI) is entitled to broad immunity from copyright law for reproducing copyrighted works without authorization due to a fair use defense. It examines fair use…

Computers and Society · Computer Science 2025-10-20 David Atkinson

We investigate the contents of web-scraped data for training AI systems, at sizes where human dataset curators and compilers no longer manually annotate every sample. Building off of prior privacy concerns in machine learning models, we…

Cryptography and Security · Computer Science 2026-04-08 Rachel Hong , Jevan Hutson , William Agnew , Imaad Huda , Tadayoshi Kohno , Jamie Morgenstern

The groundbreaking advancements around generative AI have recently caused a wave of concern culminating in a row of lawsuits, including high-profile actions against Stability AI and OpenAI. This situation of legal uncertainty has sparked a…

Information Retrieval · Computer Science 2024-04-04 Michael Dinzinger , Florian Heß , Michael Granitzer

From pre-training to query-time augmentation, web-scraped data helps to improve the quality and contextual relevancy of content generated by large language models (LLMs). However, large-scale web scraping to feed LLMs can affect site…

Cryptography and Security · Computer Science 2026-05-14 Steven Seiden , Triss Ren , Caroline Zhang , Taein Kim , Enze Liu , Emily Wenger

The World Wide Web, a ubiquitous source of information, serves as a primary resource for countless individuals, amassing a vast amount of data from global internet users. However, this online data, when scraped, indexed, and utilized for…

Networking and Internet Architecture · Computer Science 2023-11-07 Dawen Zhang , Boming Xia , Yue Liu , Xiwei Xu , Thong Hoang , Zhenchang Xing , Mark Staples , Qinghua Lu , Liming Zhu

In the present time, all know about World Wide Web and work over the Internet daily. In this paper, we introduce the search engines working for keywords that are entered by users to find something. The search engine uses different search…

Artificial Intelligence · Computer Science 2024-03-01 Piyush Vyas , Akhilesh Chauhan , Tushar Mandge , Surbhi Hardikar

Fact verification is a critical yet underexplored component of non-litigation legal practice. While existing research has examined automation in legal workflow and human-AI collaboration in high-stakes domains, little is known about how…

Human-Computer Interaction · Computer Science 2026-02-10 Sirui Han , Yuyao Zhang , Yidan Huang , Xueyan Li , Chengzhong Liu , Yike Guo
‹ Prev 1 2 3 10 Next ›