中文
相关论文

相关论文: How Do Data Owners Say No? A Case Study of Data Co…

200 篇论文

The race to train language models on vast, diverse, and inconsistently documented datasets has raised pressing concerns about the legal and ethical risks for practitioners. To remedy these practices threatening data transparency and…

We investigate the contents of web-scraped data for training AI systems, at sizes where human dataset curators and compilers no longer manually annotate every sample. Building off of prior privacy concerns in machine learning models, we…

密码学与安全 · 计算机科学 2026-04-08 Rachel Hong , Jevan Hutson , William Agnew , Imaad Huda , Tadayoshi Kohno , Jamie Morgenstern

Code datasets are of immense value for training neural-network-based code completion models, where companies or organizations have made substantial investments to establish and process these datasets. Unluckily, these datasets, either built…

软件工程 · 计算机科学 2023-08-29 Zhensu Sun , Xiaoning Du , Fu Song , Li Li

Permissive licenses like MIT, Apache-2.0, and BSD-3-Clause dominate open-source AI, signaling that artifacts like models, datasets, and code can be freely used, modified, and redistributed. However, these licenses carry mandatory…

机器学习 · 计算机科学 2026-02-10 James Jewitt , Gopi Krishnan Rajbahadur , Hao Li , Bram Adams , Ahmed E. Hassan

Multimodal AI systems integrate text generation, image generation, and other capabilities within a single conversational interface. These systems employ safety mechanisms to prevent disallowed actions, including the removal of watermarks…

计算机与社会 · 计算机科学 2026-01-13 Bentley DeVilling

The rapid advancement of general-purpose AI models has increased concerns about copyright infringement in training data, yet current regulatory frameworks remain predominantly reactive rather than proactive. This paper examines the…

计算机与社会 · 计算机科学 2026-01-21 Mariia Kyrychenko , Mykyta Mudryi , Markiyan Chaklosh

Substantial research works have shown that deep models, e.g., pre-trained models, on the large corpus can learn universal language representations, which are beneficial for downstream NLP tasks. However, these powerful models are also…

密码学与安全 · 计算机科学 2024-07-16 Yixin Liu , Hongsheng Hu , Xun Chen , Xuyun Zhang , Lichao Sun

Deep neural networks (DNNs) rely heavily on high-quality open-source datasets (e.g., ImageNet) for their success, making dataset ownership verification (DOV) crucial for protecting public dataset copyrights. In this paper, we find existing…

机器学习 · 计算机科学 2025-06-17 Ting Qiao , Yiming Li , Jianbin Li , Yingjia Wang , Leyi Qi , Junfeng Guo , Ruili Feng , Dacheng Tao

With increasingly more data and computation involved in their training, machine learning models constitute valuable intellectual property. This has spurred interest in model stealing, which is made more practical by advances in learning…

机器学习 · 统计学 2021-04-23 Pratyush Maini , Mohammad Yaghini , Nicolas Papernot

The emergence of text-to-image models has recently sparked significant interest, but the attendant is a looming shadow of potential infringement by violating the user terms. Specifically, an adversary may exploit data created by a…

计算机视觉与模式识别 · 计算机科学 2024-09-25 Likun Zhang , Hao Wu , Lingcui Zhang , Fengyuan Xu , Jin Cao , Fenghua Li , Ben Niu

We study protecting a user's data (images in this work) against a learner's unauthorized use in training neural networks. It is especially challenging when the user's data is only a tiny percentage of the learner's complete training set. We…

密码学与安全 · 计算机科学 2022-08-03 Zihang Zou , Boqing Gong , Liqiang Wang

The rapid advancement of general-purpose AI models has increased concerns about copyright infringement in training data, yet current regulatory frameworks remain predominantly reactive rather than proactive. This paper examines the…

计算机与社会 · 计算机科学 2026-01-21 Mariia Kyrychenko , Mykyta Mudryi , Markiyan Chaklosh

By and large, existing Intellectual Property (IP) protection on deep neural networks typically i) focus on image classification task only, and ii) follow a standard digital watermarking framework that was conventionally used to protect the…

计算机视觉与模式识别 · 计算机科学 2021-09-01 Jian Han Lim , Chee Seng Chan , Kam Woh Ng , Lixin Fan , Qiang Yang

As datasets become critical assets in modern machine learning systems, ensuring robust copyright protection has emerged as an urgent challenge. Traditional legal mechanisms often fail to address the technical complexities of digital data…

密码学与安全 · 计算机科学 2025-09-09 Kun Li , Cheng Wang , Minghui Xu , Yue Zhang , Xiuzhen Cheng

As there are increasing needs of sharing data for machine learning, there is growing attention for the owners of the data to claim the ownership. Visible watermarking has been an effective way to claim the ownership of visual data, yet the…

密码学与安全 · 计算机科学 2019-06-05 Sanghyun Hong , Tae-hoon Kim , Tudor Dumitraş , Jonghyun Choi

The World Wide Web, a ubiquitous source of information, serves as a primary resource for countless individuals, amassing a vast amount of data from global internet users. However, this online data, when scraped, indexed, and utilized for…

网络与互联网体系结构 · 计算机科学 2023-11-07 Dawen Zhang , Boming Xia , Yue Liu , Xiwei Xu , Thong Hoang , Zhenchang Xing , Mark Staples , Qinghua Lu , Liming Zhu

The prosperity of deep neural networks (DNNs) is largely benefited from open-source datasets, based on which users can evaluate and improve their methods. In this paper, we revisit backdoor-based dataset ownership verification (DOV), which…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Junfeng Guo , Yiming Li , Lixu Wang , Shu-Tao Xia , Heng Huang , Cong Liu , Bo Li
‹ 上一页 1 2 3 10 下一页 ›