中文
相关论文

相关论文: Dataset Ownership Verification for Pre-trained Mas…

200 篇论文

With the increasing adoption of deep learning in speaker verification, large-scale speech datasets have become valuable intellectual property. To audit and prevent the unauthorized usage of these valuable released datasets, especially in…

密码学与安全 · 计算机科学 2025-04-08 Yiming Li , Kaiying Yan , Shuo Shao , Tongqing Zhai , Shu-Tao Xia , Zhan Qin , Dacheng Tao

Document Visual Question Answering (DocVQA) has introduced a new paradigm for end-to-end document understanding, and quickly became one of the standard benchmarks for multimodal LLMs. Automating document processing workflows, driven by…

机器学习 · 计算机科学 2025-02-07 Khanh Nguyen , Raouf Kerkouche , Mario Fritz , Dimosthenis Karatzas

Currently, deep neural networks (DNNs) are widely adopted in different applications. Despite its commercial values, training a well-performing DNN is resource-consuming. Accordingly, the well-trained model is valuable intellectual property…

密码学与安全 · 计算机科学 2025-03-04 Yiming Li , Linghui Zhu , Xiaojun Jia , Yang Bai , Yong Jiang , Shu-Tao Xia , Xiaochun Cao , Kui Ren

Training high performance Deep Neural Networks (DNNs) models require large-scale and high-quality datasets. The expensive cost of collecting and annotating large-scale datasets make the valuable datasets can be considered as the…

密码学与安全 · 计算机科学 2023-05-26 Mingfu Xue , Yinghao Wu , Yushu Zhang , Jian Wang , Weiqiang Liu

With the development of practical deep learning models like generative AI, their excellent performance has brought huge economic value. For instance, ChatGPT has attracted more than 100 million users in three months. Since the model…

密码学与安全 · 计算机科学 2023-12-04 Yihao Li , Yanyi Lai , Tianchi Liao , Chuan Chen , Zibin Zheng

The rapid advancement of deep neural networks (DNNs) heavily relies on large-scale, high-quality datasets. However, unauthorized commercial use of these datasets severely violates the intellectual property rights of dataset owners. Existing…

密码学与安全 · 计算机科学 2025-10-31 Yingjia Wang , Ting Qiao , Xing Liu , Chongzuo Li , Sixing Wu , Jianbin Li

Annotating medical data for training AI models is often costly and limited due to the shortage of specialists with relevant clinical expertise. This challenge is further compounded by privacy and ethical concerns associated with sensitive…

Deploying Machine Learning as a Service gives rise to model plagiarism, leading to copyright infringement. Ownership testing techniques are designed to identify model fingerprints for verifying plagiarism. However, previous works often rely…

密码学与安全 · 计算机科学 2023-10-18 Aoting Hu , Zhigang Lu , Renjie Xie , Minhui Xue

Deep neural networks are extensively applied to real-world tasks, such as face recognition and medical image classification, where privacy and data protection are critical. Image data, if not protected, can be exploited to infer personal or…

计算机视觉与模式识别 · 计算机科学 2024-02-15 Weiheng Chai , Brian Testa , Huantao Ren , Asif Salekin , Senem Velipasalar

Watermarking has become a plausible candidate for ownership verification and intellectual property protection of deep neural networks. Regarding image classification neural networks, current watermarking schemes uniformly resort to backdoor…

密码学与安全 · 计算机科学 2022-04-12 Fangqi Li , Shilin Wang

Deep neural networks have had enormous impact on various domains of computer science, considerably outperforming previous state of the art machine learning techniques. To achieve this performance, neural networks need large quantities of…

密码学与安全 · 计算机科学 2018-09-05 Dorjan Hitaj , Luigi V. Mancini

Being trained on large and vast datasets, visual foundation models (VFMs) can be fine-tuned for diverse downstream tasks, achieving remarkable performance and efficiency in various computer vision applications. The high computation cost of…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Anna Chistyakova , Mikhail Pautov

The rapid advancement of deep learning has turned models into highly valuable assets due to their reliance on massive data and costly training processes. However, these models are increasingly vulnerable to leakage and theft, highlighting…

密码学与安全 · 计算机科学 2026-05-01 Yunfei Yang , Xiaojun Chen , Zhendong Zhao , Yu Zhou , Xiaoyan Gu , Juan Cao

As large language models (LLMs) are trained on increasingly vast and opaque text corpora, determining which data contributed to training has become essential for copyright enforcement, compliance auditing, and user trust. While prior work…

计算与语言 · 计算机科学 2026-03-30 Pranav Shetty , Mirazul Haque , Zhiqiang Ma , Xiaomo Liu

Masked Image Modeling (MIM) has achieved significant success in the realm of self-supervised learning (SSL) for visual recognition. The image encoder pre-trained through MIM, involving the masking and subsequent reconstruction of input…

密码学与安全 · 计算机科学 2024-08-14 Zheng Li , Xinlei He , Ning Yu , Yang Zhang

Code datasets are of immense value for training neural-network-based code completion models, where companies or organizations have made substantial investments to establish and process these datasets. Unluckily, these datasets, either built…

软件工程 · 计算机科学 2023-08-29 Zhensu Sun , Xiaoning Du , Fu Song , Li Li

High-quality medical imaging datasets are essential for training deep learning models, but their unauthorized use raises serious copyright and ethical concerns. Medical imaging presents a unique challenge for existing dataset ownership…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Pranav Kulkarni , Junfeng Guo , Heng Huang

Detecting whether a given text is a member of the pre-training data of Large Language Models (LLMs) is crucial for ensuring data privacy and copyright protection. Most existing methods rely on the LLM's hidden information (e.g., model…

计算与语言 · 计算机科学 2025-06-25 Ruihan Hu , Yu-Ming Shang , Jiankun Peng , Wei Luo , Yazhe Wang , Xi Zhang

With the widespread use of deep neural networks (DNNs) in many areas, more and more studies focus on protecting DNN models from intellectual property (IP) infringement. Many existing methods apply digital watermarking to protect the DNN…

密码学与安全 · 计算机科学 2022-07-11 Lina Lin , Hanzhou Wu

To safely deploy deep learning-based computer vision models for computer-aided detection and diagnosis, we must ensure that they are robust and reliable. Towards that goal, algorithmic auditing has received substantial attention. To guide…

机器学习 · 计算机科学 2023-04-07 Mitchell Pavlak , Nathan Drenkow , Nicholas Petrick , Mohammad Mehdi Farhangi , Mathias Unberath