English
Related papers

Related papers: GLiNER2-PII: A Multilingual Model for Personally I…

200 papers

Large language models (LLMs) do not preserve privacy at inference-time. The LLM's outputs can inadvertently reveal information about the model's context, which presents a privacy challenge when the LLM is augmented via tools or databases…

Computation and Language · Computer Science 2026-02-03 Rushil Thareja , Preslav Nakov , Praneeth Vepakomma , Nils Lukas

We present MULTICONER V2, a dataset for fine-grained Named Entity Recognition covering 33 entity classes across 12 languages, in both monolingual and multilingual settings. This dataset aims to tackle the following practical challenges in…

Computation and Language · Computer Science 2023-10-23 Besnik Fetahu , Zhiyu Chen , Sudipta Kar , Oleg Rokhlenko , Shervin Malmasi

We curated WikiPII, an automatically labeled dataset composed of Wikipedia biography pages, annotated for personal information extraction. Although automatic annotation can lead to a high degree of label noise, it is an inexpensive process…

Computation and Language · Computer Science 2021-05-20 Rajitha Hathurusinghe , Isar Nejadgholi , Miodrag Bolic

Users can divulge sensitive information to proprietary LLM providers, raising significant privacy concerns. While open-source models, hosted locally on the user's machine, alleviate some concerns, models that users can host locally are…

Cryptography and Security · Computer Science 2025-03-27 Li Siyan , Vethavikashini Chithrra Raghuram , Omar Khattab , Julia Hirschberg , Zhou Yu

The de-identification of private information in medical data is a crucial process to mitigate the risk of confidentiality breaches, particularly when patient personal details are not adequately removed before the release of medical records.…

Cryptography and Security · Computer Science 2025-04-29 Guanchen Wu , Linzhi Zheng , Han Xie , Zhen Xiang , Jiaying Lu , Darren Liu , Delgersuren Bold , Bo Li , Xiao Hu , Carl Yang

Large-scale sharing of dialogue data is key to advancing the science of teaching and learning, yet rigorous de-identification remains a major barrier. In mathematics tutoring transcripts, numeric expressions frequently resemble structured…

Browsing-enabled LLM assistants can fetch webpages and answer contact-seeking queries, creating a practical channel for scraping contact-style personally identifiable information (PII) from public pages. Many prior defenses are deployed at…

Cryptography and Security · Computer Science 2026-05-06 Mingshuo Liu , Yiwei Zha , Min Chen

Named Entity Recognition (NER) is a foundational task in Natural Language Processing (NLP) and Information Retrieval (IR), which facilitates semantic search and structured data extraction. We introduce \textbf{AWED-FiNER}, an open-source…

Computation and Language · Computer Science 2026-02-23 Prachuryya Kaushik , Ashish Anand

Prediction of protein-ligand interactions (PLI) plays a crucial role in drug discovery as it guides the identification and optimization of molecules that effectively bind to target proteins. Despite remarkable advances in deep…

Biomolecules · Quantitative Biology 2023-07-18 Seokhyun Moon , Sang-Yeon Hwang , Jaechang Lim , Woo Youn Kim

Semantic identifier (ID) is an important concept in information retrieval that aims to preserve the semantics of objects such as documents and items inside their IDs. Previous studies typically adopt a two-stage pipeline to learn semantic…

Information Retrieval · Computer Science 2024-06-14 Bowen Jin , Hansi Zeng , Guoyin Wang , Xiusi Chen , Tianxin Wei , Ruirui Li , Zhengyang Wang , Zheng Li , Yang Li , Hanqing Lu , Suhang Wang , Jiawei Han , Xianfeng Tang

Open Cyber threat intelligence (OpenCTI) information is available in an unstructured format from heterogeneous sources on the Internet. We present CyNER, an open-source python library for cybersecurity named entity recognition (NER). CyNER…

Cryptography and Security · Computer Science 2022-04-13 Md Tanvirul Alam , Dipkamal Bhusal , Youngja Park , Nidhi Rastogi

Aligning large language models (LLMs) typically aim to reflect general human values and behaviors, but they often fail to capture the unique characteristics and preferences of individual users. To address this gap, we introduce the concept…

Computation and Language · Computer Science 2025-03-11 Minjun Zhu , Yixuan Weng , Linyi Yang , Yue Zhang

The increasing use of Online Vision Language Models (OVLMs) for processing images has introduced significant privacy risks, as individuals frequently upload images for various utilities, unaware of the potential for privacy violations.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Karmesh Siddharam Chaudhari , Youxiang Zhu , Amy Feng , Xiaohui Liang , Honggang Zhang

Continuous monitoring of bipolar disorder agitation via voice biomarkers requires disentangling stable speaker traits from volatile affective states on resource-constrained edge devices. We introduce MP-IB, the first framework to treat…

Machine Learning · Computer Science 2026-05-06 Joydeep Chandra

Advancing AI in computational pathology requires large, high-quality, and diverse datasets, yet existing public datasets are often limited in organ diversity, class coverage, or annotation quality. To bridge this gap, we introduce SPIDER…

Image and Video Processing · Electrical Eng. & Systems 2025-04-08 Dmitry Nechaev , Alexey Pchelnikov , Ekaterina Ivanova

Accurate recognition of specific categories, such as persons' names, dates or other identifiers is critical in many Automatic Speech Recognition (ASR) applications. As these categories represent personal information, ethical use of this…

With the growing use of camera devices, the industry has many image datasets that provide more opportunities for collaboration between the machine learning community and industry. However, the sensitive information in the datasets…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Jia-Wei Chen , Li-Ju Chen , Chia-Mu Yu , Chun-Shien Lu

Deploying large language model (LLM) on edge device enables personalized LLM agents for various users. The growing availability of diverse personalized agents presents a unique opportunity for peer-to-peer (P2P) collaboration, wherein each…

Computation and Language · Computer Science 2026-05-19 Zile Wang , Qianli Liu , Kaibin Guo , Haodong Wang , Jian Lin , Zicong Hong , Song Guo

The use of Natural Language Processing (NLP) in highstakes AI-based applications has increased significantly in recent years, especially since the emergence of Large Language Models (LLMs). However, despite their strong performance, LLMs…

De-identification of clinical text remains essential for secondary use of electronic health records (EHRs), yet public benchmarks such as i2b2 2006/2014 are over a decade old and lack the semantic and demographic diversity of modern…

Computation and Language · Computer Science 2026-05-06 Jose D. Posada , David Love , Somalee Datta , Priya Desai