English
Related papers

Related papers: An Evaluation Study of Hybrid Methods for Multilin…

200 papers

The memorization of sensitive and personally identifiable information (PII) by large language models (LLMs) poses growing privacy risks as models scale and are increasingly deployed in real-world applications. Existing efforts to study…

Computation and Language · Computer Science 2025-05-20 Sriram Selvam , Anneswa Ghosh

We propose an end-to-end one-step person search approach with learnable proposals, named LEAPS. Given a set of sparse and learnable proposals, LEAPS employs a dynamic person search head to directly perform person detection and corresponding…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Zhiqiang Dong , Jiale Cao , Rao Muhammad Anwer , Jin Xie , Fahad Khan , Yanwei Pang

Street-level imagery contains personally identifiable information (PII), some of which is context-dependent. Existing anonymization methods either over-process images or miss subtle identifiers, while API-based solutions compromise data…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Robert Aufschläger , Jakob Folz , Gautam Savaliya , Manjitha D Vidanalage , Michael Heigl , Martin Schramm

Key Information Extraction (KIE) from visually-rich documents (VrDs) is a critical task, for which recent Large Language Models (LLMs) and Multi-Modal Large Language Models (MLLMs) have demonstrated strong potential. However, their reliance…

Computation and Language · Computer Science 2026-01-28 Xinzhong Wang , Ya Guo , Jing Li , Huan Chen , Yi Tu , Yijie Hong , Gongshen Liu , Huijia Zhu

Regulatory texts are inherently long and complex, presenting significant challenges for information retrieval systems in supporting regulatory officers with compliance tasks. This paper introduces a hybrid information retrieval system that…

Computation and Language · Computer Science 2025-02-25 Jhon Rayo , Raul de la Rosa , Mario Garrido

Binary decompilation plays an important role in software security analysis, reverse engineering, and malware understanding when source code is unavailable. However, existing decompilation techniques often fail to produce source code that…

Software Engineering · Computer Science 2026-04-14 Xiaohan Wang , Yuxin Hu , Kevin Leach

Protecting sensitive information is crucial in today's world of Large Language Models (LLMs) and data-driven services. One common method used to preserve privacy is by using data perturbation techniques to reduce overreaching utility of…

Computation and Language · Computer Science 2023-07-19 Ajinkya Deshmukh , Saumya Banthia , Anantha Sharma

Large-scale sharing of dialogue data is key to advancing the science of teaching and learning, yet rigorous de-identification remains a major barrier. In mathematics tutoring transcripts, numeric expressions frequently resemble structured…

Background: Leaking sensitive information - such as API keys, tokens, and credentials - in source code remains a persistent security threat. Traditional regex and entropy-based tools often generate high false positives due to limited…

Software Engineering · Computer Science 2025-07-29 Md Nafiu Rahman , Sadif Ahmed , Zahin Wahab , S M Sohan , Rifat Shahriyar

There has been limited success for dense retrieval models in multilingual retrieval, due to uneven and scarce training data available across multiple languages. Synthetic training data generation is promising (e.g., InPars or Promptagator),…

Information Retrieval · Computer Science 2024-04-17 Nandan Thakur , Jianmo Ni , Gustavo Hernández Ábrego , John Wieting , Jimmy Lin , Daniel Cer

High relevance of retrieved and re-ranked items to the search query is the cornerstone of successful product search, yet measuring relevance of items to queries is one of the most challenging tasks in product information retrieval, and…

Despite advances in the use of large language models (LLMs) in downstream tasks, their ability to memorize information has raised privacy concerns. Therefore, protecting personally identifiable information (PII) during LLM training remains…

Machine Learning · Computer Science 2025-12-02 Stella Etuk , Ashraf Matrawy

Large language models (LLMs) are widely used, but concerns about data contamination challenge the reliability of LLM evaluations. Existing contamination detection methods are often task-specific or require extra prerequisites, limiting…

Computation and Language · Computer Science 2024-10-22 Yi Zhao , Jing Li , Linyi Yang

AI chatbots have quietly become the world's most popular therapists, coaches, and confidants. Users of cloud-based LLM services are increasingly shifting from simple queries like idea generation and poem writing, to deeply personal…

Human-Computer Interaction · Computer Science 2026-04-08 Max Holschneider , Saetbyeol LeeYouk

Large Language Models (LLMs) have transformed natural language processing, yet improving their problem-solving capabilities, particularly for complex, reasoning-intensive tasks, remains a persistent challenge. This paper introduces the REAP…

Computation and Language · Computer Science 2024-09-17 Ryan Lingo , Martin Arroyo , Rajeev Chhajer

The rapid advancement and widespread use of large language models (LLMs) have raised significant concerns regarding the potential leakage of personally identifiable information (PII). These models are often trained on vast quantities of…

Cryptography and Security · Computer Science 2023-07-06 Siwon Kim , Sangdoo Yun , Hwaran Lee , Martin Gubri , Sungroh Yoon , Seong Joon Oh

Text-based person anomaly retrieval has emerged as a challenging task, with most existing approaches relying on complex deep-learning techniques. This raises a research question: How can the model be optimized to achieve greater…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Tien-Huy Nguyen , Huu-Loc Tran , Huu-Phong Phan-Nguyen , Quang-Vinh Dinh

Coding agents and LLM-powered applications routinely send potentially sensitive content to cloud LLM APIs where it may be logged, retained, used for training, or subpoenaed. Existing privacy tooling focuses on network-level encryption and…

Large Language Models (LLMs) generate responses based on user prompts. Often, these prompts may contain highly sensitive information, including personally identifiable information (PII), which could be exposed to third parties hosting these…

Cryptography and Security · Computer Science 2026-03-30 Shashie Dilhara Batan Arachchige , Hassan Jameel Asghar , Benjamin Zi Hao Zhao , Dinusha Vatsalan , Dali Kaafar

Large language models (LLMs) exhibit remarkable performance across various NLP tasks. However, they often generate incorrect or hallucinated information, which hinders their practical applicability in real-world scenarios. Human feedback…

Computation and Language · Computer Science 2023-05-24 Wenhao Yu , Zhihan Zhang , Zhenwen Liang , Meng Jiang , Ashish Sabharwal