中文
相关论文

相关论文: AspirinSum: an Aspect-based utility-preserved de-i…

200 篇论文

With the proliferation of open-sourced Large Language Models (LLMs) and efficient finetuning techniques, we are on the cusp of the emergence of numerous domain-specific LLMs that have been finetuned for expertise across specialized fields…

计算与语言 · 计算机科学 2023-06-28 Teo Susnjak

Aspect-based summarization aims to generate summaries that highlight specific aspects of a text, enabling more personalized and targeted summaries. However, its application to books remains unexplored due to the difficulty of constructing…

计算与语言 · 计算机科学 2025-11-11 Ryuhei Miyazato , Ting-Ruen Wei , Xuyang Wu , Hsin-Tai Wu , Kei Harada

Analyzing vast textual data and summarizing key information from electronic health records imposes a substantial burden on how clinicians allocate their time. Although large language models (LLMs) have shown promise in natural language…

In many countries, personal information that can be published or shared between organizations is regulated and, therefore, documents must undergo a process of de-identification to eliminate or obfuscate confidential data. Our work focuses…

计算与语言 · 计算机科学 2019-10-10 Diego Garat , Dina Wonsever

Substance use disorder (SUD) poses a major concern due to its detrimental effects on health and society. SUD identification and treatment depend on a variety of factors such as severity, co-determinants (e.g., withdrawal symptoms), and…

计算与语言 · 计算机科学 2024-03-20 Maria Mahbub , Gregory M. Dams , Sudarshan Srinivasan , Caitlin Rizy , Ioana Danciu , Jodie Trafton , Kathryn Knight

Legal documents are often long, dense, and difficult to comprehend, not only for laypeople but also for legal experts. While automated document summarization has great potential to improve access to legal knowledge, prevailing task-based…

计算与语言 · 计算机科学 2026-03-24 Tsz Fung Pang , Maryam Berijanian , Thomas Orth , Breanna Shi , Charlotte S. Alexander

Unstructured textual data is at the heart of healthcare systems. For obvious privacy reasons, these documents are not accessible to researchers as long as they contain personally identifiable information. One way to share this data while…

密码学与安全 · 计算机科学 2022-11-03 Yakini Tchouka , Jean-François Couchot , David Laiymani

The current winning recipe for automatic summarization is using proprietary large-scale language models (LLMs) such as ChatGPT as is, or imitation learning from them as teacher models. While increasingly ubiquitous dependence on such…

计算与语言 · 计算机科学 2024-08-21 Jaehun Jung , Ximing Lu , Liwei Jiang , Faeze Brahman , Peter West , Pang Wei Koh , Yejin Choi

Cross-Domain Sequential Recommendation (CDSR) aims to mine and transfer users' sequential preferences across different domains to alleviate the long-standing cold-start issue. Traditional CDSR models capture collaborative information…

机器学习 · 计算机科学 2024-06-06 Tingjia Shen , Hao Wang , Jiaqing Zhang , Sirui Zhao , Liangyue Li , Zulong Chen , Defu Lian , Enhong Chen

Recent advancements in Large Language Models (LLMs) and Prompt Engineering have made chatbot customization more accessible, significantly reducing barriers to tasks that previously required programming skills. However, prompt evaluation,…

人机交互 · 计算机科学 2025-08-13 Sam Yu-Te Lee , Aryaman Bahukhandi , Dongyu Liu , Kwan-Liu Ma

As LLMs rapidly advance and enter real-world use, their privacy implications are increasingly important. We study an authorship de-anonymization threat: using LLMs to link anonymous documents to their authors, potentially compromising…

密码学与安全 · 计算机科学 2026-04-17 Lirui Zhang , Huishuai Zhang

The objective of this study is to address the critical issue of de-identification of clinical reports in order to allow access to data for research purposes, while ensuring patient privacy. The study highlights the difficulties faced in…

计算与语言 · 计算机科学 2023-03-24 Xavier Tannier , Perceval Wajsbürt , Alice Calliger , Basile Dura , Alexandre Mouchet , Martin Hilka , Romain Bey

Large Language Models (LLMs) suffer from a critical limitation: their knowledge is static and quickly becomes outdated. Retraining these massive models is computationally prohibitive, while existing knowledge editing techniques can be slow…

计算与语言 · 计算机科学 2025-12-30 Kabir Khan , Priya Sharma , Arjun Mehta , Neha Gupta , Ravi Narayanan

De-identification in the healthcare setting is an application of NLP where automated algorithms are used to remove personally identifying information of patients (and, sometimes, providers). With the recent rise of generative large language…

计算与语言 · 计算机科学 2025-09-19 Kiana Aghakasiri , Noopur Zambare , JoAnn Thai , Carrie Ye , Mayur Mehta , J. Ross Mitchell , Mohamed Abdalla

The rise of pre-trained language models has yielded substantial progress in the vast majority of Natural Language Processing (NLP) tasks. However, a generic approach towards the pre-training procedure can naturally be sub-optimal in some…

计算与语言 · 计算机科学 2021-09-03 Entony Lekhtman , Yftah Ziser , Roi Reichart

Named entity recognition is an important task when constructing knowledge bases from unstructured data sources. Whereas entity detection methods mostly rely on extensive training data, Large Language Models (LLMs) have paved the way towards…

计算与语言 · 计算机科学 2025-01-09 Lu Gan , Martin Blum , Danilo Dessi , Brigitte Mathiak , Ralf Schenkel , Stefan Dietze

Large language models (LLMs) are increasingly being used as decision aids. However, users have diverse values and preferences that can affect their decision-making, which requires novel methods for LLM alignment and personalization.…

Online product reviews contain rich but noisy signals that overwhelm users and hinder effective decision-making. Existing LLM-based summarizers remain generic and fail to account for individual preferences, limiting their practical utility.…

计算与语言 · 计算机科学 2025-12-15 Yuming Feng , Xinrui Jiang

The de-identification of private information in medical data is a crucial process to mitigate the risk of confidentiality breaches, particularly when patient personal details are not adequately removed before the release of medical records.…

密码学与安全 · 计算机科学 2025-04-29 Guanchen Wu , Linzhi Zheng , Han Xie , Zhen Xiang , Jiaying Lu , Darren Liu , Delgersuren Bold , Bo Li , Xiao Hu , Carl Yang

Face anonymization aims to conceal identity information while preserving non-identity attributes. Mainstream diffusion models rely on inference-time interventions such as negative guidance or energy-based optimization, which are applied…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Haoxin Yang , Yihong Lin , Jingdan Kang , Xuemiao Xu , Yue Li , Cheng Xu , Shengfeng He