中文
相关论文

相关论文: Detection of Personal Data in Structured Datasets …

200 篇论文

Software developers frequently hard-code credentials such as passwords, generic secrets, private keys, and generic tokens in software repositories, even though it is strictly advised against due to the severe threat to the security of the…

密码学与安全 · 计算机科学 2025-06-17 Chidera Biringa , Gokhan Kul

With the rise of multimodal large language models, GPT-4o stands out as a pioneering model, driving us to evaluate its capabilities. This report assesses GPT-4o across various tasks to analyze its audio processing and reasoning abilities.…

计算与语言 · 计算机科学 2025-02-17 Yu-Xiang Lin , Chih-Kai Yang , Wei-Chih Chen , Chen-An Li , Chien-yu Huang , Xuanjun Chen , Hung-yi Lee

The primary objective of this study is to demonstrate the impact of data augmentation using ChatGPT-4o-mini on food hazard and product analysis. The augmented data is generated using ChatGPT-4o-mini and subsequently used to train two large…

计算与语言 · 计算机科学 2025-02-14 Areeg Fahad Rasheed , M. Zarkoosh , Shimam Amer Chasib , Safa F. Abbas

Mental health care poses an increasingly serious challenge to modern societies. In this context, there has been a surge in research that utilizes information technologies to address mental health problems, including those aiming to develop…

计算与语言 · 计算机科学 2024-02-21 Michimasa Inaba , Mariko Ukiyo , Keiko Takamizo

The detection of political fake statements is crucial for maintaining information integrity and preventing the spread of misinformation in society. Historically, state-of-the-art machine learning models employed various methods for…

计算与语言 · 计算机科学 2023-06-16 Mars Gokturk Buchholz

Current approaches to music emotion annotation remain heavily reliant on manual labelling, a process that imposes significant resource and labour burdens, severely limiting the scale of available annotated data. This study examines the…

声音 · 计算机科学 2025-08-19 Meng Yang , Jon McCormack , Maria Teresa Llano , Wanchao Su

Traditional dataset retrieval systems rely on metadata for indexing, rather than on the underlying data values. However, high-quality metadata creation and enrichment often require manual annotations, which is a labour-intensive and…

数据库 · 计算机科学 2024-09-09 Margherita Martorana , Tobias Kuhn , Lise Stork , Jacco van Ossenbruggen

Humor is a fundamental facet of human cognition and interaction. Yet, despite recent advances in natural language processing, humor detection remains a challenging task that is complicated by the scarcity of datasets that pair humorous…

计算与语言 · 计算机科学 2024-06-24 Zachary Horvitz , Jingru Chen , Rahul Aditya , Harshvardhan Srivastava , Robert West , Zhou Yu , Kathleen McKeown

Recently, using a powerful proprietary Large Language Model (LLM) (e.g., GPT-4) as an evaluator for long-form responses has become the de facto standard. However, for practitioners with large-scale evaluation tasks and custom criteria in…

Engineering educational curriculum and standards cover many material and manufacturing options. However, engineers and designers are often unfamiliar with certain composite materials or manufacturing techniques. Large language models (LLMs)…

The coding capabilities of large language models (LLMs) have opened up new opportunities for automatic statistical analysis in machine learning and data science. However, before their widespread adoption, it is crucial to assess the…

应用统计 · 统计学 2025-02-26 Xinyi Song , Lina Lee , Kexin Xie , Xueying Liu , Xinwei Deng , Yili Hong

How to generate a large, realistic set of tables along with joinability relationships, to stress-test dataset discovery methods? Dataset discovery methods aim to automatically identify related data assets in a data lake. The development and…

数据库 · 计算机科学 2025-07-09 Zhenwei Dai , Chuan Lei , Asterios Katsifodimos , Xiao Qin , Christos Faloutsos , Huzefa Rangwala

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an…

音频与语音处理 · 电气工程与系统科学 2025-01-22 Ryandhimas E. Zezario , Sabato M. Siniscalchi , Hsin-Min Wang , Yu Tsao

Animal ethology is an crucial aspect of animal research, and animal behavior labeling is the foundation for studying animal behavior. This process typically involves labeling video clips with behavioral semantic tags, a task that is…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Yiqi Wu , Xiaodan Hu , Ziming Fu , Siling Zhou , Jiangong Li

In this work, we explore how multimodal large language models can support real-time context- and value-aware decision-making. To do so, we combine the GPT-4o language model with a TurtleBot 4 platform simulating a smart vacuum cleaning…

机器人学 · 计算机科学 2026-02-03 Giulio Antonio Abbo , Senne Lenaerts , Tony Belpaeme

Social media datasets are essential for research on disinformation, influence operations, social sensing, hate speech detection, cyberbullying, and other significant topics. However, access to these datasets is often restricted due to costs…

计算机与社会 · 计算机科学 2024-07-12 Henry Tari , Danial Khan , Justus Rutten , Darian Othman , Rishabh Kaushal , Thales Bertaglia , Adriana Iamnitchi

Court transcripts and judgments are rich repositories of legal knowledge, detailing the intricacies of cases and the rationale behind judicial decisions. The extraction of key information from these documents provides a concise overview of…

计算与语言 · 计算机科学 2024-03-20 Joana Ribeiro de Faria , Huiyuan Xie , Felix Steffek

Leveraging the power of multimodal large language models (LLMs) offers a promising approach to enhancing the accuracy and interpretability of morphing attack detection (MAD), especially in real-world biometric applications. This work…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Ria Shekhawat , Hailin Li , Raghavendra Ramachandra , Sushma Venkatesh

The increasing demand for personalized interactions with large language models (LLMs) calls for methodologies capable of accurately and efficiently identifying user opinions and preferences. Retrieval augmentation emerges as an effective…

计算与语言 · 计算机科学 2025-02-04 Chenkai Sun , Ke Yang , Revanth Gangi Reddy , Yi R. Fung , Hou Pong Chan , Kevin Small , ChengXiang Zhai , Heng Ji

Recently, using large language models (LLMs) for data augmentation has led to considerable improvements in unsupervised sentence embedding models. However, existing methods encounter two primary challenges: limited data diversity and high…

计算与语言 · 计算机科学 2025-10-07 Peichao Lai , Zhengfeng Zhang , Wentao Zhang , Fangcheng Fu , Bin Cui