English
Related papers

Related papers: Socially Aware Synthetic Data Generation for Suici…

200 papers

Large Language Models (LLMs) are increasingly integrated into critical decision-making pipelines, a trend that raises the demand for robust and automated data analysis. Current approaches to dataset risk analysis are limited to manual…

Artificial Intelligence · Computer Science 2026-05-28 Panteleimon Rodis

We describe our contribution to the Strict and Strict-Small tracks of the 2nd iteration of the BabyLM Challenge. The shared task is centered around efficient pre-training given data constraints motivated by human development. In response,…

Computation and Language · Computer Science 2025-08-04 Nikitas Theodoropoulos , Giorgos Filandrianos , Vassilis Lyberatos , Maria Lymperaiou , Giorgos Stamou

Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Tianchen Zhao , Xuanbai Chen , Zhihua Li , Jun Fang , Dongsheng An , Xiang Xu , Zhuowen Tu , Yifan Xing

Large language model-powered chatbots have transformed how people seek information, especially in high-stakes contexts like mental health. Despite their support capabilities, safe detection and response to crises such as suicidal ideation…

Computation and Language · Computer Science 2026-04-09 Adrian Arnaiz-Rodriguez , Miguel Baidal , Erik Derner , Jenn Layton Annable , Mark Ball , Mark Ince , Elvira Perez Vallejos , Nuria Oliver

We present a novel approach to automatically generate non-trivial task-specific synthetic datasets for hallucination detection. Our approach features a two-step generation-selection pipeline, using hallucination pattern guidance and a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Yong Xie , Karan Aggarwal , Aitzaz Ahmad , Stephen Lau

Clinical data is fundamental to advance neurosurgical research, but access is often constrained by data availability, small sample sizes, privacy regulations, and resource-intensive preprocessing and de-identification procedures. Synthetic…

Computation and Language · Computer Science 2025-10-21 Austin A. Barr , Eddie Guo , Emre Sezgin

This study investigates the automation of meta-analysis in scientific documents using large language models (LLMs). Meta-analysis is a robust statistical method that synthesizes the findings of multiple studies support articles to provide a…

Computation and Language · Computer Science 2024-11-19 Jawad Ibn Ahad , Rafeed Mohammad Sultan , Abraham Kaikobad , Fuad Rahman , Mohammad Ruhul Amin , Nabeel Mohammed , Shafin Rahman

Data augmentation has been widely used to improve deep neural networks in many research fields, such as computer vision. However, less work has been done in the context of text, partially due to its discrete nature and the complexity of…

Computation and Language · Computer Science 2021-01-12 Ping Yu , Ruiyi Zhang , Yang Zhao , Yizhe Zhang , Chunyuan Li , Changyou Chen

Large language models (LLMs) are increasingly used for mental health support, yet existing safety evaluations rely primarily on small, simulation-based test sets that have an unknown relationship to the linguistic distribution of real…

Computers and Society · Computer Science 2026-01-27 Caitlin A. Stamatis , Jonah Meyerhoff , Richard Zhang , Olivier Tieleman , Matteo Malgaroli , Thomas D. Hull

The latest large language models (LLMs) such as ChatGPT, exhibit strong capabilities in automated mental health analysis. However, existing relevant studies bear several limitations, including inadequate evaluations, lack of prompting…

Computation and Language · Computer Science 2024-10-03 Kailai Yang , Shaoxiong Ji , Tianlin Zhang , Qianqian Xie , Ziyan Kuang , Sophia Ananiadou

A growing body of work has focused on text classification methods for detecting the increasing amount of hate speech posted online. This progress has been limited to only a select number of highly-resourced languages causing detection…

Computation and Language · Computer Science 2023-10-05 Aman Khullar , Daniel Nkemelu , Cuong V. Nguyen , Michael L. Best

Obtaining data in the medical field is challenging, making the adoption of AI technology within the space slow and high-risk. We evaluate whether we can overcome this obstacle with synthetic data generated by large language models (LLMs).…

Machine Learning · Computer Science 2024-09-04 Paulo Soares , Sean McCurdy , Andrew J. Gerber , Peter Fonagy

Psychological support hotlines serve as critical lifelines for crisis intervention but encounter significant challenges due to rising demand and limited resources. Large language models (LLMs) offer potential support in crisis assessments,…

Computation and Language · Computer Science 2025-12-19 Guifeng Deng , Shuyin Rao , Tianyu Lin , Anlu Dai , Pan Wang , Junyi Xie , Haidong Song , Ke Zhao , Dongwu Xu , Zhengdong Cheng , Tao Li , Haiteng Jiang

Mental health conditions affect hundreds of millions globally, yet early detection remains limited. While large language models (LLMs) have shown promise in mental health applications, their size and computational demands hinder practical…

Artificial Intelligence · Computer Science 2025-12-17 Tianyi Zhang , Xiangyuan Xue , Lingyan Ruan , Shiya Fu , Feng Xia , Simon D'Alfonso , Vassilis Kostakos , Ting Dang , Hong Jia

The early detection of suicide risk is important since it enables the intervention to prevent potential suicide attempts. This paper studies the automatic detection of suicide risk based on spontaneous speech from adolescents, and collects…

Computation and Language · Computer Science 2024-09-25 Ziyun Cui , Chang Lei , Wen Wu , Yinan Duan , Diyang Qu , Ji Wu , Runsen Chen , Chao Zhang

Automatic hate speech detection using deep neural models is hampered by the scarcity of labeled datasets, leading to poor generalization. To mitigate this problem, generative AI has been utilized to generate large amounts of synthetic hate…

Computation and Language · Computer Science 2023-11-17 Sagi Pendzel , Tomer Wullach , Amir Adler , Einat Minkov

We introduce Synthetic Alignment data Generation for Safety Evaluation and Red Teaming (SAGE-RT or SAGE) a novel pipeline for generating synthetic alignment and red-teaming data. Existing methods fall short in creating nuanced and diverse…

Artificial Intelligence · Computer Science 2024-08-23 Anurakt Kumar , Divyanshu Kumar , Jatan Loya , Nitin Aravind Birur , Tanay Baswa , Sahil Agarwal , Prashanth Harshangi

We study, from an empirical standpoint, the efficacy of synthetic data in real-world scenarios. Leveraging synthetic data for training perception models has become a key strategy embraced by the community due to its efficiency, scalability,…

Machine Learning · Computer Science 2024-03-26 Che-Jui Chang , Danrui Li , Seonghyeon Moon , Mubbasir Kapadia

Speech dysfluency detection is crucial for clinical diagnosis and language assessment, but existing methods are limited by the scarcity of high-quality annotated data. Although recent advances in TTS model have enabled synthetic dysfluency…

The growing number of pretrained models in Machine Learning (ML) presents significant challenges for practitioners. Given a new dataset, they need to determine the most suitable deep learning (DL) pipeline, consisting of the pretrained…

Machine Learning · Computer Science 2025-06-17 Fabio Ferreira