中文
相关论文

相关论文: SynSym: A Synthetic Data Generation Framework for …

200 篇论文

Modern studies of societal phenomena rely on the availability of large datasets capturing attributes and activities of synthetic, city-level, populations. For instance, in epidemiology, synthetic population datasets are necessary to study…

数据库 · 计算机科学 2016-02-26 Hao Wu , Yue Ning , Prithwish Chakraborty , Jilles Vreeken , Nikolaj Tatti , Naren Ramakrishnan

The usage of medical image data for the training of large-scale machine learning approaches is particularly challenging due to its scarce availability and the costly generation of data annotations, typically requiring the engagement of…

计算机视觉与模式识别 · 计算机科学 2024-06-26 Joshua Niemeijer , Jan Ehrhardt , Hristina Uzunova , Heinz Handels

In recent years, deep learning (DL) has shown great potential in the field of dermatological image analysis. However, existing datasets in this domain have significant limitations, including a small number of image samples, limited disease…

图像与视频处理 · 电气工程与系统科学 2024-04-23 Ashish Sinha , Jeremy Kawahara , Arezou Pakzad , Kumar Abhishek , Matthieu Ruthven , Enjie Ghorbel , Anis Kacem , Djamila Aouada , Ghassan Hamarneh

Large language models (LLMs) have been widely adopted for synthetic data generation, significantly reducing annotation costs. However, most existing studies treat synthesis as a set of isolated tasks and overlook a more fundamental…

人工智能 · 计算机科学 2026-05-29 Zhenlin Hu , Yan Wang , Zhen Bi , Zihao Xue , Bingyu Zhu , Longtao Huang , Xiongtao Zhang , Zeyu Yang , Zhixuan Chu , Jungang Lou

Speech dysfluency detection is crucial for clinical diagnosis and language assessment, but existing methods are limited by the scarcity of high-quality annotated data. Although recent advances in TTS model have enabled synthetic dysfluency…

Facial expression datasets remain limited in scale due to the subjectivity of annotations and the labor-intensive nature of data collection. This limitation poses a significant challenge for developing modern deep learning-based facial…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Xilin He , Cheng Luo , Xiaole Xian , Bing Li , Muhammad Haris Khan , Zongyuan Ge , Weicheng Xie , Siyang Song , Linlin Shen , Bernard Ghanem , Xiangyu Yue

We introduce SynBullying, a synthetic multi-LLM conversational dataset for studying and detecting cyberbullying (CB). SynBullying provides a scalable and ethically safe alternative to human data collection by leveraging large language…

Our ability to synthesize sensory data that preserves specific statistical properties of the real data has had tremendous implications on data privacy and big data analytics. The synthetic data can be used as a substitute for selective real…

机器学习 · 计算机科学 2017-02-01 Moustafa Alzantot , Supriyo Chakraborty , Mani B. Srivastava

Mental health is a significant and growing public health concern. As language usage can be leveraged to obtain crucial insights into mental health conditions, there is a need for large-scale, labeled, mental health-related datasets of users…

计算与语言 · 计算机科学 2018-07-12 Arman Cohan , Bart Desmet , Andrew Yates , Luca Soldaini , Sean MacAvaney , Nazli Goharian

Depression is a pervasive mental health condition that affects hundreds of millions of individuals worldwide, yet many cases remain undiagnosed due to barriers in traditional clinical access and pervasive stigma. Social media platforms, and…

计算与语言 · 计算机科学 2025-08-06 Eliseo Bao , Anxo Pérez , Javier Parapar

To train deep learning models for vision-based action recognition of elders' daily activities, we need large-scale activity datasets acquired under various daily living environments and conditions. However, most public datasets used in…

计算机视觉与模式识别 · 计算机科学 2020-11-22 Hochul Hwang , Cheongjae Jang , Geonwoo Park , Junghyun Cho , Ig-Jae Kim

Recent approaches in skill matching, employing synthetic training data for classification or similarity model training, have shown promising results, reducing the need for time-consuming and expensive annotations. However, previous…

计算与语言 · 计算机科学 2024-02-06 Antoine Magron , Anna Dai , Mike Zhang , Syrielle Montariol , Antoine Bosselut

A fundamental component of user-level social media language based clinical depression modelling is depression symptoms detection (DSD). Unfortunately, there does not exist any DSD dataset that reflects both the clinical insights and the…

计算与语言 · 计算机科学 2022-09-30 Nawshad Farruque , Randy Goebel , Sudhakar Sivapalan , Osmar Zaiane

State-of-the-art face recognition networks are often computationally expensive and cannot be used for mobile applications. Training lightweight face recognition models also requires large identity-labeled datasets. Meanwhile, there are…

计算机视觉与模式识别 · 计算机科学 2023-08-30 Hatef Otroshi Shahreza , Anjith George , Sébastien Marcel

Deep neural networks have become prevalent in human analysis, boosting the performance of applications, such as biometric recognition, action recognition, as well as person re-identification. However, the performance of such networks scales…

计算机视觉与模式识别 · 计算机科学 2022-08-22 Indu Joshi , Marcel Grimmer , Christian Rathgeb , Christoph Busch , Francois Bremond , Antitza Dantcheva

Large Language Models (LLMs) with extended context windows promise direct reasoning over long documents, reducing the need for chunking or retrieval. Constructing annotated resources for training and evaluation, however, remains costly.…

计算与语言 · 计算机科学 2025-11-13 Mohamed Elaraby , Jyoti Prakash Maheswari

Since the COVID-19 pandemic, clinicians have seen a large and sustained influx in patient portal messages, significantly contributing to clinician burnout. To the best of our knowledge, there are no large-scale public patient portal…

人工智能 · 计算机科学 2024-11-12 Joseph Gatto , Parker Seegmiller , Timothy E. Burdick , Sarah Masud Preum

Individual-level data (microdata) that characterizes a population, is essential for studying many real-world problems. However, acquiring such data is not straightforward due to cost and privacy constraints, and access is often limited to…

机器学习 · 计算机科学 2022-12-13 Angeela Acharya , Siddhartha Sikdar , Sanmay Das , Huzefa Rangwala

Persona-driven simulations are increasingly used in computational social science, yet their validity critically depends on the fidelity of the underlying personas. Constructing virtual populations that are both authentic and scalable…

计算与语言 · 计算机科学 2026-04-21 Vahid Rahimzadeh , Erfan Moosavi Monazzah , Mohammad Taher Pilehvar , Yadollah Yaghoobzadeh

Generative models capable of capturing nuanced clinical features in medical images hold great promise for facilitating clinical data sharing, enhancing rare disease datasets, and efficiently synthesizing annotated medical images at scale.…

图像与视频处理 · 电气工程与系统科学 2023-06-23 Shenghuan Sun , Gregory M. Goldgof , Atul Butte , Ahmed M. Alaa