中文
相关论文

相关论文: A Semi-Synthetic Dataset Generation Framework for …

200 篇论文

In the fundamental statistics course, students are taught to remember the well-known saying: "Correlation is not Causation". Till now, statistics (i.e., correlation) have developed various successful frameworks, such as Transformer and…

人工智能 · 计算机科学 2023-11-22 Ning Xu , Yifei Gao , Hongshuo Tian , Yongdong Zhang , An-An Liu

Commercial establishments like restaurants, service centres and retailers have several sources of customer feedback about products and services, most of which need not be as structured as rated reviews provided by services like Yelp, or…

计算与语言 · 计算机科学 2017-03-28 Vineet John

The importance of Synthetic Data Generation (SDG) has increased significantly in domains where data quality is poor or access is limited due to privacy and regulatory constraints. One such domain is recruitment, where publicly available…

机器学习 · 计算机科学 2025-11-24 Andrea Iommi , Antonio Mastropietro , Riccardo Guidotti , Anna Monreale , Salvatore Ruggieri

A huge amount of user generated content related to movies is created with the popularization of web 2.0. With these continues exponential growth of data, there is an inevitable need for recommender systems as people find it difficult to…

信息检索 · 计算机科学 2019-06-04 Lasitha Uyangoda , Supunmali Ahangama , Tharindu Ranasinghe

Code smell is a great challenge in software refactoring, which indicates latent design or implementation flaws that may degrade the software maintainability and evolution. Over the past of decades, the research on code smell has received…

软件工程 · 计算机科学 2026-04-21 Hanyu Zhang , Tomoji Kishi

Creating challenging tabular inference data is essential for learning complex reasoning. Prior work has mostly relied on two data generation strategies. The first is human annotation, which yields linguistically diverse data but is…

计算与语言 · 计算机科学 2022-11-24 Aashna Jena , Vivek Gupta , Manish Shrivastava , Julian Martin Eisenschlos

Knowledge Graph (KG), as a side-information, tends to be utilized to supplement the collaborative filtering (CF) based recommendation model. By mapping items with the entities in KGs, prior studies mostly extract the knowledge information…

信息检索 · 计算机科学 2022-12-21 Yinwei Wei , Xiang Wang , Liqiang Nie , Shaoyu Li , Dingxian Wang , Tat-Seng Chua

Causal discovery from observational data is an important tool in many branches of science. Under certain assumptions it allows scientists to explain phenomena, predict, and make decisions. In the large sample limit, sound and complete…

机器学习 · 统计学 2021-07-13 Shami Nisimov , Yaniv Gurwicz , Raanan Y. Rohekar , Gal Novik

Sequential recommendation (SR) is traditionally formulated as next-item prediction over a chronological sequence of interacted items. Although recent generative recommendation (GR) methods introduce new machinery, such as semantic IDs,…

Recommender-system datasets are used for recommender-system evaluations, training machine-learning algorithms, and exploring user behavior. While there are many datasets for recommender systems in the domains of movies, books, and music,…

信息检索 · 计算机科学 2017-06-21 Joeran Beel , Zeljko Carevic , Johann Schaible , Gabor Neusch

Robust causal discovery in time series datasets depends on reliable benchmark datasets with known ground-truth causal relationships. However, such datasets remain scarce, and existing synthetic alternatives often overlook critical temporal…

机器学习 · 计算机科学 2025-06-03 Muhammad Hasan Ferdous , Emam Hossain , Md Osman Gani

Recommendation is a prevalent and critical service in information systems. To provide personalized suggestions to users, industry players embrace machine learning, more specifically, building predictive models based on the click behavior…

信息检索 · 计算机科学 2021-05-25 Wenjie Wang , Fuli Feng , Xiangnan He , Hanwang Zhang , Tat-Seng Chua

Conversational recommender systems (CRS) aim to recommend high-quality items to users through interactive conversations. To develop an effective CRS, the support of high-quality datasets is essential. Existing CRS datasets mainly focus on…

计算与语言 · 计算机科学 2020-11-03 Kun Zhou , Yuanhang Zhou , Wayne Xin Zhao , Xiaoke Wang , Ji-Rong Wen

Simulating a recommendation system in a controlled environment, to identify specific behaviors and user preferences, requires highly flexible synthetic data generation models capable of mimicking the patterns and trends of real datasets. In…

信息检索 · 计算机科学 2025-05-19 Simone Mungari , Erica Coppolillo , Ettore Ritacco , Giuseppe Manco

Most of the existing recommender systems are based only on the rating data, and they ignore other sources of information that might increase the quality of recommendations, such as textual reviews, or user and item characteristics.…

信息检索 · 计算机科学 2021-11-17 Tatev Karen Aslanyan , Flavius Frasincar

In biomedical research, repeated measurements within each subject are often processed to remove artifacts and unwanted sources of variation. The resulting data are used to construct derived outcomes that act as proxies for scientific…

统计方法学 · 统计学 2026-02-03 Zihang Wang , Razieh Nabi , Benjamin B. Risk

Traditional recommender systems aim to estimate a user's rating to an item based on observed ratings from the population. As with all observational studies, hidden confounders, which are factors that affect both item exposures and user…

机器学习 · 计算机科学 2022-11-22 Yaochen Zhu , Jing Yi , Jiayi Xie , Zhenzhong Chen

Click-through rate (CTR) prediction plays an indispensable role in online platforms. Numerous models have been proposed to capture users' shifting preferences by leveraging user behavior sequences. However, these historical sequences often…

信息检索 · 计算机科学 2024-04-16 Junjie Huang , Guohao Cai , Jieming Zhu , Zhenhua Dong , Ruiming Tang , Weinan Zhang , Yong Yu

Albeit, the implicit feedback based recommendation problem - when only the user history is available but there are no ratings - is the most typical setting in real-world applications, it is much less researched than the explicit feedback…

机器学习 · 计算机科学 2013-04-05 Balázs Hidasi , Domonkos Tikk

Large Language Models (LLMs) generate realistic synthetic data but offer no guarantee that their outputs respect the causal mechanisms governing the target domain. We introduce CausalSynth, a framework that decouples causal structure…

机器学习 · 计算机科学 2026-05-19 Zehua Cheng , Wei Dai , Jiahao Sun , Thomas Lukasiewicz