中文
相关论文

相关论文: Collaborative Filtering and the Missing at Random …

200 篇论文

With the rapid development of modern technology, the Web has become an important platform for users to make friends and acquire information. However, since information on the Web is over-abundant, information filtering becomes a key task…

社会与信息网络 · 计算机科学 2022-07-28 Hao Liao , Qi-xin Liu , Ze-cheng Huang , Chi Ho Yeung , Yi-Cheng Zhang

Missing data in supervised learning is well-studied, but the specific issue of missing labels during model evaluation has been overlooked. Ignoring samples with missing values, a common solution, can introduce bias, especially when data is…

机器学习 · 计算机科学 2025-04-28 Danial Dervovic , Michael Cashmore

Forecasting the popularity of new songs has become a standard practice in the music industry and provides a comparative advantage for those that do it well. Considerable efforts were put into machine learning prediction models for that…

物理与社会 · 物理学 2022-11-29 Niklas Reisz , Vito D. P. Servedio , Stefan Thurner

Voting online with explicit ratings could largely reflect people's preferences and objects' qualities, but ratings are always irrational, because they may be affected by many unpredictable factors like mood, weather, as well as other…

数据分析、统计与概率 · 物理学 2013-05-03 Zimo Yang , Zi-Ke Zhang , Tao Zhou

Recommendation algorithms are susceptible to popularity bias: a tendency to recommend popular items even when they fail to meet user needs. A related issue is that the recommendation quality can vary by demographic groups. Marginalized…

信息检索 · 计算机科学 2021-10-19 Nicola Neophytou , Bhaskar Mitra , Catherine Stinson

Missing Not At Random (MNAR) values lead to significant biases in the data, since the probability of missingness depends on the unobserved values.They are ''not ignorable'' in the sense that they often require defining a model for the…

统计理论 · 数学 2020-06-11 Aude Sportisse , Claire Boyer , Julie Josse

Classifying samples in incomplete datasets is a common aim for machine learning practitioners, but is non-trivial. Missing data is found in most real-world datasets and these missing values are typically imputed using established methods,…

High quality user feedback data is essential to training and evaluating a successful music recommendation system, particularly one that has to balance the needs of multiple stakeholders. Most existing music datasets suffer from noisy…

信息检索 · 计算机科学 2021-09-17 Sasha Stoikov , Hongyi Wen

Artificial Intelligence (AI ) has been very successful in creating and predicting music playlists for online users based on their data; data received from users experience using the app such as searching the songs they like. There are lots…

信息检索 · 计算机科学 2021-12-21 Marissa Baxter , Lisa Ha , Kirill Perfiliev , Natalie Sayre

Random forest (RF) missing data algorithms are an attractive approach for dealing with missing data. They have the desirable properties of being able to handle mixed types of missing data, they are adaptive to interactions and nonlinearity,…

机器学习 · 统计学 2017-01-23 Fei Tang , Hemant Ishwaran

Reviews contain rich information about product characteristics and user interests and thus are commonly used to boost recommender system performance. Specifically, previous work show that jointly learning to perform review generation…

信息检索 · 计算机科学 2022-09-13 Zhouhang Xie , Julian McAuley , Bodhisattwa Prasad Majumder

Missing data arise in most applied settings and are ubiquitous in electronic health records (EHR). When data are missing not at random (MNAR) with respect to measured covariates, sensitivity analyses are often considered. These post-hoc…

统计方法学 · 统计学 2023-07-11 Alexander W. Levis , Rajarshi Mukherjee , Rui Wang , Heidi Fischer , Sebastien Haneuse

Missing data can be informative. Ignoring this information can lead to misleading conclusions when the data model does not allow information to be extracted from the missing data. We propose a co-clustering model, based on the Latent Block…

机器学习 · 计算机科学 2020-10-26 Gabriel Frisch , Jean-Benoist Léger , Yves Grandvalet

Evaluation of search engines relies on assessments of search results for selected test queries, from which we would ideally like to draw conclusions in terms of relevance of the results for general (e.g., future, unknown) users. In practice…

信息检索 · 计算机科学 2015-11-24 Thomas Demeester , Robin Aly , Djoerd Hiemstra , Dong Nguyen , Chris Develder

Understanding user preference is essential to the optimization of recommender systems. As a feedback of user's taste, rating scores can directly reflect the preference of a given user to a given product. Uncovering the latent components of…

信息检索 · 计算机科学 2017-10-20 Junhua Chen , Wei Zeng , Junming Shao , Ge Fan

Recommendation systems rely on user-provided data to learn about item quality and provide personalized recommendations. An implicit assumption when aggregating ratings into item quality is that ratings are strong indicators of item quality.…

信息检索 · 计算机科学 2023-07-27 Rana Shahout , Yehonatan Peisakhovsky , Sasha Stoikov , Nikhil Garg

Comprehensive evaluation of machine learning models is the key to make sure that they perform as robustly and consistently as desired. In order to summarize the experimental results and pick a winner, Critical Difference (CD) diagrams are…

机器学习 · 计算机科学 2026-05-25 Muhammad Rajabinasab , Afsaneh M. Nejad , Arthur Zimek

Most modern recommendation algorithms are data-driven: they generate personalized recommendations by observing users' past behaviors. A common assumption in recommendation is that how a user interacts with a piece of content (e.g., whether…

计算机与社会 · 计算机科学 2024-05-12 Sarah H. Cen , Andrew Ilyas , Jennifer Allen , Hannah Li , Aleksander Madry

Crowdsourcing systems aggregate decisions of many people to help users quickly identify high-quality options, such as the best answers to questions or interesting news stories. A long-standing issue in crowdsourcing is how option quality…

社会与信息网络 · 计算机科学 2020-10-28 Keith Burghardt , Tad Hogg , Raissa M. D'Souza , Kristina Lerman , Marton Posfai

In the Internet era the information overload and the challenge to detect quality content has raised the issue of how to rank both resources and users in online communities. In this paper we develop a general ranking method that can…

物理与社会 · 物理学 2016-09-23 Hao Liao , Giulio Cimini , Matus Medo