English
Related papers

Related papers: Collaborative Filtering and the Missing at Random …

200 papers

With the rapid development of modern technology, the Web has become an important platform for users to make friends and acquire information. However, since information on the Web is over-abundant, information filtering becomes a key task…

Social and Information Networks · Computer Science 2022-07-28 Hao Liao , Qi-xin Liu , Ze-cheng Huang , Chi Ho Yeung , Yi-Cheng Zhang

Missing data in supervised learning is well-studied, but the specific issue of missing labels during model evaluation has been overlooked. Ignoring samples with missing values, a common solution, can introduce bias, especially when data is…

Machine Learning · Computer Science 2025-04-28 Danial Dervovic , Michael Cashmore

Forecasting the popularity of new songs has become a standard practice in the music industry and provides a comparative advantage for those that do it well. Considerable efforts were put into machine learning prediction models for that…

Physics and Society · Physics 2022-11-29 Niklas Reisz , Vito D. P. Servedio , Stefan Thurner

Voting online with explicit ratings could largely reflect people's preferences and objects' qualities, but ratings are always irrational, because they may be affected by many unpredictable factors like mood, weather, as well as other…

Data Analysis, Statistics and Probability · Physics 2013-05-03 Zimo Yang , Zi-Ke Zhang , Tao Zhou

Recommendation algorithms are susceptible to popularity bias: a tendency to recommend popular items even when they fail to meet user needs. A related issue is that the recommendation quality can vary by demographic groups. Marginalized…

Information Retrieval · Computer Science 2021-10-19 Nicola Neophytou , Bhaskar Mitra , Catherine Stinson

Missing Not At Random (MNAR) values lead to significant biases in the data, since the probability of missingness depends on the unobserved values.They are ''not ignorable'' in the sense that they often require defining a model for the…

Statistics Theory · Mathematics 2020-06-11 Aude Sportisse , Claire Boyer , Julie Josse

Classifying samples in incomplete datasets is a common aim for machine learning practitioners, but is non-trivial. Missing data is found in most real-world datasets and these missing values are typically imputed using established methods,…

High quality user feedback data is essential to training and evaluating a successful music recommendation system, particularly one that has to balance the needs of multiple stakeholders. Most existing music datasets suffer from noisy…

Information Retrieval · Computer Science 2021-09-17 Sasha Stoikov , Hongyi Wen

Artificial Intelligence (AI ) has been very successful in creating and predicting music playlists for online users based on their data; data received from users experience using the app such as searching the songs they like. There are lots…

Information Retrieval · Computer Science 2021-12-21 Marissa Baxter , Lisa Ha , Kirill Perfiliev , Natalie Sayre

Random forest (RF) missing data algorithms are an attractive approach for dealing with missing data. They have the desirable properties of being able to handle mixed types of missing data, they are adaptive to interactions and nonlinearity,…

Machine Learning · Statistics 2017-01-23 Fei Tang , Hemant Ishwaran

Reviews contain rich information about product characteristics and user interests and thus are commonly used to boost recommender system performance. Specifically, previous work show that jointly learning to perform review generation…

Information Retrieval · Computer Science 2022-09-13 Zhouhang Xie , Julian McAuley , Bodhisattwa Prasad Majumder

Missing data arise in most applied settings and are ubiquitous in electronic health records (EHR). When data are missing not at random (MNAR) with respect to measured covariates, sensitivity analyses are often considered. These post-hoc…

Methodology · Statistics 2023-07-11 Alexander W. Levis , Rajarshi Mukherjee , Rui Wang , Heidi Fischer , Sebastien Haneuse

Missing data can be informative. Ignoring this information can lead to misleading conclusions when the data model does not allow information to be extracted from the missing data. We propose a co-clustering model, based on the Latent Block…

Machine Learning · Computer Science 2020-10-26 Gabriel Frisch , Jean-Benoist Léger , Yves Grandvalet

Evaluation of search engines relies on assessments of search results for selected test queries, from which we would ideally like to draw conclusions in terms of relevance of the results for general (e.g., future, unknown) users. In practice…

Information Retrieval · Computer Science 2015-11-24 Thomas Demeester , Robin Aly , Djoerd Hiemstra , Dong Nguyen , Chris Develder

Understanding user preference is essential to the optimization of recommender systems. As a feedback of user's taste, rating scores can directly reflect the preference of a given user to a given product. Uncovering the latent components of…

Information Retrieval · Computer Science 2017-10-20 Junhua Chen , Wei Zeng , Junming Shao , Ge Fan

Recommendation systems rely on user-provided data to learn about item quality and provide personalized recommendations. An implicit assumption when aggregating ratings into item quality is that ratings are strong indicators of item quality.…

Information Retrieval · Computer Science 2023-07-27 Rana Shahout , Yehonatan Peisakhovsky , Sasha Stoikov , Nikhil Garg

Comprehensive evaluation of machine learning models is the key to make sure that they perform as robustly and consistently as desired. In order to summarize the experimental results and pick a winner, Critical Difference (CD) diagrams are…

Machine Learning · Computer Science 2026-05-25 Muhammad Rajabinasab , Afsaneh M. Nejad , Arthur Zimek

Most modern recommendation algorithms are data-driven: they generate personalized recommendations by observing users' past behaviors. A common assumption in recommendation is that how a user interacts with a piece of content (e.g., whether…

Computers and Society · Computer Science 2024-05-12 Sarah H. Cen , Andrew Ilyas , Jennifer Allen , Hannah Li , Aleksander Madry

Crowdsourcing systems aggregate decisions of many people to help users quickly identify high-quality options, such as the best answers to questions or interesting news stories. A long-standing issue in crowdsourcing is how option quality…

Social and Information Networks · Computer Science 2020-10-28 Keith Burghardt , Tad Hogg , Raissa M. D'Souza , Kristina Lerman , Marton Posfai

In the Internet era the information overload and the challenge to detect quality content has raised the issue of how to rank both resources and users in online communities. In this paper we develop a general ranking method that can…

Physics and Society · Physics 2016-09-23 Hao Liao , Giulio Cimini , Matus Medo