中文
相关论文

相关论文: Clustering and Labelling Auction Fraud Data

200 篇论文

Experimental evaluation is a major research methodology for investigating clustering algorithms and many other machine learning algorithms. For this purpose, a number of benchmark datasets have been widely used in the literature and their…

机器学习 · 计算机科学 2019-10-21 Tiantian Zhang , Li Zhong , Bo Yuan

In supervised machine learning, use of correct labels is extremely important to ensure high accuracy. Unfortunately, most datasets contain corrupted labels. Machine learning models trained on such datasets do not generalize well. Thus,…

机器学习 · 计算机科学 2023-09-14 Chang Yue , Niraj K. Jha

Market segmentation of an online auction site is studied by analyzing the users' bidding behavior. The distribution of user activity is investigated and a network of bidders connected by common interest in individual articles is…

物理与社会 · 物理学 2008-12-11 Joerg Reichardt , Stefan Bornholdt

Clustering is an important data mining technique where we will be interested in maximizing intracluster distance and also minimizing intercluster distance. We have utilized clustering techniques for detecting deviation in product sales and…

数据库 · 计算机科学 2013-12-11 S. Hanumanth Sastry , Prof. M. S. Prasada Babu

Online auctions play a central role in online advertising, and are one of the main reasons for the industry's scalability and growth. With great changes in how auctions are being organized, such as changing the second- to first-price…

计算机科学与博弈论 · 计算机科学 2020-09-04 Djordje Gligorijevic , Tian Zhou , Bharatbhushan Shetty , Brendan Kitts , Shengjun Pan , Junwei Pan , Aaron Flores

In this paper we propose a measure of clustering quality or accuracy that is appropriate in situations where it is desirable to evaluate a clustering algorithm by somehow comparing the clusters it produces with ``ground truth' consisting of…

机器学习 · 计算机科学 2013-01-07 Byron E Dom

In online advertising, a set of potential advertisements can be ranked by a certain auction system where usually the top-1 advertisement would be selected and displayed at an advertising space. In this paper, we show a selection bias issue…

信息检索 · 计算机科学 2022-06-09 Shinya Suzumura , Hitoshi Abe

There are many cluster analysis methods that can produce quite different clusterings on the same dataset. Cluster validation is about the evaluation of the quality of a clustering; "relative cluster validation" is about using such criteria…

统计方法学 · 统计学 2020-09-10 Christian Hennig

We propose a learning setting in which unlabeled data is free, and the cost of a label depends on its value, which is not known in advance. We study binary classification in an extreme case, where the algorithm only pays for negative…

机器学习 · 计算机科学 2015-07-14 Sivan Sabato , Anand D. Sarwate , Nathan Srebro

Clustering algorithms are widely utilized for many modern data science applications. This motivates the need to make outputs of clustering algorithms fair. Traditionally, new fair algorithmic variants to clustering algorithms are developed…

机器学习 · 计算机科学 2021-10-26 Anshuman Chhabra , Adish Singla , Prasant Mohapatra

Given a set of financial transactions (who buys from whom, when, and for how much), as well as prior information from buyers and sellers, how can we find fraudulent transactions? If we have labels for some transactions for known types of…

机器学习 · 计算机科学 2025-10-07 Robson L. F. Cordeiro , Meng-Chieh Lee , Christos Faloutsos

With the increasing scale of search engine marketing, designing an efficient bidding system is becoming paramount for the success of e-commerce companies. The critical challenges faced by a modern industrial-level bidding system include: 1.…

计算与语言 · 计算机科学 2021-08-09 Cheng Jie , Da Xu , Zigeng Wang , Lu Wang , Wei Shen

Clustering a graph, i.e., assigning its nodes to groups, is an important operation whose best known application is the discovery of communities in social networks. Graph clustering and community detection have traditionally focused on…

社会与信息网络 · 计算机科学 2015-01-09 Cecile Bothorel , Juan David Cruz , Matteo Magnani , Barbora Micenkova

Auctions are important mechanisms extensively implemented in various markets, e.g., search engines' keyword auctions, antique auctions, etc. Finding an optimal auction mechanism is extremely difficult due to the constraints of imperfect…

机器学习 · 计算机科学 2025-07-28 Jiayin Liu , Chenglong Zhang

In an illiquid stock, traders can collude and place orders on a predetermined price and quantity at a fixed schedule. This is usually done to manipulate the price of the stock or to create artificial liquidity in the stock, which may…

交易与市场微观结构 · 定量金融 2016-10-18 Suneel Sarswat , Kandathil Mathew Abraham , Subir Kumar Ghosh

Mapping of spatial hotspots, i.e., regions with significantly higher rates of generating cases of certain events (e.g., disease or crime cases), is an important task in diverse societal domains, including public health, public safety,…

机器学习 · 统计学 2021-10-12 Yiqun Xie , Shashi Shekhar , Yan Li

Learning to detect real-world anomalous events using video-level annotations is a difficult task mainly because of the noise present in labels. An anomalous labelled video may actually contain anomaly only in a short duration while the rest…

计算机视觉与模式识别 · 计算机科学 2021-05-03 Muhammad Zaigham Zaheer , Jin-ha Lee , Marcella Astrid , Arif Mahmood , Seung-Ik Lee

Clustering is considered a non-supervised learning setting, in which the goal is to partition a collection of data points into disjoint clusters. Often a bound $k$ on the number of clusters is given or assumed by the practitioner. Many…

机器学习 · 计算机科学 2012-02-01 Nir Ailon , Ron Begleiter

Donations to charity-based crowdfunding environments have been on the rise in the last few years. Unsurprisingly, deception and fraud in such platforms have also increased, but have not been thoroughly studied to understand what…

计算机与社会 · 计算机科学 2020-07-01 Beatrice Perez , Sara R. Machado , Jerone T. A. Andrews , Nicolas Kourtellis

With recent advances in data collection from multiple sources, multi-view data has received significant attention. In multi-view data, each view represents a different perspective of data. Since label information is often expensive to…

机器学习 · 计算机科学 2021-05-10 Mehrnaz Najafi , Lifang He , Philip S. Yu