English
Related papers

Related papers: Sidorenko-Inspired Pessimistic Estimation

200 papers

Networked representations of real-world phenomena are often partially observed, which lead to incomplete networks. Analysis of such incomplete networks can lead to skewed results. We examine the following problem: given an incomplete…

Social and Information Networks · Computer Science 2015-11-23 Sucheta Soundarajan , Tina Eliassi-Rad , Brian Gallagher , Ali Pinar

A novel method to obtain hierarchical and overlapping clusters from network data -i.e., a set of nodes endowed with pairwise dissimilarities- is presented. The introduced method is hierarchical in the sense that it outputs a nested…

Social and Information Networks · Computer Science 2017-12-13 Fernando Gama , Santiago Segarra , Alejandro Ribeiro

Innovation records often exhibit "hockey-stick" patterns of abrupt, near-singular growth at the collective level. However, this macroscopic explosiveness stands in stark contrast to individual discovery, which remains bounded by cognitive…

Physics and Society · Physics 2026-02-17 Alessandro Bellina , Gabriele Di Bona , Giordano De Marzo , Vittorio Loreto

Recently, data augmentation in the semi-supervised regime, where unlabeled data vastly outnumbers labeled data, has received a considerable attention. In this paper, we describe an efficient technique for this task, exploiting a recent…

Machine Learning · Statistics 2019-06-21 Indro Spinelli , Simone Scardapane , Michele Scarpiniti , Aurelio Uncini

In this paper, we propose a novel framework that combines ensemble learning with augmented graph structures to improve the performance and robustness of semi-supervised node classification in graphs. By creating multiple augmented views of…

Machine Learning · Computer Science 2025-03-25 Maryam Abdolali , Romina Zakerian , Behnam Roshanfekr , Fardin Ayar , Mohammad Rahmati

An applied problem facing all areas of data science is harmonizing data sources. Joining data from multiple origins with unmapped and only partially overlapping features is a prerequisite to developing and testing robust, generalizable…

Convex optimization is an essential tool for modern data analysis, as it provides a framework to formulate and solve many problems in machine learning and data mining. However, general convex optimization solvers do not scale well, and…

Social and Information Networks · Computer Science 2015-07-02 David Hallac , Jure Leskovec , Stephen Boyd

We study how well we can reconstruct the 2-point clustering of galaxies on linear scales, as a function of mass and luminosity, using the halo occupation distribution (HOD) in several semi-analytical models (SAMs) of galaxy formation from…

Cosmology and Nongalactic Astrophysics · Physics 2014-06-20 Arnau Pujol , Enrique Gaztanaga

We study distributed protocols for finding all pairs of similar vectors in a large dataset. Our results pertain to a variety of discrete metrics, and we give concrete instantiations for Hamming distance. In particular, we give improved…

Data Structures and Algorithms · Computer Science 2016-11-16 Paul Beame , Cyrus Rashtchian

We study the asymptotics of large directed graphs, constrained to have certain densities of edges and/or outward $p$-stars. Our models are close cousins of exponential random graph models (ERGMs), in which edges and certain other subgraph…

Probability · Mathematics 2015-08-24 David Aristoff , Lingjiong Zhu

We present an automated data augmentation approach for image classification. We formulate the problem as Monte Carlo sampling where our goal is to approximate the optimal augmentation policies. We propose a particle filtering scheme for the…

Machine Learning · Computer Science 2021-10-18 Alexander Tsaregorodtsev , Vasileios Belagiannis

One of the best-known models in network science is preferential attachment. In this model, the probability of attaching to a node depends on the degree of all nodes in the population, and thus depends on global information. In many…

Physics and Society · Physics 2022-09-22 Watson Levens , Alex Szorkovszky , David J. T. Sumpter

Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process…

Methodology · Statistics 2015-12-03 Jeffrey Miller , Brenda Betancourt , Abbas Zaidi , Hanna Wallach , Rebecca C. Steorts

Network structure is growing popular for capturing the intrinsic relationship between large-scale variables. In the paper we propose to improve the estimation accuracy for large-dimensional factor model when a network structure between…

Methodology · Statistics 2020-01-30 Long Yu , Yong He , Xinsheng Zhang , Ji Zhu

Joining records with all other records that meet a linkage condition can result in an astronomically large number of combinations due to many-to-many relationships. For such challenging (acyclic) joins, a random sample over the join result…

Databases · Computer Science 2022-01-11 Michael Shekelyan , Graham Cormode , Peter Triantafillou , Ali Shanghooshabad , Qingzhi Ma

Recent literature in self-supervised has demonstrated significant progress in closing the gap between supervised and unsupervised methods in the image and text domains. These methods rely on domain-specific augmentations that are not…

Machine Learning · Computer Science 2021-09-02 Sajad Darabi , Shayan Fazeli , Ali Pazoki , Sriram Sankararaman , Majid Sarrafzadeh

In order to reduce overfitting, neural networks are typically trained with data augmentation, the practice of artificially generating additional training data via label-preserving transformations of existing training examples. While these…

Computer Vision and Pattern Recognition · Computer Science 2019-01-23 Cecilia Summers , Michael J. Dinneen

The fundamental result of Li, Long, and Srinivasan on approximations of set systems has become a key tool across several communities such as learning theory, algorithms, computational geometry, combinatorics and data analysis. The goal of…

Machine Learning · Computer Science 2022-09-02 Mónika Csikós , Nabil H. Mustafa

Based on a factorization of an input covariance matrix, we define a mild generalization of an upper bound of Nikolov (2015) and Li and Xie (2020) for the NP-Hard constrained maximum-entropy sampling problem (CMESP). We demonstrate that this…

Optimization and Control · Mathematics 2022-07-11 Zhongzhu Chen , Marcia Fampa , Jon Lee

For the last few decades, optimization has been developing at a fast rate. Bio-inspired optimization algorithms are metaheuristics inspired by nature. These algorithms have been applied to solve different problems in engineering, economics,…

Artificial Intelligence · Computer Science 2014-07-17 Muhammad Marwan Muhammad Fuad