English
Related papers

Related papers: Neighbor Embedding for High-Dimensional Sparse Poi…

200 papers

The rise of internet has resulted in an explosion of data consisting of millions of articles, images, songs, and videos. Most of this data is high dimensional and sparse. The need to perform an efficient search for similar objects in such…

Data Structures and Algorithms · Computer Science 2016-12-20 Raghav Kulkarni , Rameshwar Pratap

Unsupervised dimensionality reduction is one of the commonly used techniques in the field of high dimensional data recognition problems. The deep autoencoder network which constrains the weights to be non-negative, can learn a low…

Computer Vision and Pattern Recognition · Computer Science 2020-09-18 Anyong Qin , Zhaowei Shang , Zhuolin Tan , Taiping Zhang , Yuan Yan Tang

We are motivated by problems that arise in a number of applications such as Online Marketing and explosives detection, where the observations are usually modeled using Poisson statistics. We model each observation as a Poisson random…

Statistics Theory · Mathematics 2016-11-17 D. Motamedvaziri , M. H. Rohban , V. Saligrama

We present the performance of a semantic segmentation network, SparseSSNet, that provides pixel-level classification of MicroBooNE data. The MicroBooNE experiment employs a liquid argon time projection chamber for the study of neutrino…

Instrumentation and Detectors · Physics 2021-04-07 MicroBooNE collaboration , P. Abratenko , M. Alrashed , R. An , J. Anthony , J. Asaadi , A. Ashkenazi , S. Balasubramanian , B. Baller , C. Barnes , G. Barr , V. Basque , L. Bathe-Peters , O. Benevides Rodrigues , S. Berkman , A. Bhanderi , A. Bhat , M. Bishai , A. Blake , T. Bolton , L. Camilleri , D. Caratelli , I. Caro Terrazas , R. Castillo Fernandez , F. Cavanna , G. Cerati , Y. Chen , E. Church , D. Cianci , J. M. Conrad , M. Convery , L. Cooper-Troendle , J. I. Crespo-Anadon , M. Del Tutto , S. R. Dennis , D. Devitt , R. Diurba , R. Dorrill , K. Duffy , S. Dytman , B. Eberly , A. Ereditato , J. J. Evans , G. A. Fiorentini Aguirre , R. S. Fitzpatrick , B. T. Fleming , N. Foppiani , D. Franco , A. P. Furmanski , D. Garcia-Gamez , S. Gardiner , G. Ge , S. Gollapinni , O. Goodwin , E. Gramellini , P. Green , H. Greenlee , W. Gu , R. Guenette , P. Guzowski , L. Hagaman , E. Hall , P. Hamilton , O. Hen , G. A. Horton-Smith , A. Hourlier , R. Itay , C. James , J. Jan de Vries , X. Ji , L. Jiang , J. H. Jo , R. A. Johnson , Y. J. Jwa , N. Kamp , N. Kaneshige , G. Karagiorgi , W. Ketchum , B. Kirby , M. Kirby , T. Kobilarcik , I. Kreslo , R. LaZur , I. Lepetic , K. Li , Y. Li , B. R. Littlejohn , W. C. Louis , X. Luo , A. Marchionni , C. Mariani , D. Marsden , J. Marshall , J. Martin-Albo , D. A. Martinez Caicedo , K. Mason , A. Mastbaum , N. McConkey , V. Meddage , T. Mettler , K. Miller , J. Mills , K. Mistry , T. Mohayai , A. Mogan , J. Moon , M. Mooney , A. F. Moor , C. D. Moore , L. Mora Lepin , J. Mousseau , M. Murphy , D. Naples , A. Navrer-Agasson , R. K. Neely , P. Nienaber , J. Nowak , O. Palamara , V. Paolone , A. Papadopoulou , V. Papavassiliou , S. F. Pate , A. Paudel , Z. Pavlovic , E. Piasetzky , I. Ponce-Pinto , S. Prince , X. Qian , J. L. Raaf , V. Radeka , A. Rafique , M. Reggiani-Guzzo , L. Ren , L. Rochester , J. Rodriguez Rondon , H. E. Rogers , M. Rosenberg , M. Ross-Lonergan , B. Russell , G. Scanavini , D. W. Schmitz , A. Schukraft , W. Seligman , M. H. Shaevitz , R. Sharankova , J. Sinclair , A. Smith , E. L. Snider , M. Soderberg , S. Soldner-Rembold , S. R. Soleti , P. Spentzouris , J. Spitz , M. Stancari , J. St. John , T. Strauss , K. Sutton , S. Sword-Fehlberg , A. M. Szelc , N. Tagg , W. Tang , K. Terao , C. Thorpe , M. Toups , Y. -T. Tsai , M. A. Uchida , T. Usher , W. Van De Pontseele , B. Viren , M. Weber , H. Wei , Z. Williams , S. Wolbers , T. Wongjirad , M. Wospakrik , W. Wu , E. Yandel , T. Yang , G. Yarbrough , L. E. Yates , G. P. Zeller , J. Zennamo , C. Zhang

Previous research on word embeddings has shown that sparse representations, which can be either learned on top of existing dense embeddings or obtained through model constraints during training time, have the benefit of increased…

Computation and Language · Computer Science 2018-09-26 Valentin Trifonov , Octavian-Eugen Ganea , Anna Potapenko , Thomas Hofmann

Information networks are ubiquitous and are ideal for modeling relational data. Networks being sparse and irregular, network embedding algorithms have caught the attention of many researchers, who came up with numerous embeddings algorithms…

Machine Learning · Computer Science 2020-09-25 Junshan Wang , Yilun Jin , Guojie Song , Xiaojun Ma

The sparse modeling is an evident manifestation capturing the parsimony principle just described, and sparse models are widespread in statistics, physics, information sciences, neuroscience, computational mathematics, and so on. In…

Machine Learning · Computer Science 2023-08-29 Jianyi Lin

The development of data-dependent heuristics and representations for biological sequences that reflect their evolutionary distance is critical for large-scale biological research. However, popular machine learning approaches, based on…

Quantitative Methods · Quantitative Biology 2021-10-13 Gabriele Corso , Rex Ying , Michal Pándy , Petar Veličković , Jure Leskovec , Pietro Liò

Dimensionality reduction methods, also known as projections, are frequently used for exploring multidimensional data in machine learning, data science, and information visualization. Among these, t-SNE and its variants have become very…

Machine Learning · Computer Science 2019-02-22 Mateus Espadoto , Nina S. T. Hirata , Alexandru C. Telea

In Near-Neighbor Search (NNS), a new client queries a database (held by a server) for the most similar data (near-neighbors) given a certain similarity metric. The Privacy-Preserving variant (PP-NNS) requires that neither server nor the…

Cryptography and Security · Computer Science 2019-10-18 M. Sadegh Riazi , Beidi Chen , Anshumali Shrivastava , Dan Wallach , Farinaz Koushanfar

Information distances like the Hellinger distance and the Jensen-Shannon divergence have deep roots in information theory and machine learning. They are used extensively in data analysis especially when the objects being compared are high…

Data Structures and Algorithms · Computer Science 2015-03-19 Amirali Abdullah , Ravi Kumar , Andrew McGregor , Sergei Vassilvitskii , Suresh Venkatasubramanian

We propose a new method for estimating the intrinsic dimension of a dataset by applying the principle of regularized maximum likelihood to the distances between close neighbors. We propose a regularization scheme which is motivated by…

Machine Learning · Computer Science 2012-03-19 Mithun Das Gupta , Thomas S. Huang

Calculation of near-neighbor interactions among high dimensional, irregularly distributed data points is a fundamental task to many graph-based or kernel-based machine learning algorithms and applications. Such calculations, involving…

Dense high dimensional vectors are becoming increasingly vital in fields such as computer vision, machine learning, and large language models (LLMs), serving as standard representations for multimodal data. Now the dimensionality of these…

Machine Learning · Computer Science 2024-10-10 Zhonghan Chen , Ruiyuan Zhang , Xi Zhao , Xiaojun Cheng , Xiaofang Zhou

Deep neural networks often require copious amount of labeled-data to train their scads of parameters. Training larger and deeper networks is hard without appropriate regularization, particularly while using a small dataset. Laterally,…

Computer Vision and Pattern Recognition · Computer Science 2019-05-31 Xiang Xu , Xiong Zhou , Ragav Venkatesan , Gurumurthy Swaminathan , Orchid Majumder

Recent advancements in discrete image generation showed that scaling the VQ codebook size significantly improves reconstruction fidelity. However, training generative models with a large VQ codebook remains challenging, typically requiring…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Shufan Li , Jiuxiang Gu , Kangning Liu , Zhe Lin , Aditya Grover , Jason Kuen

Today's densely instrumented world offers tremendous opportunities for continuous acquisition and analysis of multimodal sensor data providing temporal characterization of an individual's behaviors. Is it possible to efficiently couple such…

Machine Learning · Computer Science 2018-09-03 Homa Hosseinmardi , Amir Ghasemian , Shrikanth Narayanan , Kristina Lerman , Emilio Ferrara

Neural networks have seen limited use in prediction for high-dimensional data with small sample sizes, because they tend to overfit and require tuning many more hyperparameters than existing off-the-shelf machine learning methods. With…

Machine Learning · Statistics 2020-05-12 Jean Feng , Noah Simon

High-energy large-scale particle colliders generate data at extraordinary rates. Developing real-time high-throughput data compression algorithms to reduce data volume and meet the bandwidth requirement for storage has become increasingly…

Network Embeddings (NEs) map the nodes of a given network into $d$-dimensional Euclidean space $\mathbb{R}^d$. Ideally, this mapping is such that `similar' nodes are mapped onto nearby points, such that the NE can be used for purposes such…

Machine Learning · Statistics 2018-10-17 Bo Kang , Jefrey Lijffijt , Tijl De Bie