中文
相关论文

相关论文: A Weighted Similarity Metric for Community Detecti…

200 篇论文

The growing environmental footprint of artificial intelligence (AI), especially in terms of storage and computation, calls for more frugal and interpretable models. Sparse models (e.g., linear, neural networks) offer a promising solution by…

机器学习 · 统计学 2025-09-23 Sylvain Sardy , Maxime van Cutsem , Xiaoyu Ma

Many real-world networks such as the gene networks, protein-protein interaction networks and metabolic networks exhibit community structures, meaning the existence of groups of densely connected vertices in the networks. Many local…

物理与社会 · 物理学 2016-03-25 Ju Xiang , Ke Hu , Yan Zhang , Mei-Hua Bao , Liang Tang , Yan-Ni Tang , Yuan-Yuan Gao , Jian-Ming Li , Benyan Chen , Jing-Bo Hu

Recent developments in natural language processing (NLP) have highlighted the need for substantial amounts of data for models to capture textual information accurately. This raises concerns regarding the computational resources and time…

机器学习 · 计算机科学 2024-02-26 Roxana Petcu , Subhadeep Maji

We present skweak, a versatile, Python-based software toolkit enabling NLP developers to apply weak supervision to a wide range of NLP tasks. Weak supervision is an emerging machine learning paradigm based on a simple idea: instead of…

计算与语言 · 计算机科学 2021-08-18 Pierre Lison , Jeremy Barnes , Aliaksandr Hubin

Neural network models are widely used in solving many challenging problems, such as computer vision, personalized recommendation, and natural language processing. Those models are very computationally intensive and reach the hardware limit…

机器学习 · 计算机科学 2020-04-28 Fei Sun , Minghai Qin , Tianyun Zhang , Liu Liu , Yen-Kuang Chen , Yuan Xie

We consider message-efficient continuous random sampling from a distributed stream, where the probability of inclusion of an item in the sample is proportional to a weight associated with the item. The unweighted version, where all weights…

数据结构与算法 · 计算机科学 2019-04-09 Rajesh Jayaram , Gokarna Sharma , Srikanta Tirthapura , David P. Woodruff

Sparse modelling or model selection with categorical data is challenging even for a moderate number of variables, because one parameter is roughly needed to encode one category or level. The Group Lasso is a well known efficient algorithm…

统计方法学 · 统计学 2022-11-14 Szymon Nowakowski , Piotr Pokarowski , Wojciech Rejchel , Agnieszka Sołtys

Optimal data aggregation aimed at maximizing IoT network lifetime by minimizing constrained on-board resource utilization continues to be a challenging task. The existing data aggregation methods have proven that compressed sensing is…

信号处理 · 电气工程与系统科学 2018-06-14 Amarlingam M , Pradeep Kumar Mishra , P Rajalakshmi , Sumohana S. Channappayya , C. S. Sastry

Language model fusion helps smart assistants recognize words which are rare in acoustic data but abundant in text-only corpora (typed search logs). However, such corpora have properties that hinder downstream performance, including being…

计算与语言 · 计算机科学 2022-06-16 W. Ronny Huang , Cal Peyser , Tara N. Sainath , Ruoming Pang , Trevor Strohman , Shankar Kumar

For some or all of the data instances a number of independent-world clustering issues suffer from incomplete data characterization due to losing or absent attributes. Typical clustering approaches cannot be applied directly to such data…

机器学习 · 计算机科学 2020-02-25 Y. A. Joarder , Emran Hossain , Al Faisal Mahmud

Demixing problems in many areas such as hyperspectral imaging and differential optical absorption spectroscopy (DOAS) often require finding sparse nonnegative linear combinations of dictionary elements that match observed data. We show how…

机器学习 · 统计学 2013-01-04 Ernie Esser , Yifei Lou , Jack Xin

Not identical but similar objects are ubiquitous in our world, ranging from four-legged animals such as dogs and cats to cars of different models and flowers of various colors. This study addresses a novel task of matching such…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Yusuke Marumo , Kazuhiko Kawamoto , Satomi Tanaka , Shigenobu Hirano , Hiroshi Kera

The short text has been the prevalent format for information of Internet in recent decades, especially with the development of online social media, whose millions of users generate a vast number of short messages everyday. Although…

计算与语言 · 计算机科学 2014-12-18 Yuan Zuo , Jichang Zhao , Ke Xu

Weak topic correlation across document collections with different numbers of topics in individual collections presents challenges for existing cross-collection topic models. This paper introduces two probabilistic topic models, Correlated…

计算与语言 · 计算机科学 2015-08-20 Jingwei Zhang , Aaron Gerow , Jaan Altosaar , James Evans , Richard Jean So

Evaluating LLMs and text-to-image models is a computationally intensive task often overlooked. Efficient evaluation is crucial for understanding the diverse capabilities of these models and enabling comparisons across a growing number of…

One of the major issues in signed networks is to use network structure to predict the missing sign of an edge. In this paper, we introduce a novel probabilistic approach for the sign prediction problem. The main characteristic of the…

社会与信息网络 · 计算机科学 2018-02-20 Amin Javari , HongXiang Qiu , Elham Barzegaran , Mahdi Jalili , Kevin Chen-Chuan Chang

As one of the most commonly seen data challenges, missing data, in particular, multiple, non-monotone missing patterns, complicates estimation and inference due to the fact that missingness mechanisms are often not missing at random, and…

统计方法学 · 统计学 2025-04-21 Jianing Dong , Raymond K. W. Wong , Kwun Chuen Gary Chan

Objective: Social-environmental data obtained from the U.S. Census is an important resource for understanding health disparities, but rarely is the full dataset utilized for analysis. A barrier to incorporating the full data is a lack of…

应用统计 · 统计学 2020-09-02 Elizabeth Handorf , Yinuo Yin , Michael Slifker , Shannon Lynch

We present Sampled Weighted Min-Hashing (SWMH), a randomized approach to automatically mine topics from large-scale corpora. SWMH generates multiple random partitions of the corpus vocabulary based on term co-occurrence and agglomerates…

机器学习 · 计算机科学 2015-09-09 Gibran Fuentes-Pineda , Ivan Vladimir Meza-Ruiz

Personality refers to individual differences in behavior, thinking, and feeling. With the growing availability of digital footprints, especially from social media, automated methods for personality assessment have become increasingly…

计算与语言 · 计算机科学 2025-10-06 Matej Gjurković