中文
相关论文

相关论文: The hypergeometric test performs comparably to TF-…

200 篇论文

Journal Impact Factors (IFs) can be considered historically as the first attempt to normalize citation distributions by using averages over two years. However, it has been recognized that citation distributions vary among fields of science…

数字图书馆 · 计算机科学 2012-02-07 Loet Leydesdorff

Deep neural networks have achieved significant improvements in information retrieval (IR). However, most existing models are computational costly and can not efficiently scale to long documents. This paper proposes a novel End-to-End neural…

计算与语言 · 计算机科学 2019-08-13 Chen Zheng , Yu Sun , Shengxian Wan , Dianhai Yu

The demand for accurate and efficient verification of information in texts generated by large language models (LMs) is at an all-time high, but remains unresolved. Recent efforts have focused on extracting and verifying atomic facts from…

计算与语言 · 计算机科学 2024-09-25 Vasileios Katranidis , Gabor Barany

A strategy is presented to incorporate prior information from conceptual geological models in probabilistic inversion of geophysical data. The conceptual geological models are represented by multiple-point statistics training images (TIs)…

地球物理 · 物理学 2017-01-06 T. Lochbühler , J. A. Vrugt , M. Sadegh , N. Linde

Graphical models have found widespread applications in many areas of modern statistics and machine learning. Iterative Proportional Fitting (IPF) and its variants have become the default method for undirected graphical model estimation, and…

统计方法学 · 统计学 2024-08-22 Kshitij Khare , Syed Rahman , Bala Rajaratnam , Jiayuan Zhou

We develop a test of normality for spatially indexed functions. The assumption of normality is common in spatial statistics, yet no significance tests, or other means of assessment, have been available for functional data. This paper aims…

统计方法学 · 统计学 2021-07-01 Thomas Kuenzer , Siegfried Hörmann , Piotr Kokoszka

Comparison of two probability density/mass functions (PDF/PMFs) is ubiquitous in various forms of scientific analysis, including machine learning, optimization problems, and hypothesis tests. A copious amount of distance metrics have…

核实验 · 物理学 2026-04-16 Nafis Fuad

There have been multiple attempts to resolve various inflection matching problems in information retrieval. Stemming is a common approach to this end. Among many techniques for stemming, statistical stemming has been shown to be effective…

信息检索 · 计算机科学 2016-06-22 Javid Dadashkarimi , Hossein Nasr Esfahani , Heshaam Faili , Azadeh Shakery

Term frequency is a common method for identifying the importance of a term in a query or document. But it is a weak signal, especially when the frequency distribution is flat, such as in long queries or short documents where the text is of…

信息检索 · 计算机科学 2019-11-28 Zhuyun Dai , Jamie Callan

Item difficulty plays a crucial role in test performance, interpretability of scores, and equity for all test-takers, especially in large-scale assessments. Traditional approaches to item difficulty modeling rely on field testing and…

计算与语言 · 计算机科学 2025-09-30 Sydney Peters , Nan Zhang , Hong Jiao , Ming Li , Tianyi Zhou , Robert Lissitz

Statistical depth is the act of gauging how representative a point is compared to a reference probability measure. The depth allows introducing rankings and orderings to data living in multivariate, or function spaces. Though widely applied…

统计理论 · 数学 2021-05-28 George Wynne , Stanislav Nagy

Nowadays impact factor is the significant indicator for journal evaluation. In impact factor calculation is used number of all citations to journal, regardless of the prestige of cited journals, however, scientific units (paper, researcher,…

数字图书馆 · 计算机科学 2015-06-10 Rasim Alguliyev , Ramiz Aliguliyev , Nigar Ismayilova

The semantics are derived from textual data that provide representations for Machine Learning algorithms. These representations are interpretable form of high dimensional sparse matrix that are given as an input to the machine learning…

机器学习 · 计算机科学 2022-02-08 Sayali Tambe , Raunak Joshi , Abhishek Gupta , Nandan Kanvinde , Vidya Chitre

Real-world complex systems often comprise many distinct types of elements as well as many more types of networked interactions between elements. When the relative abundances of types can be measured well, we often observe heavy-tailed…

物理与社会 · 物理学 2025-03-17 P. S. Dodds , J. R. Minot , M. V. Arnold , T. Alshaabi , J. L. Adams , A. J. Reagan , C. M. Danforth

We aim to construct a class of learning algorithms that are of practical value to applied researchers in fields such as biostatistics, epidemiology and econometrics, where the need to learn from incompletely observed information is…

统计方法学 · 统计学 2021-02-09 Alicia Curth , Ahmed M. Alaa , Mihaela van der Schaar

This paper considers derivation of $f$-divergence inequalities via the approach of functional domination. Bounds on an $f$-divergence based on one or several other $f$-divergences are introduced, dealing with pairs of probability measures…

信息论 · 计算机科学 2016-10-31 Igal Sason , Sergio Verdú

The diversity across outputs generated by LLMs shapes perception of their quality and utility. High lexical diversity is often desirable, but there is no standard method to measure this property. Templated answer structures and ``canned''…

计算与语言 · 计算机科学 2026-02-19 Chantal Shaib , Venkata S. Govindarajan , Joe Barrow , Jiuding Sun , Alexa F. Siu , Byron C. Wallace , Ani Nenkova

A method providing optimal estimate of probability density functions (PDFs) from time series is proposed. It allows almost arbitrary resolution PDFs when applied to either, sampled analytic functions or digitized data from experiments. When…

数据分析、统计与概率 · 物理学 2007-05-30 R. Labbé

This paper proposes a novel statistical approach to intelligent document retrieval. It seeks to offer a more structured and extensible mathematical approach to the term generalization done in the popular Latent Semantic Analysis (LSA)…

信息检索 · 计算机科学 2011-11-30 Scott Hand

We analyse correspondence of a text to a simple probabilistic model. The model assumes that the words are selected independently from an infinite dictionary. The probability distribution correspond to the Zipf---Mandelbrot law. We count…