English
Related papers

Related papers: A Comparison of Two Smoothing Methods for Word Big…

200 papers

The detection of machine-generated text, especially from large language models (LLMs), is crucial in preventing serious social problems resulting from their misuse. Some methods train dedicated detectors on specific datasets but fall short…

Machine Learning · Computer Science 2024-06-05 Yibo Miao , Hongcheng Gao , Hao Zhang , Zhijie Deng

Widely used methods for analyzing missing data can be biased in small samples. To understand these biases, we evaluate in detail the situation where a small univariate normal sample, with values missing at random, is analyzed using either…

Statistics Theory · Mathematics 2017-03-27 Paul T. von Hippel

This paper addresses the weak instruments problem in linear instrumental variable models from a Bayesian perspective. The new approach has two components. First, a novel predictor-dependent shrinkage prior is developed for the many…

Methodology · Statistics 2014-08-05 P. Richard Hahn , Hedibert Lopes

Smoothing algorithms for state-space models, i.e., fixed-interval smoothing, fixed-lag smoothing, and two-filter formula for smoothing, are examined using real examples. For linear and Gaussian state-space models, it is observed that…

Computation · Statistics 2023-07-10 G. Kitagawa

We examine a new form of smooth approximation to the zero one loss in which learning is performed using a reformulation of the widely used logistic function. Our approach is based on using the posterior mean of a novel generalized…

Computer Vision and Pattern Recognition · Computer Science 2015-11-19 Md Kamrul Hasan , Christopher J. Pal

A recent line of work has shown that end-to-end optimization of Bayesian filters can be used to learn state estimators for systems whose underlying models are difficult to hand-design or tune, while retaining the core advantages of…

Robotics · Computer Science 2021-08-24 Brent Yi , Michelle A. Lee , Alina Kloss , Roberto Martín-Martín , Jeannette Bohg

Single-channel deep speech enhancement approaches often estimate a single multiplicative mask to extract clean speech without a measure of its accuracy. Instead, in this work, we propose to quantify the uncertainty associated with clean…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-16 Huajian Fang , Timo Gerkmann

Many text corpora exhibit socially problematic biases, which can be propagated or amplified in the models trained on such data. For example, doctor cooccurs more frequently with male pronouns than female pronouns. In this study we (i)…

Computation and Language · Computer Science 2019-04-08 Shikha Bordia , Samuel R. Bowman

Measuring sentence semantic similarity using pre-trained language models such as BERT generally yields unsatisfactory zero-shot performance, and one main reason is ineffective token aggregation methods such as mean pooling. In this paper,…

Computation and Language · Computer Science 2020-10-23 M. Li , H. Bai , L. Tan , K. Xiong , M. Li , J. Lin

In recent years, word embeddings have been widely used to measure biases in texts. Even if they have proven to be effective in detecting a wide variety of biases, metrics based on word embeddings lack transparency and interpretability. We…

Computation and Language · Computer Science 2023-07-19 Francisco Valentini , Germán Rosati , Damián Blasi , Diego Fernandez Slezak , Edgar Altszyler

We present a sparse estimation and dictionary learning framework for compressed fiber sensing based on a probabilistic hierarchical sparse model. To handle severe dictionary coherence, selective shrinkage is achieved using a Weibull prior,…

Machine Learning · Statistics 2016-10-24 Christian Weiss , Abdelhak M. Zoubir

We present a simple method to obtain optimal posterior distributions and improve the quality of Bayesian inference with reduced human and computational effort. Bayes' Theorem is reformulated in the language of statistical mechanics, wherein…

Methodology · Statistics 2026-04-28 Alfred C. K. Farris

Recent approaches to cross-lingual word embedding have generally been based on linear transformations between the sets of embedding vectors in the two languages. In this paper, we propose an approach that instead expresses the two…

Computation and Language · Computer Science 2019-10-08 Chunting Zhou , Xuezhe Ma , Di Wang , Graham Neubig

This paper, which is Part 1 of a two-part paper series, considers a simulation-based inference with learned summary statistics, in which such a learned summary statistic serves as an empirical-likelihood with ameliorative effects in the…

Machine Learning · Statistics 2026-02-02 Getachew K. Befekadu

The category of figurative language contains many varieties, some of which are non-compositional in nature. This type of phrase or multi-word expression (MWE) includes idioms, which represent a single meaning that does not consist of the…

Computation and Language · Computer Science 2026-03-25 Blake Matheny , Phuong Minh Nguyen , Minh Le Nguyen

We compare the accuracy, precision and reliability of different methods for estimating key system parameters for two-level systems subject to Hamiltonian evolution and decoherence. It is demonstrated that the use of Bayesian modelling and…

Quantum Physics · Physics 2019-10-15 Sophie Schirmer , Frank Langbein

Computing sample means on Riemannian manifolds is typically computationally costly as exemplified by computation of the Fr\'echet mean which often requires finding minimizing geodesics to each data point for each step of an iterative…

Methodology · Statistics 2022-05-25 Mathias Højgaard Jensen , Stefan Sommer

In automatic speech recognition, many studies have shown performance improvements using language models (LMs). Recent studies have tried to use bidirectional LMs (biLMs) instead of conventional unidirectional LMs (uniLMs) for rescoring the…

Computation and Language · Computer Science 2019-05-17 Joongbo Shin , Yoonhyung Lee , Kyomin Jung

The relationship between written and spoken words is convoluted in languages with a deep orthography such as English and therefore it is difficult to devise explicit rules for generating the pronunciations for unseen words. Pronunciation by…

Computation and Language · Computer Science 2011-09-22 Janne V. Kujala , Aleksi Keurulainen

Many statistical models can be simulated forwards but have intractable likelihoods. Approximate Bayesian Computation (ABC) methods are used to infer properties of these models from data. Traditionally these methods approximate the posterior…

Machine Learning · Statistics 2018-04-03 George Papamakarios , Iain Murray