English
Related papers

Related papers: Conditional Equivalence Testing: an alternative re…

200 papers

We reassess a recent study (Hassan et al., 2018) that claimed that machine translation (MT) has reached human parity for the translation of news from Chinese into English, using pairwise ranking and considering three variables that were not…

Computation and Language · Computer Science 2018-08-31 Antonio Toral , Sheila Castilho , Ke Hu , Andy Way

We consider the problem of refuting equivalence of probabilistic programs, i.e., the problem of proving that two probabilistic programs induce different output distributions. We study this problem in the context of programs with…

Programming Languages · Computer Science 2025-01-14 Krishnendu Chatterjee , Ehsan Kafshdar Goharshady , Petr Novotný , Đorđe Žikelić

Randomized controlled trials often enroll participants whose characteristics differ from those of a target population, which can limit the generalizability of the estimated treatment effects when effect modifiers differ across populations.…

Methodology · Statistics 2026-05-15 Lan Wen , Issa J. Dahabreh , Yu-Han Chiu

The COMET metric has blazed a trail in the machine translation community, given its strong correlation with human judgements of translation quality. Its success stems from being a modified pre-trained multilingual model finetuned for…

Computation and Language · Computer Science 2024-10-01 Vilém Zouhar , Pinzhen Chen , Tsz Kin Lam , Nikita Moghe , Barry Haddow

Although there are many ideas for the formulations of statistical hypothesis testing, we consider that the likelihood ratio test is the most reasonable and orthodox. However, it is not handy, and thus, it is not usual in elementary books.…

Statistics Theory · Mathematics 2014-01-15 Shiro Ishikawa

Combinatorial Testing (CT) tools are essential to test properly a wide range of systems (train systems, Graphical User Interfaces (GUIs), autonomous driving systems, etc). While there is an active research community working on developing CT…

Software Engineering · Computer Science 2023-01-25 Carlos Ansotegui , Eduard Torres

When permutation methods are used in practice, often a limited number of random permutations are used to decrease the computational burden. However, most theoretical literature assumes that the whole permutation group is used, and methods…

Statistics Theory · Mathematics 2018-08-20 Jesse Hemerik , Jelle Goeman

This article concerns testing for equality of distribution between groups. We focus on screening variables with shared distributional features such as common support, modes and patterns of skewness. We propose a Bayesian testing method…

Methodology · Statistics 2016-02-19 Eric F. Lock , David B. Dunson

In this paper, we propose a simple and easy-to-implement Bayesian hypothesis test for the presence of an association, described by Kendall's \tau coefficient, between two variables measured on at least an ordinal scale. Owing to the absence…

Methodology · Statistics 2022-09-09 Shen Zhang , Keying Ye , Min Wang

In this article, we survey some controversial problems concerning the idea of erasing Which Way information proposed in recent years. A statistical examination of these proposals suggests that whenever the Bayesian rule is taken into…

Quantum Physics · Physics 2007-08-20 Mohammad Bahrami , Afshin Shafiee

Gender bias in machine translation (MT) systems has been extensively documented, but bias in automatic quality estimation (QE) metrics remains comparatively underexplored. Existing studies suggest that QE metrics can also exhibit gender…

Computation and Language · Computer Science 2025-10-09 Giorgos Filandrianos , Orfeas Menis Mastromichalakis , Wafaa Mohammed , Giuseppe Attanasio , Chrysoula Zerva

Tech companies (e.g., Google or Facebook) often use randomized online experiments and/or A/B testing primarily based on the average treatment effects to compare their new product with an old one. However, it is also critically important to…

Methodology · Statistics 2021-11-09 Chengchun Shi , Shikai Luo , Hongtu Zhu , Rui Song

High-dimensional tests are applied to find relevant sets of variables and relevant models. If variables are selected by analyzing the sums of products matrices and a corresponding mean-value test is performed, there is the danger that the…

Methodology · Statistics 2012-02-10 Juergen Laeuter , Maciej Rosolowski , Ekkehard Glimm

The high-quality translation results produced by machine translation (MT) systems still pose a huge challenge for automatic evaluation. Current MT evaluation pays the same attention to each sentence component, while the questions of…

Computation and Language · Computer Science 2021-08-02 Runzhe Zhan , Xuebo Liu , Derek F. Wong , Lidia S. Chao

Automatic evaluation in grammatical error correction (GEC) is crucial for selecting the best-performing systems. Currently, reference-based metrics are a popular choice, which basically measure the similarity between hypothesis and…

Computation and Language · Computer Science 2026-02-06 Takumi Goto , Yusuke Sakai , Taro Watanabe

This paper addresses an important problem in Example-Based Machine Translation (EBMT), namely how to measure similarity between a sentence fragment and a set of stored examples. A new method is proposed that measures similarity according to…

cmp-lg · Computer Science 2008-02-03 Lambros Cranias , Harris Papageorgiou , Stelios Piperidis

Publication bias is a major concern in conducting systematic reviews and meta-analyses. Various sensitivity analysis or bias-correction methods have been developed based on selection models and they have some advantages over the widely used…

Methodology · Statistics 2021-09-28 Ao Huang , Kosuke Morikawa , Tim Friede , Satoshi Hattori

This paper presents a Bayesian framework for assessing the adequacy of a model without the necessity of explicitly enumerating a specific alternate model. A test statistic is developed for tracking the performance of the model across…

Artificial Intelligence · Computer Science 2013-03-25 Kathryn Blackmond Laskey

We consider a quantum system that is being continuously monitored, giving rise to a measurement signal. From such a stream of data, information needs to be inferred about the underlying system's dynamics. Here we focus on hypothesis testing…

Quantum Physics · Physics 2024-03-27 Giulio Gasbarri , Matias Bilkis , Elisabet Roda-Salichs , John Calsamiglia

Style is an integral part of natural language. However, evaluation methods for style measures are rare, often task-specific and usually do not control for content. We propose the modular, fine-grained and content-controlled similarity-based…

Computation and Language · Computer Science 2021-09-13 Anna Wegmann , Dong Nguyen