中文
相关论文

相关论文: Better Smatch = Better Parser? AMR evaluation is n…

200 篇论文

Many constructs that characterize language, like its complexity or emotionality, have a naturally continuous semantic structure; a public speech is not just "simple" or "complex," but exists on a continuum between extremes. Although large…

计算与语言 · 计算机科学 2025-09-23 Hauke Licht , Rupak Sarkar , Patrick Y. Wu , Pranav Goel , Niklas Stoehr , Elliott Ash , Alexander Miserlis Hoyle

Machine Translation (MT) and automatic MT evaluation have improved dramatically in recent years, enabling numerous novel applications. Automatic evaluation techniques have evolved from producing scalar quality scores to precisely locating…

计算与语言 · 计算机科学 2026-03-23 Stefano Perrella , Eric Morales Agostinho , Hugo Zaragoza

Identifying semantically equivalent sentences is important for many cross-lingual and mono-lingual NLP tasks. Current approaches to semantic equivalence take a loose, sentence-level approach to "equivalence," despite previous evidence that…

计算与语言 · 计算机科学 2022-10-07 Shira Wein , Zhuxin Wang , Nathan Schneider

Numerous studies have assessed the proficiency of AI systems, particularly large language models (LLMs), in facilitating everyday tasks such as email writing, question answering, and creative content generation. However, researchers face…

Automatic machine translation metrics typically rely on human translations to determine the quality of system translations. Common wisdom in the field dictates that the human references should be of very high quality. However, there are no…

计算与语言 · 计算机科学 2024-04-11 Vilém Zouhar , Ondřej Bojar

We present a transition-based AMR parser that directly generates AMR parses from plain text. We use Stack-LSTMs to represent our parser state and make decisions greedily. In our experiments, we show that our parser achieves very competitive…

计算与语言 · 计算机科学 2017-08-03 Miguel Ballesteros , Yaser Al-Onaizan

Normally, summary quality measures are compared with quality scores produced by human annotators. A higher correlation with human scores is considered to be a fair indicator of a better measure. We discuss observations that cast doubt on…

计算与语言 · 计算机科学 2021-01-01 Oleg Vasilyev , John Bohannon

Many studies have shown automatic speech processing (ASR) systems have unequal performance across speakergroups (SG's). However, the manner in which such studies arrive at this conclusion is inconsistent. To pave the wayfor more reliable…

计算与语言 · 计算机科学 2026-05-12 Felix Herron , Ange Richard , François Portet , Alexandre Allauzen , Solange Rossato

Many predicted structured objects (e.g., sequences, matchings, trees) are evaluated using the F-score, alignment error rate (AER), or other multivariate performance measures. Since inductively optimizing these measures using training data…

机器学习 · 统计学 2017-12-22 Hong Wang , Ashkan Rezaei , Brian D. Ziebart

Starting from the 1950s, Machine Translation (MT) was challenged by different scientific solutions, which included rule-based methods, example-based and statistical models (SMT), to hybrid models, and very recent years the neural models…

计算与语言 · 计算机科学 2025-08-07 Lifeng Han , Serge Gladkoff

Modern consumer products are full of interconnected electrical and electronic modules to fulfill direct and indirect needs. In an automated assembly line still, most of these interconnections are required to be done manually due to the…

信号处理 · 电气工程与系统科学 2025-10-29 Brian Skoglind , Travis Roberts , Sourabh Karmakar , Cameron Turner , Laine Mears

The use of machine learning (ML) models to assess and score textual data has become increasingly pervasive in an array of contexts including natural language processing, information retrieval, search and recommendation, and credibility…

计算与语言 · 计算机科学 2023-09-27 Marialena Bevilacqua , Kezia Oketch , Ruiyang Qin , Will Stamey , Xinyuan Zhang , Yi Gan , Kai Yang , Ahmed Abbasi

We present a review of high-performance automatic modulation recognition (AMR) models proposed in the literature to classify various Radio Frequency (RF) modulation schemes. We replicated these models and compared their performance in terms…

Improvements in text generation technologies such as machine translation have necessitated more costly and time-consuming human evaluation procedures to ensure an accurate signal. We investigate a simple way to reduce cost by reducing the…

计算与语言 · 计算机科学 2022-04-12 Belén Saldías , George Foster , Markus Freitag , Qijun Tan

Machine translation has made rapid advances in recent years. Millions of people are using it today in online translation systems and mobile applications in order to communicate across language barriers. The question naturally arises whether…

Word reordering is one of the most difficult aspects of statistical machine translation (SMT), and an important factor of its quality and efficiency. Despite the vast amount of research published to date, the interest of the community in…

计算与语言 · 计算机科学 2017-02-27 Arianna Bisazza , Marcello Federico

The Abstraction Reasoning Corpus (ARC) is a visual analogical reasoning test designed for humans and machines (Chollet, 2019). We compared human and large language model (LLM) performance on a new child-friendly set of ARC items. Results…

计算与语言 · 计算机科学 2024-05-14 Gustaw Opiełka , Hannes Rosenbusch , Veerle Vijverberg , Claire E. Stevenson

Translated texts bear several hallmarks distinct from texts originating in the language. Though individual translated texts are often fluent and preserve meaning, at a large scale, translated texts have statistical tendencies which…

计算与语言 · 计算机科学 2024-01-31 Shira Wein , Nathan Schneider

As AI becomes more integral in our lives, the need for transparency and responsibility grows. While natural language explanations (NLEs) are vital for clarifying the reasoning behind AI decisions, evaluating them through human judgments is…

计算与语言 · 计算机科学 2024-03-27 Fan Huang , Haewoon Kwak , Kunwoo Park , Jisun An

Inferring evaluation scores based on human judgments is invaluable compared to using current evaluation metrics which are not suitable for real-time applications e.g. post-editing. However, these judgments are much more expensive to collect…

计算与语言 · 计算机科学 2013-07-09 Ibrahim Sabek , Noha A. Yousri , Nagwa Elmakky , Mona Habib