中文
相关论文

相关论文: Euskarazko lehen C1 ebaluatzaile automatikoa

200 篇论文

This paper introduces the first publicly available dataset for Automatic Essay Scoring (AES) and feedback generation in Basque, targeting the CEFR C1 proficiency level. The dataset comprises 3,200 essays from HABE, each annotated by expert…

计算与语言 · 计算机科学 2026-03-24 Ekhi Azurmendi , Xabier Arregi , Oier Lopez de Lacalle

Studies on evaluation metrics and LLM-as-a-Judge models for automatic text summarization have largely been focused on English, limiting our understanding of their effectiveness in other languages. Through our new dataset BASSE (BAsque and…

计算与语言 · 计算机科学 2025-04-15 Jeremy Barnes , Naiara Perez , Alba Bonet-Jover , Begoña Altuna

Reflective thinking is a key competency in education, but assessing reflective writing remains a time-consuming and subjective task for education experts. While automated reflective analysis has been explored in several languages, Hungarian…

计算与语言 · 计算机科学 2026-05-05 Zsolt Csibi , Mónika Sándor , Mónika Serfőző , Kinga Gyöngy , Kristian Fenech

The widespread availability of Question Answering (QA) datasets in English has greatly facilitated the advancement of the Natural Language Processing (NLP) field. However, the scarcity of such resources for minority languages, such as…

计算与语言 · 计算机科学 2024-06-05 Aitor García-Pablos , Naiara Perez , Montse Cuadros , Jaione Bengoetxea

Readability assessment is the task of determining how difficult or easy a text is or which level/grade it has. Traditionally, language dependent readability formula have been used, but these formulae take few text characteristics into…

计算与语言 · 计算机科学 2021-09-13 Kepa Bengoetxea , Itziar Gonzalez-Dios

In this paper we describe the Senseval 2 Basque lexical-sample task. The task comprised 40 words (15 nouns, 15 verbs and 10 adjectives) selected from Euskal Hiztegia, the main Basque dictionary. Most examples were taken from the Egunkaria…

计算与语言 · 计算机科学 2007-05-23 Eneko Agirre , Elena Garcia , Mikel Lersundi , David Martinez , Eli Pociello

Large Language Models (LLMs) have demonstrated significant potential in various engineering tasks, including software development, digital logic generation, and companion document maintenance. However, their ability to perform board-level…

硬件体系结构 · 计算机科学 2026-03-20 Weibo Qiu , Yinhao Xiao , Runyu Pan

Evaluating statement autoformalization, translating natural language mathematics into formal languages like Lean 4, remains a significant challenge, with few metrics, datasets, and standards to robustly measure progress. In this work, we…

计算与语言 · 计算机科学 2025-10-30 Auguste Poiroux , Gail Weiss , Viktor Kunčak , Antoine Bosselut

In grammatical error correction (GEC), automatic evaluation is an important factor for research and development of GEC systems. Previous studies on automatic evaluation have demonstrated that quality estimation models built from datasets…

Automated content analysis increasingly supports communication research, yet scaling manual coding into computational pipelines raises concerns about measurement reliability and validity. We introduce a Hierarchical Error Correction (HEC)…

计算与语言 · 计算机科学 2025-10-27 Zhilong Zhao , Yindi Liu

Language models (LMs) are capable of acquiring elements of human-like syntactic knowledge. Targeted syntactic evaluation tests have been employed to measure how well they form generalizations about syntactic phenomena in high-resource…

计算与语言 · 计算机科学 2024-12-13 Daria Kryvosheieva , Roger Levy

Cross-lingual transfer-learning is widely used in Event Extraction for low-resource languages and involves a Multilingual Language Model that is trained in a source language and applied to the target language. This paper studies whether the…

计算与语言 · 计算机科学 2024-04-10 Mikel Zubillaga , Oscar Sainz , Ainara Estarrona , Oier Lopez de Lacalle , Eneko Agirre

As large language models continue to develop, the feasibility and significance of text-based symbolic music tasks have become increasingly prominent. While symbolic music has been widely used in generation tasks, LLM capabilities in…

声音 · 计算机科学 2025-09-30 Jiahao Zhao , Yunjia Li , Wei Li , Kazuyoshi Yoshii

In this paper we present the ADAPT system built for the Basque to English Low Resource MT Evaluation Campaign. Basque is a low-resourced, morphologically-rich language. This poses a challenge for Neural Machine Translation models which…

计算与语言 · 计算机科学 2018-11-15 Alberto Poncelas , Andy Way , Kepa Sarasola

This paper presents CRACQ, a multi-dimensional evaluation framework tailored to evaluate documents across f i v e specific traits: Coherence, Rigor, Appropriateness, Completeness, and Quality. Building on insights from traitbased Automated…

计算与语言 · 计算机科学 2025-10-06 Ishak Soltani , Francisco Belo , Bernardo Tavares

Reliable evaluation is essential for developing and deploying large language models, yet in practice it often requires substantial manual effort: practitioners must identify appropriate benchmarks, reproduce heterogeneous evaluation…

Patients with low health literacy usually have difficulty understanding medical jargon and the complex structure of professional medical language. Although some studies are proposed to automatically translate expert language into…

计算与语言 · 计算机科学 2024-02-09 Junyu Luo , Zifei Zheng , Hanzhong Ye , Muchao Ye , Yaqing Wang , Quanzeng You , Cao Xiao , Fenglong Ma

The availability of LLM benchmarks for the Estonian language is limited, and a comprehensive evaluation comparing the performance of different LLMs on Estonian tasks has yet to be conducted. We introduce a new benchmark for evaluating LLMs…

计算与语言 · 计算机科学 2026-02-20 Helena Grete Lillepalu , Tanel Alumäe

Word-level quality estimation (WQE) aims to automatically identify fine-grained error spans in machine-translated outputs and has found many uses, including assisting translators during post-editing. Modern WQE techniques are often…

计算与语言 · 计算机科学 2025-11-18 Gabriele Sarti , Vilém Zouhar , Malvina Nissim , Arianna Bisazza

Current evaluation frameworks for foundation models rely heavily on static, manually curated benchmarks, limiting their ability to capture the full breadth of model capabilities. This paper introduces Active learning for Capability…

‹ 上一页 1 2 3 10 下一页 ›