中文
相关论文

相关论文: HELFI: a Hebrew-Greek-Finnish Parallel Bible Corpu…

200 篇论文

This paper presents the first Swedish evaluation benchmark for textual semantic similarity. The benchmark is compiled by simply running the English STS-B dataset through the Google machine translation API. This paper discusses potential…

计算与语言 · 计算机科学 2020-12-01 Tim Isbister , Magnus Sahlgren

We present PhiloBERTA, a cross-lingual transformer model that measures semantic relationships between ancient Greek and Latin lexicons. Through analysis of selected term pairs from classical texts, we use contextual embeddings and angular…

计算与语言 · 计算机科学 2025-08-26 Rumi Allbert , Makai L. Allbert

In today's modern wide-field galaxy surveys, there is the necessity for parametric surface brightness decomposition codes characterised by accuracy, small degree of user intervention, and high degree of parallelisation. We try to address…

星系天体物理 · 物理学 2023-03-03 Luca Tortorelli , Amata Mercurio

Entity resolution is a widely studied problem with several proposals to match records across relations. Matching textual content is a widespread task in many applications, such as question answering and search. While recent methods achieve…

数据库 · 计算机科学 2021-12-17 Naser Ahmadi , Hansjorg Sand , Paolo Papotti

In recent years, a flurry of morphological datasets had emerged, most notably UniMorph, a multi-lingual repository of inflection tables. However, the flat structure of the current morphological annotation schema makes the treatment of some…

计算与语言 · 计算机科学 2022-03-22 David Guriel , Omer Goldman , Reut Tsarfaty

Multilingual alignment of sentence representations has mostly required bitexts to bridge the gap between languages. We investigate whether visual information can bridge this gap instead. Image caption datasets are very easy to create…

计算与语言 · 计算机科学 2025-05-21 Nathaniel Krasner , Nicholas Lanuzo , Antonios Anastasopoulos

We introduce FIN-bench-v2, a unified benchmark suite for evaluating large language models in Finnish. FIN-bench-v2 consolidates Finnish versions of widely used benchmarks together with an updated and expanded version of the original…

计算与语言 · 计算机科学 2025-12-16 Joona Kytöniemi , Jousia Piha , Akseli Reunamo , Fedor Vitiugin , Farrokh Mehryary , Sampo Pyysalo

We present preliminary results about Legistix, a tool we are developing to automatically consolidate the French and European law. Legistix is based both on regular expressions used in several compound grammars, similar to the successive…

计算与语言 · 计算机科学 2023-01-18 Georges-André Silber

Judeo-Arabic refers to Arabic variants historically spoken by Jewish communities across the Arab world, primarily during the Middle Ages. Unlike standard Arabic, it is written in Hebrew script by Jewish writers and for Jewish audiences.…

计算与语言 · 计算机科学 2026-01-30 Juan Moreno Gonzalez , Bashar Alhafni , Nizar Habash

Foundational Hebrew NLP tasks such as segmentation, tagging and parsing, have relied to date on various versions of the Hebrew Treebank (HTB, Sima'an et al. 2001). However, the data in HTB, a single-source newswire corpus, is now over 30…

计算与语言 · 计算机科学 2022-10-19 Amir Zeldes , Nick Howell , Noam Ordan , Yifat Ben Moshe

This paper introduces an updated and combined version of the bidirectional English-German EPIC-UdS (spoken) and EuroParl-UdS (written) corpora containing original European Parliament speeches as well as their translations and…

计算与语言 · 计算机科学 2026-03-17 Maria Kunilovskaya , Christina Pollkläsener

In light of recent legal allegations brought by publishers, newspapers, and other creators of copyrighted corpora against large language model developers who use their copyrighted materials for training or fine-tuning purposes, we propose a…

计算与语言 · 计算机科学 2024-08-05 Devam Mondal , Carlo Lipizzi

Word alignment over parallel corpora has a wide variety of applications, including learning translation lexicons, cross-lingual transfer of language processing tools, and automatic evaluation or analysis of translation outputs. The great…

计算与语言 · 计算机科学 2021-08-13 Zi-Yi Dou , Graham Neubig

We study unsupervised multilingual alignment, the problem of finding word-to-word translations between multiple languages without using any parallel data. One popular strategy is to reduce multilingual alignment to the much simplified…

计算与语言 · 计算机科学 2020-07-30 Xin Lian , Kshitij Jain , Jakub Truszkowski , Pascal Poupart , Yaoliang Yu

Document alignment is necessary for the hierarchical mining (Ba\~n\'on et al., 2020; Morishita et al., 2022), which aligns documents across source and target languages within the same web domain. Several high precision sentence…

计算与语言 · 计算机科学 2025-10-20 Xiaotian Wang , Takehito Utsuro , Masaaki Nagata

This paper proposes a method for extracting translations of morphologically constructed terms from comparable corpora. The method is based on compositional translation and exploits translation equivalences at the morpheme-level, which…

计算与语言 · 计算机科学 2012-10-23 Estelle Delpech , Béatrice Daille , Emmanuel Morin , Claire Lemaire

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for mining such data from…

计算与语言 · 计算机科学 2015-11-20 Krzysztof Wołk , Emilia Rejmund , Krzysztof Marasek

This paper addresses the problem of providing a novel approach to sourcing significant training data for LLMs focused on science and engineering. In particular, a crucial challenge is sourcing parallel scientific codes in the ranges of…

软件工程 · 计算机科学 2025-05-06 Matthew T. Dearing , Yiheng Tao , Xingfu Wu , Zhiling Lan , Valerie Taylor

This paper presents a constraint-based morphological disambiguation approach that is applicable languages with complex morphology--specifically agglutinative languages with productive inflectional and derivational morphological phenomena.…

cmp-lg · 计算机科学 2008-02-03 Kemal Oflazer , Gokhan Tur

Machine translation has been a major motivation of development in natural language processing. Despite the burgeoning achievements in creating more efficient machine translation systems thanks to deep learning methods, parallel corpora have…

计算与语言 · 计算机科学 2020-10-06 Sina Ahmadi , Hossein Hassani , Daban Q. Jaff