中文
相关论文

相关论文: Discriminating Between Similar Nordic Languages

200 篇论文

Identifying closely related languages at sentence level is difficult, in particular because it is often impossible to assign a sentence to a single language. In this paper, we focus on multi-label sentence-level Scandinavian language…

We classify twenty-one Indo-European languages starting from written text. We use neural networks in order to define a distance among different languages, construct a dendrogram and analyze the ultrametric structure that emerges. Four or…

无序系统与神经网络 · 物理学 2015-07-16 Angelo Mariano , Giorgio Parisi , Saverio Pascazio

Finnish is a language with multiple dialects that not only differ from each other in terms of accent (pronunciation) but also in terms of morphological forms and lexical choice. We present the first approach to automatically detect the…

计算与语言 · 计算机科学 2021-11-09 Mika Hämäläinen , Khalid Alnajjar , Niko Partanen , Jack Rueter

In this paper, we present several baselines for automatic speech recognition (ASR) models for the two official written languages in Norway: Bokm{\aa}l and Nynorsk. We compare the performance of models of varying sizes and pre-training…

计算与语言 · 计算机科学 2023-07-06 Javier de la Rosa , Rolv-Arild Braaten , Per Egil Kummervold , Freddy Wetjen , Svein Arne Brygfjeld

Agglutinative languages such as Turkish, Finnish and Hungarian require morphological disambiguation before further processing due to the complex morphology of words. A morphological disambiguator is used to select the correct morphological…

计算与语言 · 计算机科学 2017-02-14 Eray Yildiz , Caglar Tirkaz , H. Bahadir Sahin , Mustafa Tolga Eren , Ozan Sonmez

Background: Clinical natural language processing (NLP) refers to the use of computational methods for extracting, processing, and analyzing unstructured clinical text data, and holds a huge potential to transform healthcare in various…

In text processing, deep neural networks mostly use word embeddings as an input. Embeddings have to ensure that relations between words are reflected through distances in a high-dimensional numeric space. To compare the quality of different…

计算与语言 · 计算机科学 2022-06-01 Matej Ulčar , Kristiina Vaik , Jessica Lindström , Milda Dailidėnaitė , Marko Robnik-Šikonja

Closed-set spoken language identification is the task of recognizing the language being spoken in a recorded audio clip from a set of known languages. In this study, a language identification system was built and trained to distinguish…

计算与语言 · 计算机科学 2022-05-20 Benjamin Kepecs , Homayoon Beigi

We investigate strategies for adapting small, efficient language models to Faroese, a low-resource North Germanic language. Starting from English-pretrained models, we apply continued pre-training on related Scandinavian languages --…

计算与语言 · 计算机科学 2026-03-27 Jenny Kunz , Iben Nyholm Debess , Annika Simonsen

Pre-training Large Language Models (LLMs) require massive amounts of text data, and the performance of the LLMs typically correlates with the scale and quality of the datasets. This means that it may be challenging to build LLMs for smaller…

Identification of the languages written using cuneiform symbols is a difficult task due to the lack of resources and the problem of tokenization. The Cuneiform Language Identification task in VarDial 2019 addresses the problem of…

计算与语言 · 计算机科学 2020-09-24 Ehsan Doostmohammadi , Minoo Nassajian

Norwegian, spoken by approximately five million people, remains underrepresented in many of the most significant breakthroughs in Natural Language Processing (NLP). To address this gap, the NorLLM team at NorwAI has developed a family of…

计算与语言 · 计算机科学 2026-01-09 Jon Atle Gulla , Peng Liu , Lemei Zhang

This paper introduces NorEval, a new and comprehensive evaluation suite for large-scale standardized benchmarking of Norwegian generative language models (LMs). NorEval consists of 24 high-quality human-created datasets -- of which five are…

This paper proposes the first multilingual (French, English and Arabic) and multicultural (Indo-European languages vs. less culturally close languages) irony detection system. We employ both feature-based models and neural architectures…

计算与语言 · 计算机科学 2020-02-07 Bilal Ghanem , Jihen Karoui , Farah Benamara , Paolo Rosso , Véronique Moriceau

This paper provides language identification models for low- and under-resourced languages in the Pacific region with a focus on previously unavailable Austronesian languages. Accurate language identification is an important part of…

计算与语言 · 计算机科学 2022-06-10 Jonathan Dunn , Wikke Nijhof

What is here called controlled natural language (CNL) has traditionally been given many different names. Especially during the last four decades, a wide variety of such languages have been designed. They are applied to improve communication…

计算与语言 · 计算机科学 2015-07-08 Tobias Kuhn

In language identification, a common first step in natural language processing, we want to automatically determine the language of some input text. Monolingual language identification assumes that the given document is written in one…

计算与语言 · 计算机科学 2017-08-01 Tom Kocmi , Ondřej Bojar

In historical linguistics, the affiliation of languages to a common language family is traditionally carried out using a complex workflow that relies on manually comparing individual languages. Large-scale standardized collections of…

计算与语言 · 计算机科学 2025-12-09 Frederic Blum , Steffen Herbold , Johann-Mattis List

Language identification is used as the first step in many data collection and crawling efforts because it allows us to sort online text into language-specific buckets. However, many modern languages, such as Konkani, Kashmiri, Punjabi etc.,…

计算与语言 · 计算机科学 2024-06-27 Milind Agarwal , Joshua Otten , Antonios Anastasopoulos

Language Identification (LI) is an important first step in several speech processing systems. With a growing number of voice-based assistants, speech LI has emerged as a widely researched field. To approach the problem of identifying…

计算与语言 · 计算机科学 2019-10-11 Sarthak , Shikhar Shukla , Govind Mittal
‹ 上一页 1 2 3 10 下一页 ›