中文
相关论文

相关论文: Discriminating Between Similar Nordic Languages

200 篇论文

Language is, as commonly theorized, largely arbitrary. Yet, systematic relationships between phonetics and semantics have been observed in many specific cases. To what degree could those systematic relationships manifest themselves in large…

计算与语言 · 计算机科学 2025-10-30 George Flint , Kaustubh Kislay

Lexical ambiguity, a challenging phenomenon in all natural languages, is particularly prevalent for languages with diacritics that tend to be omitted in writing, such as Arabic. Omitting diacritics leads to an increase in the number of…

计算与语言 · 计算机科学 2019-12-11 Sawsan Alqahtani , Hanan Aldarmaki , Mona Diab

Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may be more appropriate to consider it a data mining problem. For…

计算与语言 · 计算机科学 2025-03-11 Rasul Dent , Pedro Ortiz Suarez , Thibault Clérice , Benoît Sagot

In this paper we analyze features to classify human- and AI-generated text for English, French, German and Spanish and compare them across languages. We investigate two scenarios: (1) The detection of text generated by AI from scratch, and…

计算与语言 · 计算机科学 2024-01-31 Kristina Schaaff , Tim Schlippe , Lorenz Mindner

Hate speech detection is a challenging problem with most of the datasets available in only one language: English. In this paper, we conduct a large scale analysis of multilingual hate speech in 9 languages from 16 different sources. We…

社会与信息网络 · 计算机科学 2020-12-10 Sai Saketh Aluru , Binny Mathew , Punyajoy Saha , Animesh Mukherjee

KnowNER is a multilingual Named Entity Recognition (NER) system that leverages different degrees of external knowledge. A novel modular framework divides the knowledge into four categories according to the depth of knowledge they convey.…

计算与语言 · 计算机科学 2017-09-13 Dominic Seyler , Tatiana Dembelova , Luciano Del Corro , Johannes Hoffart , Gerhard Weikum

Current decoder-based pre-trained language models (PLMs) successfully demonstrate multilingual capabilities. However, it is unclear how these models handle multilingualism. We analyze the neuron-level internal behavior of multilingual…

计算与语言 · 计算机科学 2024-04-04 Takeshi Kojima , Itsuki Okimura , Yusuke Iwasawa , Hitomi Yanaka , Yutaka Matsuo

Deep neural networks have been employed for various spoken language recognition tasks, including tasks that are multilingual by definition such as spoken language identification. In this paper, we present a neural model for Slavic language…

计算与语言 · 计算机科学 2020-10-26 Badr M. Abdullah , Jacek Kudera , Tania Avgustinova , Bernd Möbius , Dietrich Klakow

Detecting fine-grained differences in content conveyed in different languages matters for cross-lingual NLP and multilingual corpora analysis, but it is a challenging machine learning problem since annotation is expensive and hard to scale.…

计算与语言 · 计算机科学 2020-10-09 Eleftheria Briakou , Marine Carpuat

In language recognition, the task of rejecting/differentiating closely spaced versus acoustically far spaced languages remains a major challenge. For confusable closely spaced languages, the system needs longer input test duration material…

声音 · 计算机科学 2016-09-22 Suwon Shon , Seongkyu Mun , John H. L. Hansen , Hanseok Ko

A broad goal in natural language processing (NLP) is to develop a system that has the capacity to process any natural language. Most systems, however, are developed using data from just one language such as English. The SIGMORPHON 2020…

Syllabification describes the task of dividing words into syllables. Due to many rules and exceptions, training an algorithm to perform syllabification with high accuracy remains a challenge. Throughout the last decades, different…

计算与语言 · 计算机科学 2026-05-29 Gus Lathouwers , Wieke Harmsen , Catia Cucchiarini , Helmer Strik

Both research and commercial machine translation have so far neglected the importance of properly handling the spelling, lexical and grammar divergences occurring among language varieties. Notable cases are standard national varieties such…

计算与语言 · 计算机科学 2018-11-06 Surafel M. Lakew , Aliia Erofeeva , Marcello Federico

We present the first open-set language identification experiments using one-class classification. We first highlight the shortcomings of traditional feature extraction methods and propose a hashing-based feature vectorization approach as a…

计算与语言 · 计算机科学 2017-07-18 Shervin Malmasi

Despite their high predictive accuracies, current machine learning systems often exhibit systematic biases stemming from annotation artifacts or insufficient support for certain classes in the dataset. Recent work proposes automatic methods…

计算与语言 · 计算机科学 2024-10-30 Rakesh R. Menon , Shashank Srivastava

By training on text in various languages, large language models (LLMs) typically possess multilingual support and demonstrate remarkable capabilities in solving tasks described in different languages. However, LLMs can exhibit linguistic…

计算与语言 · 计算机科学 2024-05-13 Guoliang Dong , Haoyu Wang , Jun Sun , Xinyu Wang

Arabic has diverse dialects, where one dialect can be substantially different from the others. In the NLP literature, some assumptions about these dialects are widely adopted (e.g., ``Arabic dialects can be grouped into distinguishable…

计算与语言 · 计算机科学 2025-05-29 Amr Keleg , Sharon Goldwater , Walid Magdy

When learning a new skill, you take advantage of your preexisting skills and knowledge. For instance, if you are a skilled violinist, you will likely have an easier time learning to play cello. Similarly, when learning a new language you…

计算与语言 · 计算机科学 2017-11-06 Johannes Bjerva

Language identification is an important Natural Language Processing task. It has been thoroughly researched in the literature. However, some issues are still open. This work addresses the identification of the related low-resource languages…

计算与语言 · 计算机科学 2022-03-10 Olha Dovbnia , Anna Wróblewska

Language identification has become a prerequisite for all kinds of automated text processing systems. In this paper, we present a rule-based language identifier tool for two closely related Indo-Aryan languages: Hindi and Magahi. This…

计算与语言 · 计算机科学 2018-04-17 Priya Rani , Atul Kr. Ojha , Girish Nath Jha