中文
相关论文

相关论文: Detecting Structured Language Alternations in Hist…

200 篇论文

With the availability of virtually infinite number text documents in digital format, automatic comparison of textual data is essential for extracting meaningful insights that are difficult to identify manually. Many existing tools,…

信息检索 · 计算机科学 2025-03-25 Akhil Joshi , Sai Teja Erukude , Lior Shamir

This paper reports on the preliminary phase of our ongoing research towards developing an intelligent tutoring environment for Turkish grammar. One of the components of this environment is a corpus search tool which, among other aspects of…

cmp-lg · 计算机科学 2016-08-31 H. Altay Guvenir , Kemal Oflazer

The task of determining the similarity of text documents has received considerable attention in many areas such as Information Retrieval, Text Mining, Natural Language Processing (NLP) and Computational Linguistics. Transferring data to…

信息检索 · 计算机科学 2022-11-23 Bakhyt Bakiyev

Informal transliteration from other languages to English is prevalent in social media threads, instant messaging, and discussion forums. Without identifying the language of such transliterated text, users who do not speak that language…

音频与语音处理 · 电气工程与系统科学 2020-04-29 Sourav Sen

The dissemination of Large Language Models (LLMs), trained at scale, and endowed with powerful text-generating abilities, has made it easier for all to produce harmful, toxic, faked or forged content. In response, various proposals have…

计算与语言 · 计算机科学 2025-06-12 Matthieu Dubois , François Yvon , Pablo Piantanida

Handwritten Arabic manuscripts preserve the Arab world's intellectual and cultural heritage, and writer identification supports provenance, authenticity verification, and historical analysis. Using the Muharaf dataset of historical Arabic…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Hamza A. Abushahla , Ariel Justine N. Panopio , Layth Al-Khairulla , Mohamed I. AlHajri

In today's world, we follow news which is distributed globally. Significant events are reported by different sources and in different languages. In this work, we address the problem of tracking of events in a large multilingual stream.…

信息检索 · 计算机科学 2015-12-23 Jan Rupnik , Andrej Muhic , Gregor Leban , Primoz Skraba , Blaz Fortuna , Marko Grobelnik

Accessibility to historical documents is mostly limited to scholars. This is due to the language barrier inherent in human language and the linguistic properties of these documents. Given a historical document, modernization aims to…

计算与语言 · 计算机科学 2020-03-05 Miguel Domingo , Francisco Casacuberta

Any large organisation, be it public or private, monitors the media for information to keep abreast of developments in their field of interest, and usually also to become aware of positive or negative opinions expressed towards them. At…

Most state-of-the-art systems today produce morphological analysis based only on orthographic patterns. In contrast, we propose a model for unsupervised morphological analysis that integrates orthographic and semantic views of words. We…

计算与语言 · 计算机科学 2015-03-10 Karthik Narasimhan , Regina Barzilay , Tommi Jaakkola

This paper elaborates on the notion of uncertainty in the context of annotation in large text corpora, specifically focusing on (but not limited to) historical languages. Such uncertainty might be due to inherent properties of the language,…

计算与语言 · 计算机科学 2021-05-31 Marie-Luis Merten , Marcel Wever , Michaela Geierhos , Doris Tophinke , Eyke Hüllermeier

We present a new method to detect anomalies in texts (in general: in sequences of any data), using language models, in a totally unsupervised manner. The method considers probabilities (likelihoods) generated by a language model, but…

计算与语言 · 计算机科学 2024-09-06 Filip Graliński , Ryszard Staruch , Krzysztof Jurkiewicz

This paper proposes the first multilingual (French, English and Arabic) and multicultural (Indo-European languages vs. less culturally close languages) irony detection system. We employ both feature-based models and neural architectures…

计算与语言 · 计算机科学 2020-02-07 Bilal Ghanem , Jihen Karoui , Farah Benamara , Paolo Rosso , Véronique Moriceau

We introduce a new pretraining approach geared for multi-document language modeling, incorporating two key ideas into the masked language modeling self-supervised objective. First, instead of considering documents in isolation, we pretrain…

计算与语言 · 计算机科学 2021-09-06 Avi Caciularu , Arman Cohan , Iz Beltagy , Matthew E. Peters , Arie Cattan , Ido Dagan

We use the persistent homology method of topological data analysis and dimensional analysis techniques to study data of syntactic structures of world languages. We analyze relations between syntactic parameters in terms of dimensionality,…

计算与语言 · 计算机科学 2019-03-14 Alexander Port , Taelin Karidi , Matilde Marcolli

Deep neural networks have been employed for various spoken language recognition tasks, including tasks that are multilingual by definition such as spoken language identification. In this paper, we present a neural model for Slavic language…

计算与语言 · 计算机科学 2020-10-26 Badr M. Abdullah , Jacek Kudera , Tania Avgustinova , Bernd Möbius , Dietrich Klakow

Technical documents contain a fair amount of unnatural language, such as tables, formulas, pseudo-codes, etc. Unnatural language can be an important factor of confusing existing NLP tools. This paper presents an effective method of…

信息检索 · 计算机科学 2017-03-20 Myungha Jang , Jinho D. Choi , James Allan

The ambition of a character recognition system is to transform a text document typed on paper into a digital format that can be manipulated by word processor software Unlike other languages, Arabic has unique features, while other language…

计算与语言 · 计算机科学 2010-06-15 A. A Zaidan , B. B Zaidan , Hamid. A. Jalab , Hamdan. O. Alanazi , Rami Alnaqeib

This paper introduces foundational resources and models for natural language processing (NLP) of historical Turkish, a domain that has remained underexplored in computational linguistics. We present the first named entity recognition (NER)…

The debate surrounding language identification has gained renewed attention in recent years, especially with the rapid evolution of AI-powered language models. However, the non-AI-based approaches to language identification have been…

计算与语言 · 计算机科学 2025-07-24 Paul-Andrei Pogăcean , Sanda-Maria Avram