English
Related papers

Related papers: HmBlogs: A big general Persian corpus

200 papers

We describe an Arabic-Hebrew parallel corpus of TED talks built upon WIT3, the Web inventory that repurposes the original content of the TED website in a way which is more convenient for MT researchers. The benchmark consists of about 2,000…

Computation and Language · Computer Science 2016-10-04 Mauro Cettolo

Digital text is increasing day by day on the internet. It is very challenging to classify a large and heterogeneous collection of data, which require improved information processing methods to organize text. To classify large size of…

Computation and Language · Computer Science 2021-07-08 Taimoor Ahmed Javed , Waseem Shahzad , Umair Arshad

We introduce VoxPopuli, a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages. It is the largest open data to date for unsupervised representation learning as well as semi-supervised learning.…

Computation and Language · Computer Science 2021-07-28 Changhan Wang , Morgane Rivière , Ann Lee , Anne Wu , Chaitanya Talnikar , Daniel Haziza , Mary Williamson , Juan Pino , Emmanuel Dupoux

This paper presents an extension to a very low-resource parallel corpus collected in an endangered language, Griko, making it useful for computational research. The corpus consists of 330 utterances (about 20 minutes of speech) which have…

Computation and Language · Computer Science 2018-07-30 Marcely Zanon Boito , Antonios Anastasopoulos , Marika Lekakou , Aline Villavicencio , Laurent Besacier

For any deep computational processing of language we need evidences, and one such set of evidences is corpus. This paper describes the development of a text-based corpus for the Bishnupriya Manipuri language. A Corpus is considered as a…

Computation and Language · Computer Science 2013-12-12 Nayan Jyoti Kalita , Navanath Saharia , Smriti Kumar Sinha

The use of Project Gutenberg (PG) as a text corpus has been extremely popular in statistical analysis of language for more than 25 years. However, in contrast to other major linguistic datasets of similar importance, no consensual full…

Computation and Language · Computer Science 2018-12-20 Martin Gerlach , Francesc Font-Clos

Online hate speech threatens online civility, particularly in low-resource and multilingual environments. Counter-narratives offer a promising solution by promoting constructive responses to hate speech. However, automatic counter-narrative…

Social and Information Networks · Computer Science 2026-03-31 Zahra Safdari Fesaghandis , Suman Kalyan Maity

This paper aims to make up for the lack of documented baselines for Hungarian language modeling. Various approaches are evaluated on three publicly available Hungarian corpora. Perplexity values comparable to models of similar-sized English…

Computation and Language · Computer Science 2017-01-30 Dávid Márk Nemeskey

Semantic knowledge can be a great asset to natural language processing systems, but it is usually hand-coded for each application. Although some semantic information is available in general-purpose knowledge bases such as WordNet and Cyc,…

cmp-lg · Computer Science 2008-02-03 Ellen Riloff , Jessica Shepherd

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

Parallel text is required for building high-quality machine translation (MT) systems, as well as for other multilingual NLP applications. For many South Asian languages, such data is in short supply. In this paper, we described a new…

Computation and Language · Computer Science 2020-01-28 Barry Haddow , Faheem Kirefu

We analyze three critical components of word embedding training: the model, the corpus, and the training parameters. We systematize existing neural-network-based word embedding algorithms and compare them using the same corpus. We evaluate…

Computation and Language · Computer Science 2015-07-21 Siwei Lai , Kang Liu , Liheng Xu , Jun Zhao

Semantic sentence embeddings are usually supervisedly built minimizing distances between pairs of embeddings of sentences labelled as semantically similar by annotators. Since big labelled datasets are rare, in particular for non-English…

Computation and Language · Computer Science 2021-10-06 Marco Di Giovanni , Marco Brambilla

This paper reports on the preliminary phase of our ongoing research towards developing an intelligent tutoring environment for Turkish grammar. One of the components of this environment is a corpus search tool which, among other aspects of…

cmp-lg · Computer Science 2016-08-31 H. Altay Guvenir , Kemal Oflazer

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source…

Computation and Language · Computer Science 2021-06-16 Elizabeth Salesky , Matthew Wiesner , Jacob Bremerman , Roldano Cattoni , Matteo Negri , Marco Turchi , Douglas W. Oard , Matt Post

Nowadays, dialogue systems are used in many fields of industry and research. There are successful instances of these systems, such as Apple Siri, Google Assistant, and IBM Watson. Task-oriented dialogue system is a category of these, that…

Computation and Language · Computer Science 2024-01-02 Keyvan Mahmoudi , Heshaam Faili

Incorporating information from other languages can improve the results of tasks in low-resource languages. A powerful method of building functional natural language processing systems for low-resource languages is to combine multilingual…

Computation and Language · Computer Science 2022-05-19 Heydar Soudani , Mohammad Hassan Mojab , Hamid Beigy

Keyphrases are a very short summary of an input text and provide the main subjects discussed in the text. Keyphrase extraction is a useful upstream task and can be used in various natural language processing problems, for example, text…

Computation and Language · Computer Science 2020-09-28 Ehsan Doostmohammadi , Mohammad Hadi Bokaei , Hossein Sameti

Bilingual dictionaries are very important in various fields of natural language processing. In recent years, research on extracting new bilingual lexicons from non-parallel (comparable) corpora have been proposed. Almost all use a small…

Computation and Language · Computer Science 2019-09-23 Ebrahim Ansari , M. H. Sadreddini , Lucio Grandinetti , Mahsa Radinmehr , Ziba Khosravan , Mehdi Sheikhalishahi

We investigate structural traces of language contact in the intermediate representations of a monolingual language model. Focusing on Persian (Farsi) as a historically contact-rich language, we probe the representations of a Persian-trained…

Computation and Language · Computer Science 2026-01-29 Ali Basirat , Danial Namazifard , Navid Baradaran Hemmati
‹ Prev 1 8 9 10 Next ›