中文
相关论文

相关论文: Gibberish Semantics: How Good is Russian Twitter i…

200 篇论文

Many Natural Language Processing applications nowadays rely on pre-trained word representations estimated from large text corpora such as news collections, Wikipedia and Web Crawl. In this paper, we show how to train high-quality word…

计算与语言 · 计算机科学 2017-12-29 Tomas Mikolov , Edouard Grave , Piotr Bojanowski , Christian Puhrsch , Armand Joulin

Increased popularity of different text representations has also brought many improvements in Natural Language Processing (NLP) tasks. Without need of supervised data, embeddings trained on large corpora provide us meaningful relations to be…

计算与语言 · 计算机科学 2020-02-14 Gökhan Güler , A. Cüneyd Tantuğ

Text alignment and text quality are critical to the accuracy of Machine Translation (MT) systems, some NLP tools, and any other text processing tasks requiring bilingual data. This research proposes a language independent bi-sentence…

计算与语言 · 计算机科学 2015-10-16 Krzysztof Wołk

Word2vec is a popular family of algorithms for unsupervised training of dense vector representations of words on large text corpuses. The resulting vectors have been shown to capture semantic relationships among their corresponding words,…

Understanding the semantic of a collection of texts is a challenging task. Topic models are probabilistic models that aims at extracting "topics" from a corpus of documents. This task is particularly difficult when the corpus is composed of…

信息检索 · 计算机科学 2022-03-22 Hugo Schnoering

Since word embeddings have been the most popular input for many NLP tasks, evaluating their quality is of critical importance. Most research efforts are focusing on English word embeddings. This paper addresses the problem of constructing…

计算与语言 · 计算机科学 2020-04-07 Stamatis Outsios , Christos Karatsalos , Konstantinos Skianis , Michalis Vazirgiannis

We generalize the word analogy task across languages, to provide a new intrinsic evaluation method for cross-lingual semantic spaces. We experiment with six languages within different language families, including English, German, Spanish,…

计算与语言 · 计算机科学 2018-07-12 Tomáš Brychcín , Stephen Eugene Taylor , Lukáš Svoboda

We first present our work in machine translation, during which we used aligned sentences to train a neural network to embed n-grams of different languages into an $d$-dimensional space, such that n-grams that are the translation of each…

机器学习 · 计算机科学 2011-05-17 Etter Vincent

Probabilistic topic models like Latent Dirichlet Allocation (LDA) have been previously extended to the bilingual setting. A fundamental modeling assumption in several of these extensions is that the input corpora are in the form of document…

计算与语言 · 计算机科学 2021-12-01 Georgios Balikas , Massih-Reza Amini , Marianne Clausel

Word embeddings are an essential instrument in many NLP tasks. Most available resources are trained on general language from Web corpora or Wikipedia dumps. However, word embeddings for domain-specific language are rare, in particular for…

计算与语言 · 计算机科学 2023-02-14 Ricardo Schiffers , Dagmar Kern , Daniel Hienert

We use referential translation machines (RTMs) to identify the similarity between an attribute and two words in English by casting the task as machine translation performance prediction (MTPP) between the words and the attribute word and…

计算与语言 · 计算机科学 2024-07-09 Ergun Biçici

Increasing popularity of Twitter in politics is subject to commercial and academic interest. To fully exploit the merits of this platform, reaching the target audience with desired political leanings is critical. This paper extends the…

社会与信息网络 · 计算机科学 2018-09-18 Kutlu Emre Yilmaz , Osman Abul

Tasks such as semantic search and clustering on crisis-related social media texts enhance our comprehension of crisis discourse, aiding decision-making and targeted interventions. Pre-trained language models have advanced performance in…

计算与语言 · 计算机科学 2024-03-26 Rabindra Lamsal , Maria Rodriguez Read , Shanika Karunasekera

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for mining such data from…

计算与语言 · 计算机科学 2015-11-20 Krzysztof Wołk , Emilia Rejmund , Krzysztof Marasek

A set of ontology matching algorithms (for finding correspondences between concepts) is based on a thesaurus that provides the source data for the semantic distance calculations. In this wiki era, new resources may spring up and improve…

信息检索 · 计算机科学 2009-10-12 A. A. Krizhanovsky , Feiyu Lin

Distributional semantics models derive word space from linguistic items in context. Meaning is obtained by defining a distance measure between vectors corresponding to lexical entities. Such vectors present several problems. In this paper…

计算与语言 · 计算机科学 2017-12-25 Jakub Dutkiewicz , Czesław Jędrzejek

There is a huge imbalance between languages currently spoken and corresponding resources to study them. Most of the attention naturally goes to the "big" languages: those which have the largest presence in terms of media and number of…

计算与语言 · 计算机科学 2019-04-02 Albina Khusainova , Adil Khan , Adín Ramírez Rivera

In this paper, we attempt to improve Statistical Machine Translation (SMT) systems on a very diverse set of language pairs (in both directions): Czech - English, Vietnamese - English, French - English and German - English. To accomplish…

计算与语言 · 计算机科学 2015-12-08 Krzysztof Wołk , Krzysztof Marasek

The article proposes a new architecture based on Multi-head attention to solve the problem of morphological tagging for the Russian language. The preprocessing of the word vectors includes splitting the words into subtokens, followed by a…

计算与语言 · 计算机科学 2026-04-06 K. Skibin , M. Pozhidaev , S. Suschenko

This paper explores the social quality (goodness) of community structures formed across Twitter users, where social links within the structures are estimated based upon semantic properties of user-generated content (corpus). We examined the…

社会与信息网络 · 计算机科学 2016-06-01 Kuntal Dey , Sahil Agrawal , Rahul Malviya , Saroj Kaushik