中文
相关论文

相关论文: Lexical Co-occurrence, Statistical Significance, a…

200 篇论文

Beyond bibliometrics, there is interest in characterizing the evolution of the number of ideas in scientific papers. A common approach for investigating this involves analyzing the titles of publications to detect vocabulary changes over…

计算与语言 · 计算机科学 2022-08-31 James Powell , Martin Klein , Lyudmila Balakireva

In statistical classification and machine learning, as well as in social and other sciences, a number of measures of association have been proposed for assessing and comparing individual classifiers, raters, as well as their groups. In this…

机器学习 · 统计学 2020-02-04 Nadezhda Gribkova , Ričardas Zitikis

In an effort to better understand meaning from natural language texts, we explore methods aimed at organizing lexical objects into contexts. A number of these methods for organization fall into a family defined by word ordering. Unlike…

Measuring association, or the lack of it, between variables plays an important role in a variety of research areas, including education, which is of our primary interest in this paper. Given, for example, student marks on several study…

统计方法学 · 统计学 2015-06-10 Danang Teguh Qoyyimi , Ricardas Zitikis

We suggest an information-theoretic approach for measuring stylistic coordination in dialogues. The proposed measure has a simple predictive interpretation and can account for various confounding factors through proper conditioning. We…

计算与语言 · 计算机科学 2015-08-28 Shuyang Gao , Greg Ver Steeg , Aram Galstyan

Most modern computational approaches to lexical semantic change detection (LSC) rely on embedding-based distributional word representations with neural networks. Despite the strong performance on LSC benchmarks, they are often opaque. We…

计算与语言 · 计算机科学 2026-05-05 Bach Phan-Tat , Kris Heylen , Dirk Geeraerts , Stefano De Pascale , Dirk Speelman

A correlation is a binary vector that encodes all possible positions of overlaps of two words, where an overlap for an ordered pair of words (u,v) occurs if a suffix of word u matches a prefix of word v. As multiple pairs can have the same…

离散数学 · 计算机科学 2025-06-03 Eric Rivals , Pengfei Wang

Speech sounds of the languages all over the world show remarkable patterns of cooccurrence. In this work, we attempt to automatically capture the patterns of cooccurrence of the consonants across languages and at the same time figure out…

物理与社会 · 物理学 2009-11-11 Animesh Mukherjee , Monojit Choudhury , Anupam Basu , Niloy Ganguly

Regular sound correspondences constitute the principal evidence in historical language comparison. Despite the heuristic focus on regularity, it is often more an intuitive judgement than a quantified evaluation, and irregularity is more…

计算与语言 · 计算机科学 2026-02-03 Frederic Blum , Johann-Mattis List

Existing weakly supervised sound event detection (WSSED) work has not explored both types of co-occurrences simultaneously, i.e., some sound events often co-occur, and their occurrences are usually accompanied by specific background sounds,…

声音 · 计算机科学 2023-03-13 Yifei Xin , Dongchao Yang , Fan Cui , Yujun Wang , Yuexian Zou

We present a new method to detect anomalies in texts (in general: in sequences of any data), using language models, in a totally unsupervised manner. The method considers probabilities (likelihoods) generated by a language model, but…

计算与语言 · 计算机科学 2024-09-06 Filip Graliński , Ryszard Staruch , Krzysztof Jurkiewicz

Identifying the relations that exist between words (or entities) is important for various natural language processing tasks such as, relational search, noun-modifier classification and analogy detection. A popular approach to represent the…

计算与语言 · 计算机科学 2017-09-06 Huda Hakami , Danushka Bollegala

Co-occurrence matrices, such as co-citation, co-word, and co-link matrices, have been used widely in the information sciences. However, confusion and controversy have hindered the proper statistical analysis of this data. The underlying…

信息检索 · 计算机科学 2009-11-19 Loet Leydesdorff , Liwen Vaughan

Given a random text over a finite alphabet, we study the frequencies at which fixed-length words occur as subsequences. As the data size grows, the joint distribution of word counts exhibits a rich asymptotic structure. We investigate all…

概率论 · 数学 2026-05-06 Chaim Even-Zohar , Tsviqa Lakrec , Ran J. Tessler

Topic models extract representative word sets - called topics - from word counts in documents without requiring any semantic annotations. Topics are not guaranteed to be well interpretable, therefore, coherence measures have been proposed…

机器学习 · 计算机科学 2014-03-26 Frank Rosner , Alexander Hinneburg , Michael Röder , Martin Nettling , Andreas Both

Lexical features are a major source of information in state-of-the-art coreference resolvers. Lexical features implicitly model some of the linguistic phenomena at a fine granularity level. They are especially useful for representing the…

计算与语言 · 计算机科学 2017-04-25 Nafise Sadat Moosavi , Michael Strube

This paper presents a novel research problem on joint discovery of commonalities and differences between two individual documents (or document sets), called Comparative Document Analysis (CDA). Given any pair of documents from a document…

信息检索 · 计算机科学 2015-10-27 Xiang Ren , Yuanhua Lv , Kuansan Wang , Jiawei Han

With the arrival of digital era and Internet, the lack of information control provides an incentive for people to freely use any content available to them. Plagiarism occurs when users fail to credit the original owner for the content…

其他计算机科学 · 计算机科学 2010-03-26 Chien-Ying Chen , Jen-Yuan Yeh , Hao-Ren Ke

Unsupervised vector representations of sentences or documents are a major building block for many language tasks such as sentiment classification. However, current methods are uninterpretable and slow or require large training datasets.…

计算与语言 · 计算机科学 2020-11-03 Eric Zelikman , Richard Socher

Lexical ambiguity is widespread in language, allowing for the reuse of economical word forms and therefore making language more efficient. If ambiguous words cannot be disambiguated from context, however, this gain in efficiency might make…

计算与语言 · 计算机科学 2024-05-29 Tiago Pimentel , Rowan Hall Maudslay , Damián Blasi , Ryan Cotterell