English
Related papers

Related papers: A data-based classification of Slavic languages: I…

200 papers

In this article we propose a novel method to estimate the frequency distribution of linguistic variables while controlling for statistical non-independence due to shared ancestry. Unlike previous approaches, our technique uses all available…

Populations and Evolution · Quantitative Biology 2021-03-22 Gerhard Jäger , Johannes Wahle

Inflection graphs are highly complex networks representing relationships between inflectional forms of words in human languages. For so-called synthetic languages, such as Latin or Polish, they have particularly interesting structure due to…

Cellular Automata and Lattice Gases · Physics 2023-12-18 Henryk Fukś , Babak Farzad , Yi Cao

In Linguistics, a grapheme is a written unit of a writing system corresponding to a phonological sound. In Natural Language Processing tasks, written language is analysed through two different mediums, word analysis, and character analysis.…

Computation and Language · Computer Science 2024-04-03 Samuel Rose , Chandrasekhar Kambhampati

Ordered sequences of univariate or multivariate regressions provide statistical models for analysing data from randomized, possibly sequential interventions, from cohort or multi-wave panel studies, but also from cross-sectional or…

Methodology · Statistics 2015-03-19 Nanny Wermuth , Kayvan Sadeghi

Dynamical processes can be transformed into graphs through a family of mappings called visibility algorithms, enabling the possibility of (i) making empirical data analysis and signal processing and (ii) characterising classes of dynamical…

Chaotic Dynamics · Physics 2015-06-18 Lucas Lacasa

To deal with distribution shifts in graph data, various graph out-of-distribution (OOD) generalization techniques have been recently proposed. These methods often employ a two-step strategy that first creates augmented environments and…

Machine Learning · Computer Science 2025-01-09 Song Wang , Xiaodong Yang , Rashidul Islam , Huiyuan Chen , Minghua Xu , Jundong Li , Yiwei Cai

The aim of this work is to obtain new inequalities for the variable symmetric division deg index $SDD_\alpha(G) = \sum_{uv \in E(G)} (d_u^\alpha/d_v^\alpha+d_v^\alpha/d_u^\alpha)$, and to characterize graphs extremal with respect to them.…

Combinatorics · Mathematics 2021-06-03 R. Aguilar-Sanchez , J. A. Mendez-Bermudez , Jose M. Rodriguez , Jose M. Sigarreta

In the graph clustering problem with a planted solution, the input is a graph on $n$ vertices partitioned into $k$ clusters, and the task is to infer the clusters from graph structure. A standard assumption is that clusters induce…

Data Structures and Algorithms · Computer Science 2025-11-24 Hendrik Fichtenberger , Michael Kapralov , Ekaterina Kochetkova , Silvio Lattanzi , Davide Mazzali , Weronika Wrzos-Kaminska

Most existing out-of-distribution (OOD) detection benchmarks classify samples with novel labels as the OOD data. However, some marginal OOD samples actually have close semantic contents to the in-distribution (ID) sample, which makes…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Xingming Long , Jie Zhang , Shiguang Shan , Xilin Chen

In this paper, we present an algorithm for evaluating lexical similarity between a given language and several reference language clusters. As an input, we have a list of concepts and the corresponding translations in all considered…

Computation and Language · Computer Science 2025-04-10 Karol Mikula , Mariana Sarkociová Remešíková

Out-of-distribution (OOD) learning deals with scenarios in which training and test data follow different distributions. Although general OOD problems have been intensively studied in machine learning, graph OOD is only an emerging area of…

Machine Learning · Computer Science 2022-09-28 Shurui Gui , Xiner Li , Limei Wang , Shuiwang Ji

Distribution shifts on graphs -- the data distribution discrepancies between training and testing a graph machine learning model, are often ubiquitous and unavoidable in real-world scenarios. Such shifts may severely deteriorate the…

Machine Learning · Computer Science 2024-02-20 Shuhan Liu , Kaize Ding

Words are fundamental linguistic units that connect thoughts and things through meaning. However, words do not appear independently in a text sequence. The existence of syntactic rules induces correlations among neighboring words. Using an…

Computation and Language · Computer Science 2023-03-15 David Sanchez , Luciano Zunino , Juan De Gregorio , Raul Toral , Claudio Mirasso

In this paper, we consider the problem of partitioning a small data sample of size $n$ drawn from a mixture of $2$ sub-gaussian distributions. Our work is motivated by the application of clustering individuals according to their population…

Statistics Theory · Mathematics 2023-01-05 Shuheng Zhou

Out-of-distribution (OOD) detection in graphs is critical for ensuring model robustness in open-world and safety-sensitive applications. Existing graph OOD detection approaches typically train an in-distribution (ID) classifier on ID data…

Machine Learning · Computer Science 2025-05-20 Haoyan Xu , Zhengtao Yao , Ziyi Wang , Zhan Cheng , Xiyang Hu , Mengyuan Li , Yue Zhao

A generalization of the classical statistics ``maj'' and ``inv'' (the major index and number of inversions) on words is introduced, parameterized by arbitrary graphs on the underlying alphabet. The question of characterizing those graphs…

Combinatorics · Mathematics 2008-02-03 Dominique Foata , Doron Zeilberger

The degree distribution is one of the most fundamental properties used in the analysis of massive graphs. There is a large literature on graph sampling, where the goal is to estimate properties (especially the degree distribution) of a…

Social and Information Networks · Computer Science 2018-08-29 Talya Eden , Shweta Jain , Ali Pinar , Dana Ron , C. Seshadhri

This paper presents categorization of Croatian texts using Non-Standard Words (NSW) as features. Non-Standard Words are: numbers, dates, acronyms, abbreviations, currency, etc. NSWs in Croatian language are determined according to Croatian…

Computation and Language · Computer Science 2014-11-18 Slobodan Beliga , Sanda Martinčić-Ipšić

Graph-structured data commonly have node annotations. A popular approach for inference and learning involving annotated graphs is to incorporate annotations into a statistical model or algorithm. By contrast, we consider a more direct…

Social and Information Networks · Computer Science 2020-10-07 Tatsuro Kawamoto

Traditional machine learning methods heavily rely on the independent and identically distribution assumption, which imposes limitations when the test distribution deviates from the training distribution. To address this crucial issue,…

Machine Learning · Computer Science 2024-03-26 Qin Tian , Wenjun Wang , Chen Zhao , Minglai Shao , Wang Zhang , Dong Li
‹ Prev 1 2 3 10 Next ›