中文
相关论文

相关论文: Language discrimination and clustering via a neura…

200 篇论文

Technical documents contain a fair amount of unnatural language, such as tables, formulas, pseudo-codes, etc. Unnatural language can be an important factor of confusing existing NLP tools. This paper presents an effective method of…

信息检索 · 计算机科学 2017-03-20 Myungha Jang , Jinho D. Choi , James Allan

This paper introduces PhyloLM, a method adapting phylogenetic algorithms to Large Language Models (LLMs) to explore whether and how they relate to each other and to predict their performance characteristics. Our method calculates a…

计算与语言 · 计算机科学 2025-12-09 Nicolas Yax , Pierre-Yves Oudeyer , Stefano Palminteri

We describe and experimentally evaluate a method for automatically clustering words according to their distribution in particular syntactic contexts. Deterministic annealing is used to find lowest distortion sets of clusters. As the…

cmp-lg · 计算机科学 2008-02-03 Fernando Pereira , Naftali Tishby , Lillian Lee

Patterns of topological arrangement are widely used for both animal and human brains in the learning process. Nevertheless, automatic learning techniques frequently overlook these patterns. In this paper, we apply a learning technique based…

计算与语言 · 计算机科学 2013-07-09 Thiago C. Silva , Diego R. Amancio

In this paper, we describe a research method that generates Bangla word clusters on the basis of relating to meaning in language and contextual similarity. The importance of word clustering is in parts of speech (POS) tagging, word sense…

计算与语言 · 计算机科学 2017-01-31 Dipaloke Saha , Md Saddam Hossain , MD. Saiful Islam , Sabir Ismail

This paper describes our submission (named clac) to the 2016 Discriminating Similar Languages (DSL) shared task. We participated in the closed Sub-task 1 (Set A) with two separate machine learning techniques. The first approach is a…

计算与语言 · 计算机科学 2017-08-14 Andre Cianflone , Leila Kosseim

Language Identification (LID) is the task of determining the language of a given text and is a fundamental preprocessing step that affects the reliability of downstream NLP applications. While recent work has expanded LID coverage for…

计算与语言 · 计算机科学 2026-01-30 Sang Yun Kwon , AbdelRahim Elmadany , Muhammad Abdul-Mageed

We present a novel deep-learning-based method to cluster words in documents which we apply to detect and recognize tables given the OCR output. We interpret table structure bottom-up as a graph of relations between pairs of words (belonging…

机器学习 · 计算机科学 2024-05-24 Marek Polewczyk , Marco Spinaci

We analyze here a particular kind of linguistic network where vertices representwords and edges stand for syntactic relationships between words. The statisticalproperties of these networks have been recently studied and various features…

统计力学 · 物理学 2007-05-23 Ramon Ferrer i Cancho , Andrea Capocci , Guido Caldarelli

This paper is a presentation of a new method for denoising images using Haralick features and further segmenting the characters using artificial neural networks. The image is divided into kernels, each of which is converted to a GLCM (Gray…

计算机视觉与模式识别 · 计算机科学 2021-07-27 P Preethi , Hrishikesh Viswanath

Recognition of text on word or line images, without the need for sub-word segmentation has become the mainstream of research and development of text recognition for Indian languages. Modelling unsegmented sequences using Connectionist…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Minesh Mathew , Ajoy Mondal , CV Jawahar

This research conducts a comparative study on multilingual text classification methods, utilizing deep learning and embedding visualization. The study employs LangDetect, LangId, FastText, and Sentence Transformer on a dataset encompassing…

计算与语言 · 计算机科学 2023-12-08 Arinjay Wyawhare

We focus on the task of unsupervised lemmatization, i.e. grouping together inflected forms of one word under one label (a lemma) without the use of annotated training data. We propose to perform agglomerative clustering of word forms with a…

计算与语言 · 计算机科学 2019-08-23 Rudolf Rosa , Zdeněk Žabokrtský

Natural languages are complexly structured entities. They exhibit characterising regularities that can be exploited to link them one another. In this work, I compare two morphological aspects of languages: Written Patterns and Sentence…

计算与语言 · 计算机科学 2019-07-09 Alberto Calderone

Recently, the focus of complex networks research has shifted from the analysis of isolated properties of a system toward a more realistic modeling of multiple phenomena - multilayer networks. Motivated by the prosperity of multilayer…

计算与语言 · 计算机科学 2015-07-31 Domagoj Margan , Ana Meštrović , Sanda Martinčić-Ipšić

In this paper, we try to explore the evolution of language through case calculations. First, we chose the novels of eleven British writers from 1400 to 2005 and found the corresponding works; Then, we use the natural language processing…

计算与语言 · 计算机科学 2018-10-09 Zhu Gao , Yanhui Jiang , Junhui Gao

The patterns in which the syntax of different languages converges and diverges are often used to inform work on cross-lingual transfer. Nevertheless, little empirical work has been done on quantifying the prevalence of different syntactic…

计算与语言 · 计算机科学 2020-07-14 Dmitry Nikolaev , Ofir Arviv , Taelin Karidi , Neta Kenneth , Veronika Mitnik , Lilja Maria Saeboe , Omri Abend

The use of terms from natural and social scientific titles and abstracts is studied from the perspective of sublanguages and their specialized dictionaries. Different notions of sublanguage distinctiveness are explored. Objective methods…

cmp-lg · 计算机科学 2008-02-03 Robert M. Losee , Stephanie W. Haas

Hierarchical graph clustering is a common technique to reveal the multi-scale structure of complex networks. We propose a novel metric for assessing the quality of a hierarchical clustering. This metric reflects the ability to reconstruct…

社会与信息网络 · 计算机科学 2018-07-16 Thomas Bonald , Bertrand Charpentier

Language exhibits structure at different scales, ranging from subwords to words, sentences, paragraphs, and documents. To what extent do deep models capture information at these scales, and can we force them to better capture structure…

计算与语言 · 计算机科学 2020-11-11 Alex Tamkin , Dan Jurafsky , Noah Goodman