English
Related papers

Related papers: Comment on "Linguistic Features of Noncoding DNA S…

200 papers

Deciphering natural language from brain activity through non-invasive devices remains a formidable challenge. Previous non-invasive decoders either require multiple experiments with identical stimuli to pinpoint cortical regions and enhance…

Computation and Language · Computer Science 2024-03-20 Cenyuan Zhang , Xiaoqing Zheng , Ruicheng Yin , Shujie Geng , Jianhan Xu , Xuan Gao , Changze Lv , Zixuan Ling , Xuanjing Huang , Miao Cao , Jianfeng Feng

Symbolic sequences such as written language and genomic DNA display characteristic frequency distributions and long-range correlations extending over many symbols. In language, this takes the form of Zipf's law for word frequencies together…

Computation and Language · Computer Science 2026-03-04 Marcelo A. Montemurro , Mirko Degli Esposti

In contrast with animal communication systems, diversity is characteristic of almost every aspect of human language. Languages variously employ tones, clicks, or manual signs to signal differences in meaning; some languages lack the…

Physics and Society · Physics 2013-02-14 Andrea Baronchelli , Nick Chater , Romualdo Pastor-Satorras , Morten H. Christiansen

Much of the on-going statistical analysis of DNA sequences is focused on the estimation of characteristics of coding and non-coding regions that would possibly allow discrimination of these regions. In the current approach, we concentrate…

Genomics · Quantitative Biology 2009-11-10 D. Kugiumtzis , A. Provata

The natural language generation (NLG) component of a spoken dialogue system (SDS) usually needs a substantial amount of handcrafting or a well-labeled dataset to be trained on. These limitations add significantly to development costs and…

Computation and Language · Computer Science 2015-08-10 Tsung-Hsien Wen , Milica Gasic , Dongho Kim , Nikola Mrksic , Pei-Hao Su , David Vandyke , Steve Young

Self-normalizing discriminative models approximate the normalized probability of a class without having to compute the partition function. In the context of language modeling, this property is particularly appealing as it may significantly…

Computation and Language · Computer Science 2018-06-05 Jacob Goldberger , Oren Melamud

Understanding the vulnerability of linguistic features extracted from noisy text is important for both developing better health text classification models and for interpreting vulnerabilities of natural language models. In this paper, we…

Computation and Language · Computer Science 2019-10-02 Jekaterina Novikova , Aparna Balagopalan , Ksenia Shkaruta , Frank Rudzicz

Recently, several studies have investigated the transcription process associated to specific genetic regulatory networks. In this work, we present a stochastic approach for analyzing the dynamics and effect of negative feedback loops (FBL)…

Molecular Networks · Quantitative Biology 2007-08-20 J. C. Nacher , T. Ochiai

In codeswitching contexts, the language of a syntactic head determines the distribution of its complements. Mahootian 1993 derives this generalization by representing heads as the anchors of elementary trees in a lexicalized TAG. However,…

cmp-lg · Computer Science 2008-02-03 Shahrzad Mahootian , Beatrice Santorini

Scripted dialogues such as movie and TV subtitles constitute a widespread source of training data for conversational NLP models. However, there are notable linguistic differences between these dialogues and spontaneous interactions,…

Computation and Language · Computer Science 2024-10-04 Ildikó Pilán , Laurent Prévot , Hendrik Buschmeier , Pierre Lison

The United States Code (Code) is an important source of Federal law that is produced by the interactions of many heterogeneous actors in a complex, dynamic space. The Code can be represented as the union of a hierarchical network and a…

Physics and Society · Physics 2010-03-24 Michael J. Bommarito , Daniel Martin Katz

Large language models generate text through probabilistic sampling from high-dimensional distributions, yet how this process reshapes the structural statistical organization of language remains incompletely characterized. Here we show that…

Computation and Language · Computer Science 2026-02-23 Ortal Hadad , Edoardo Loru , Jacopo Nudo , Niccolò Di Marco , Matteo Cinelli , Walter Quattrociocchi

We consider the problem of coding for the substring channel, in which information strings are observed only through their (multisets of) substrings. Due to existing DNA sequencing techniques and applications in DNA-based storage systems,…

Information Theory · Computer Science 2024-03-27 Yonatan Yehezkeally , Nikita Polyanskii

Random Telegraph Noise is a ubiquitous process manifesting across technology and the natural world. It is characterized by random jumps between two distinct states with Poissonian waiting times, and is the origin of 1/f noise. Understanding…

We present a comparison of word-based and character-based sequence-to-sequence models for data-to-text natural language generation, which generate natural language descriptions for structured inputs. On the datasets of two recent generation…

Computation and Language · Computer Science 2018-10-12 Glorianna Jagfeld , Sabrina Jenne , Ngoc Thang Vu

Code-switching---the intra-utterance use of multiple languages---is prevalent across the world. Within text-to-speech (TTS), multilingual models have been found to enable code-switching. By modifying the linguistic input to…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Marlene Staib , Tian Huey Teh , Alexandra Torresquintero , Devang S Ram Mohan , Lorenzo Foglianti , Raphael Lenain , Jiameng Gao

Some systems cannot be predicted by classical theories and it is required the development of combined deterministic and stochastic theories that make used of noise for dynamical prediction. Noise is not always an interfering signal which…

Adaptation and Self-Organizing Systems · Physics 2019-05-14 Alexandra Pinto Castellanos

Large language models with a huge number of parameters, when trained on near internet-sized number of tokens, have been empirically shown to obey neural scaling laws: specifically, their performance behaves predictably as a power law in…

Machine Learning · Computer Science 2022-11-01 Alexander Maloney , Daniel A. Roberts , James Sully

A statistical classification algorithm and its application to language identification from noisy input are described. The main innovation is to compute confidence limits on the classification, so that the algorithm terminates when enough…

Computation and Language · Computer Science 2007-05-23 David Elworthy

The physics of randomness and regularities for languages (mother tongues) and their lifetimes and family trees and for the second languages are studied in terms of two opposite processes; random multiplicative noise [1], and fragmentation…

Physics and Society · Physics 2007-07-20 Caglar Tuncay
‹ Prev 1 3 4 5 6 7 10 Next ›