中文
相关论文

相关论文: Comment on "Linguistic Features of Noncoding DNA S…

200 篇论文

Context-free languages can be characterized in several ways. This article studies projective linearisations of languages of simple dependency trees, i.e., dependency trees in which a node can govern at most one node with a given syntactic…

形式语言与自动机理论 · 计算机科学 2024-01-17 Carles Cardó

Recent breakthroughs in language models (LMs) using neural networks have raised the question: how similar are these models' processing to human language processing? Results using a framework called Brain Score (BS) -- predicting fMRI…

计算与语言 · 计算机科学 2026-04-20 Jingnong Qu , Ashvin Ranjan , Shane Steinert-Threlkeld

The possibility of detecting mutations in a DNA from force measurements (as a first step towards sequence analysis) is discussed theoretically based on exact calculations. The force signal is associated with the domain wall separating the…

统计力学 · 物理学 2009-11-07 Somendra M. Bhattacharjee , D. Marenduzzo

A curious observation was made that the rank statistics of scientific citation numbers follows Zipf-Mandelbrot's law. The same pow-like behavior is exhibited by some simple random citation models. The observed regularity indicates not so…

物理与社会 · 物理学 2007-05-23 Z. K. Silagadze

Shannon information (SI) and its special case, divergence, are defined for a DNA sequence in terms of probabilities of chemical words in the sequence and are computed for a set of complete genomes highly diverse in length and composition.…

基因组学 · 定量生物学 2009-11-10 Hong-Da Chen , Chang-Heng Chang , Li-Ching Hsieh , Hoong-Chien Lee

This paper studies the effect of linguistic constraints on the large scale organization of language. It describes the properties of linguistic networks built using texts of written language with the words randomized. These properties are…

计算与语言 · 计算机科学 2011-02-16 Madhav Krishna , Ahmed Hassan , Yang Liu , Dragomir Radev

We describe an approach to robust domain-independent syntactic parsing of unrestricted naturally-occurring (English) input. The technique involves parsing sequences of part-of-speech and punctuation labels using a unification-based grammar…

cmp-lg · 计算机科学 2008-02-03 Ted Briscoe , John Carroll

The dependence of the frequency distributions due to multiple meanings of words in a text is investigated by deleting letters. By coding the words with fewer letters the number of meanings per coded word increases. This increase is measured…

计算与语言 · 计算机科学 2017-10-04 Xiaoyong Yan , Petter Minnhagen

We show that textual analysis of microbial genomes reveal telling footprints of the early evolution of the genomes. The frequencies of word occurrence of random DNA sequences considered as texts in their four nucleotides are expected to…

生物物理 · 物理学 2007-05-23 Li-Ching Hsieh , Liaofu Luo , HC Lee

Research in quantitative evolutionary genomics and systems biology led to the discovery of several universal regularities connecting genomic and molecular phenomic variables. These universals include the log-normal distribution of the…

种群与进化 · 定量生物学 2015-05-30 Eugene V. Koonin

In this paper, a new statistic feature of the discrete short-time amplitude spectrum is discovered by experiments for the signals of unvoiced pronunciation. For the random-varying short-time spectrum, this feature reveals the relationship…

声音 · 计算机科学 2016-12-22 Xiaodong Zhuang

Evaluating whether large language models (LLMs) capture the structure of natural language beyond local fluency remains an open challenge. Existing evaluation methods, largely based on task performance or short-context behavior, provide…

计算与语言 · 计算机科学 2026-05-26 Kumiko Tanaka-Ishii

We have presented the basic knowledge on the structure of molecules coding the genetic information, mechanisms of transfer of this information from DNA to proteins and phenomena connected with replication of DNA. In particular, we have…

基因组学 · 定量生物学 2009-09-30 Dorota Mackiewicz , Stanislaw Cebrat

Large Language Models (LLMs) do not differentially represent numbers, which are pervasive in text. In contrast, neuroscience research has identified distinct neural representations for numbers and words. In this work, we investigate how…

人工智能 · 计算机科学 2024-01-10 Raj Sanjay Shah , Vijay Marupudi , Reba Koenen , Khushi Bhardwaj , Sashank Varma

Background: In addition to known protein-coding genes, large amount of apparently non-coding sequence are conserved between the human and mouse genomes. It seems reasonable to assume that these conserved regions are more likely to contain…

基因组学 · 定量生物学 2007-05-23 Thomas A. Down , Tim J. P. Hubbard

We study the amount of reliable information that can be stored in a DNA-based storage system with noisy sequencing, where each codeword is composed of short DNA molecules. We analyze a concatenated coding scheme, where the outer code is…

信息论 · 计算机科学 2026-05-19 Ran Tamir , Nir Weinberger , Albert Guillén i Fàbregas

Quantifying the similarity between symbolic sequences is a traditional problem in Information Theory which requires comparing the frequencies of symbols in different sequences. In numerous modern applications, ranging from DNA over music to…

物理与社会 · 物理学 2016-04-18 Martin Gerlach , Francesc Font-Clos , Eduardo G. Altmann

We study the statistical properties of random numbers under the Martin-L\"of definition of randomness, proving that random numbers obey analogues of Strong Law of Large Numbers, the Law of the Iterated Logarithm, and that they are normal.…

逻辑 · 数学 2014-10-14 Matthew Pancia

Natural language processing has made significant inroads into learning the semantics of words through distributional approaches, however representations learnt via these methods fail to capture certain kinds of information implicit in the…

The Software Naturalness hypothesis argues that programming languages can be understood through the same techniques used in natural language processing. We explore this hypothesis through the use of a pre-trained transformer-based language…