中文
相关论文

相关论文: A Quantitative Confirmation of the Currier Languag…

200 篇论文

AI-generated text detectors have recently gained adoption in educational and professional contexts. Prior research has uncovered isolated cases of bias, particularly against English Language Learners (ELLs) however, there is a lack of…

人工智能 · 计算机科学 2025-12-15 Priyam Basu , Yunfeng Zhang , Vipul Raheja

Laser light is widely used for communication and sensing applications, so the optimal discrimination of coherent states--the quantum states of light emitted by a laser--has immense practical importance. However, quantum mechanics imposes a…

量子物理 · 物理学 2013-05-24 Marcus P. da Silva , Saikat Guha , Zachary Dutton

Given a regular language $L$, we study the language of words $\mathsf{D}(L)$, that distinguish between pairs of different left-quotients of $L$. We characterize this distinguishability operation, show that its iteration has always a fixed…

形式语言与自动机理论 · 计算机科学 2014-12-11 Cezar Câmpeanu , Nelma Moreira , Rogério Reis

Inertial measurement unit-based online handwriting recognition enables the recognition of input signals collected across different writing surfaces but remains challenged by uneven character distributions and inter-writer variability. In…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Jindong Li , Dario Zanca , Vincent Christlein , Tim Hamann , Jens Barth , Peter Kämpf , Björn Eskofier

This paper presents methods to discriminate between languages and dialects written in Cuneiform script, one of the first writing systems in the world. We report the results obtained by the PZ team in the Cuneiform Language Identification…

计算与语言 · 计算机科学 2019-04-30 Gustavo Henrique Paetzold , Marcos Zampieri

Choosing an appropriate tokenization scheme is often a bottleneck in low-resource cross-lingual transfer. To understand the downstream implications of text representation choices, we perform a comparative analysis on language models having…

计算与语言 · 计算机科学 2023-10-13 Md Mushfiqur Rahman , Fardin Ahsan Sakib , Fahim Faisal , Antonios Anastasopoulos

In this paper, on the basis of a (Fenchel) duality theory on the continuous level, we derive an $\textit{a posteriori}$ error identity for arbitrary conforming approximations of a primal formulation and a dual formulation of variational…

数值分析 · 数学 2024-10-25 Harbir Antil , Sören Bartels , Alex Kaltenbach , Rohit Khandelwal

There have been many recent advances in the structure and measurement of distributed language models: those that map from words to a vector-space that is rich in information about word choice and composition. This vector-space is the…

计算与语言 · 计算机科学 2015-07-27 Matt Taddy

Texts exhibit considerable stylistic variation. This paper reports an experiment where a corpus of documents (N= 75 000) is analyzed using various simple stylistic metrics. A subset (n = 1000) of the corpus has been previously assessed to…

cmp-lg · 计算机科学 2008-02-03 Jussi Karlgren

A new class of analog-to-digital (A/D) and digital-to-analog (D/A) converters using a flaky quantiser, called the $\beta$-encoder, has been shown to have exponential bit rate accuracy while possessing a self-correction property for…

信息论 · 计算机科学 2009-07-28 Tohru Kohda , Satoshi Hironaka , Kazuyuki Aihara

This paper provides an experimentally validated, probabilistic model of file behavior when consumed by a set of pre-existing parsers. File behavior is measured by way of a standardized set of Boolean "messages" produced as the files are…

计算工程、金融与科学 · 计算机科学 2022-09-23 Michael Robinson , Letitia W. Li , Cory Anderson , Steve Huntsman

Detection of transitions between broad phonetic classes in a speech signal is an important problem which has applications such as landmark detection and segmentation. The proposed hierarchical method detects silence to non-silence…

声音 · 计算机科学 2014-11-04 T V Ananthapadmanabha , K V Vijay Girish , A G Ramakrishnan

Language models typically tokenize text into subwords, using a deterministic, hand-engineered heuristic of combining characters into longer surface-level strings such as 'ing' or whole words. Recent literature has repeatedly shown the…

计算与语言 · 计算机科学 2023-10-19 Avijit Thawani , Saurabh Ghanekar , Xiaoyuan Zhu , Jay Pujara

Dialects exhibit a substantial degree of variation due to the lack of a standard orthography. At the same time, the ability of Large Language Models (LLMs) to process dialects remains largely understudied. To address this gap, we use…

计算与语言 · 计算机科学 2025-09-23 Robert Litschko , Verena Blaschke , Diana Burkhardt , Barbara Plank , Diego Frassinelli

Recent evidence suggests that analyzing the presence/absence of taxonomic features can offer a compelling alternative to differential abundance analysis in microbiome studies. However, standard approaches to differential prevalence analysis…

统计方法学 · 统计学 2026-05-26 Juho Pelto , Kari Auranen , Janne V. Kujala , Leo Lahti

Language Identification (LID) is a core task in multilingual NLP, yet current systems often overfit to clean, monolingual data. This work introduces DIVERS-BENCH, a comprehensive evaluation of state-of-the-art LID models across diverse…

计算与语言 · 计算机科学 2025-09-23 Jessica Ojo , Zina Kamel , David Ifeoluwa Adelani

A/B testing refers to the task of determining the best option among two alternatives that yield random outcomes. We provide distribution-dependent lower bounds for the performance of A/B testing that improve over the results currently…

统计理论 · 数学 2015-02-25 Emilie Kaufmann , Olivier Cappé , Aurélien Garivier

In this work, we use a moving Voronoi and sharp interface approach for simulating two-phase flows. At every time step, the mesh is generated anew from Voronoi seeds that behave as material points. The paper is a continuation of our previous…

数值分析 · 数学 2025-03-18 Ondřej Kincl , Ilya Peshkov , Walter Boscheri

This study investigates the register variation in texts written by humans and comparable texts produced by large language models (LLMs). Biber's multidimensional analysis (MDA) is applied to a sample of human-written texts and AI-created…

计算与语言 · 计算机科学 2025-09-22 Jiří Milička , Anna Marklová , Václav Cvrček

We propose a Bayesian approach to learn discriminative dictionaries for sparse representation of data. The proposed approach infers probability distributions over the atoms of a discriminative dictionary using a Beta Process. It also…

计算机视觉与模式识别 · 计算机科学 2015-03-30 Naveed Akhtar , Faisal Shafait , Ajmal Mian