中文
相关论文

相关论文: Should Corpora be Big, Rich, or Dense?

200 篇论文

We present an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. Surveying language documentation corpora and other resources that cover 67 languages and varieties…

计算与语言 · 计算机科学 2022-05-11 Andreas Liesenfeld , Mark Dingemanse

Objective: Large language models (LLMs) are attracting increasing interest in healthcare. This commentary evaluates the potential of LLMs to improve clinical prediction models (CPMs) for diagnostic and prognostic tasks, with a focus on…

计算机与社会 · 计算机科学 2025-11-07 Yusuf Yildiz , Goran Nenadic , Meghna Jani , David A. Jenkins

Recently, increasingly large amounts of data are generated from a variety of sources. Existing data processing technologies are not suitable to cope with the huge amounts of generated data. Yet, many research works focus on Big Data, a…

分布式、并行与集群计算 · 计算机科学 2018-06-07 Wissem Inoubli , Sabeur Aridhi , Haithem Mezni , Mondher Maddouri , Engelbert Mephu Nguifo

Recent work suggests that large language models (LLMs) can improve performance of speech tasks compared to existing systems. To support their claims, results on LibriSpeech and Common Voice are often quoted. However, this work finds that a…

音频与语音处理 · 电气工程与系统科学 2025-06-06 Yuan Tseng , Titouan Parcollet , Rogier van Dalen , Shucong Zhang , Sourav Bhattacharya

Social scientists are now using large language models to create "silicon samples": synthetic datasets intended to stand in for human respondents. However, producing these samples requires many analytic choices, including model selection,…

计算机与社会 · 计算机科学 2026-05-19 Jamie Cummins

Modern large-scale datasets are frequently said to be high-dimensional. However, their data point clouds frequently possess structures, significantly decreasing their intrinsic dimensionality (ID) due to the presence of clusters, points…

机器学习 · 计算机科学 2019-01-21 Luca Albergante , Jonathan Bac , Andrei Zinovyev

When learning from others, people tend to focus their attention on those with similar views. This is often attributed to flawed reasoning, and thought to slow learning and polarize beliefs. However, we show that echo chambers are a rational…

理论经济学 · 经济学 2025-06-10 Gabriel Martinez , Nicholas H. Tenev

Coresets have emerged as a powerful tool to summarize data by selecting a small subset of the original observations while retaining most of its information. This approach has led to significant computational speedups but the performance of…

统计理论 · 数学 2020-12-10 Paxton Turner , Jingbo Liu , Philippe Rigollet

Scaling has been proposed as a powerful tool to analyze the properties of complex systems, and in particular for cities where it describes how various properties change with population. The empirical study of scaling on a wide range of…

物理与社会 · 物理学 2018-04-18 Jules Depersin , Marc Barthelemy

This paper introduces CocoNut-Humoresque, an open-source large-scale speech likability corpus that includes speech segments and their per-listener likability scores. Evaluating voice likability is essential to designing preferable voices…

音频与语音处理 · 电气工程与系统科学 2024-07-08 Hitoshi Suda , Aya Watanabe , Shinnosuke Takamichi

We have studied both clusters and bulk systems while investigating amorphous states. We have varied the nature of interaction amongst the particles of the system under consideration in order to reveal the possible presence of universality…

统计力学 · 物理学 2008-12-31 Gurpreet Singh Matharoo

Word clouds are frequently used to analyze and communicate text data in many domains. In order to help guide research on improving the legibility of word clouds, we have conducted a survey of their usage in Digital Humanities academia and…

人机交互 · 计算机科学 2022-10-18 Rebecca M. M. Hicke , Maanya Goenka , Eric Alexander

Human languages vary widely in how they encode information within circumscribed semantic domains (e.g., time, space, color, human body parts and activities), but little is known about the global structure of semantic information and nothing…

计算与语言 · 计算机科学 2024-02-19 Pedro Aceves , James A. Evans

Over the last decade, random hyperbolic graphs have proved successful in providing geometric explanations for many key properties of real-world networks, including strong clustering, high navigability, and heterogeneous degree…

物理与社会 · 物理学 2023-03-01 Béatrice Désy , Patrick Desrosiers , Antoine Allard

High-dimensional data and high-dimensional representations of reality are inherent features of modern Artificial Intelligence systems and applications of machine learning. The well-known phenomenon of the "curse of dimensionality" states:…

机器学习 · 计算机科学 2020-01-22 Alexander N. Gorban , Valery A. Makarov , Ivan Y. Tyukin

Machine learning problems involving sparse datasets may benefit from the use of convolutional neural networks if the numbers of samples and features are very large. Such datasets are increasingly more frequently encountered in a variety of…

图像与视频处理 · 电气工程与系统科学 2020-05-21 Baris Kanber

This paper argues that large language models have a valuable scientific role to play in serving as scientific models of public languages. Linguistic study should not only be concerned with the cognitive processes behind linguistic…

计算与语言 · 计算机科学 2026-03-12 Jumbly Grindrod

The conformity bias exhibited by large language models (LLMs) can pose a significant challenge to decision-making in LLM-based multi-agent systems (LLM-MAS). While many prior studies have treated "conformity" simply as a matter of opinion…

人工智能 · 计算机科学 2026-04-22 Mikako Bito , Keita Nishimoto , Kimitaka Asatani , Ichiro Sakata

With the help of in-context learning (ICL), large language models (LLMs) have achieved impressive performance across various tasks. However, the function of descriptive instructions during ICL remains under-explored. In this work, we…

计算与语言 · 计算机科学 2025-09-09 Chenming Tang , Zhixiang Wang , Hao Sun , Yunfang Wu

Large language models (LLMs) are very performant connectionist systems, but do they exhibit more compositionality? More importantly, is that part of why they perform so well? We present empirical analyses across four LLM families (12…

计算与语言 · 计算机科学 2025-05-21 Ruchira Dhar , Anders Søgaard