中文
相关论文

相关论文: Should Corpora be Big, Rich, or Dense?

200 篇论文

While Large Language Models require more and more data to train and scale, rather than looking for any data to acquire, we should consider what types of tasks are more likely to benefit from data scaling. We should be intentional in our…

机器学习 · 计算机科学 2025-06-04 Tanya Rodchenko , Natasha Noy , Nino Scherrer

Recent studies have highlighted the potential of exploiting parallel corpora to enhance multilingual large language models, improving performance in both bilingual tasks, e.g., machine translation, and general-purpose tasks, e.g., text…

计算与语言 · 计算机科学 2025-02-11 Peiqin Lin , André F. T. Martins , Hinrich Schütze

Mental health risk prediction is a growing field in the speech community, but many studies are based on small corpora. This study illustrates how variations in test and train set sizes impact performance in a controlled study. Using a…

计算与语言 · 计算机科学 2025-01-03 Tomek Rutowski , Amir Harati , Elizabeth Shriberg , Yang Lu , Piotr Chlebek , Ricardo Oliveira

This paper presents an opinion on the potential of using large language models to query on both unstructured and structured data. It also outlines some research challenges related to the topic of building question-answering systems for both…

数据库 · 计算机科学 2023-07-07 Wang-Chiew Tan

Big data has ushered in a new wave of predictive power using machine learning models. In this work, we assess what {\it big} means in the context of typical materials-science machine-learning problems. This concerns not only data volume,…

A child's spoken ability continues to change until their adult age. Until 7-8yrs, their speech sound development and language structure evolve rapidly. This dynamic shift in their spoken communication skills and data privacy make it…

声音 · 计算机科学 2025-07-18 John Hansen , Satwik Dutta , Ellen Grand

Recent discussion of the success of feature selection methods has argued that focusing on a relatively small number of features has been counterproductive. Instead, it is suggested, the number of significant features can be in the thousands…

统计理论 · 数学 2014-07-10 Peter Hall , Jiashun Jin , Hugh Miller

Fully convolutional neural networks can process input of arbitrary size by applying a combination of downsampling and pooling. However, we find that fully convolutional image classifiers are not agnostic to the input size but rather show…

机器学习 · 计算机科学 2021-10-13 Mats L. Richter , Wolf Byttner , Ulf Krumnack , Ludwdig Schallner , Justin Shenk

Large-scale empirical data, the sample size and the dimension are high, often exhibit various characteristics. For example, the noise term follows unknown distributions or the model is very sparse that the number of critical variables is…

统计理论 · 数学 2018-06-18 Yuehan Yang , Hu Yang

Several studies have investigated the reasons behind the effectiveness of fine-tuning, usually through the lens of probing. However, these studies often neglect the role of the size of the dataset on which the model is fine-tuned. In this…

计算与语言 · 计算机科学 2022-03-21 Houman Mehrafarin , Sara Rajaee , Mohammad Taher Pilehvar

It is held as a truism that deep neural networks require large datasets to train effective models. However, large datasets, especially with high-quality labels, can be expensive to obtain. This study sets out to investigate (i) how large a…

信息检索 · 计算机科学 2019-01-31 Trond Linjordet , Krisztian Balog

The goal of the present chapter is to explore the possibility of providing the research (but also the industrial) community that commonly uses spoken corpora with a stable portfolio of well-documented standardised formats that allow a high…

计算与语言 · 计算机科学 2012-03-06 Laurent Romary , Andreas Witt

Given large datasets and sufficient compute, is it beneficial to design neural architectures for the structure and symmetries of each problem? Or is it more efficient to learn them from data? We study empirically how equivariant and…

机器学习 · 计算机科学 2025-07-29 Johann Brehmer , Sönke Behrends , Pim de Haan , Taco Cohen

The recent trend for acquiring big data assumes that possessing quantitatively more and qualitatively finer data necessarily provides an advantage that may be critical in competitive situations. Using a model complex adaptive system where…

物理与社会 · 物理学 2018-08-15 V. Sasidevan , Appilineni Kushal , Sitabhra Sinha

Understanding how size influences the internal characteristics of a system is a crucial concern across various fields. Concepts like scale invariance, universalities, and fractals are fundamental to this inquiry and find application in…

物理与社会 · 物理学 2024-04-04 Fabiano L. Ribeiro , Vinicius M. Netto

Models for text generation have become focal for many research tasks and especially for the generation of sentence corpora. However, understanding the properties of an automatically generated text corpus remains challenging. We propose a…

The material bases of information - paper, computer discs - usually scale with information quantity. Large quantities of information usually require large material bases. Conventional wisdom has it that human long-term memory locates within…

神经元与认知 · 定量生物学 2014-06-05 Donald R. Forsdyke

Recurrent neural networks can learn to predict upcoming words remarkably well on average; in syntactically complex contexts, however, they often assign unexpectedly high probabilities to ungrammatical words. We investigate to what extent…

计算与语言 · 计算机科学 2019-09-04 Marten van Schijndel , Aaron Mueller , Tal Linzen

The quantum capacity of a quantum channel is always smaller than the capacity of the channel for private communication. However, both quantities are given by the infinite regularization of respectively the coherent and the private…

量子物理 · 物理学 2015-07-28 David Elkouss , Sergii Strelchuk

A continued issue for those working with computational tools and endangered and under-resourced languages is the lower accuracy of results for languages with smaller amounts of data. We attempt to ameliorate this issue by using data…

计算与语言 · 计算机科学 2025-04-10 Alessio Tosolini , Claire Bowern
‹ 上一页 1 2 3 10 下一页 ›