中文
相关论文

相关论文: Should Corpora be Big, Rich, or Dense?

200 篇论文

The amount of information in the form of features and variables avail- able to machine learning algorithms is ever increasing. This can lead to classifiers that are prone to overfitting in high dimensions, high di- mensional models do not…

机器学习 · 计算机科学 2014-02-12 Aaron Karper

Big data features not only large volumes of data but also data with complicated structures. Complexity imposes unique challenges in big data analytics. Meeker and Hong (2014, Quality Engineering, pp. 102-116) provided an extensive…

应用统计 · 统计学 2018-03-19 Yili Hong , Man Zhang , William Q. Meeker

Collecting more diverse and representative training data is often touted as a remedy for the disparate performance of machine learning predictors across subpopulations. However, a precise framework for understanding how dataset properties…

机器学习 · 计算机科学 2021-06-08 Esther Rolf , Theodora Worledge , Benjamin Recht , Michael I. Jordan

A formula for an average connectivity between cortical areas in mammals is derived. Based on comparative neuroanatomical data, it is found, surprisingly, that this connectivity is either only weakly dependent or independent of brain size.…

神经元与认知 · 定量生物学 2007-05-23 Jan Karbowski

Urbanization promotes economy, mobility, access and availability of resources, but on the other hand, generates higher levels of pollution, violence, crime, and mental distress. The health consequences of the agglomeration of people living…

物理与社会 · 物理学 2021-02-09 Luis E. C. Rocha , Anna E. Thorson , Renaud Lambiotte

Real-world datasets are often of high dimension and effected by the curse of dimensionality. This hinders their comprehensibility and interpretability. To reduce the complexity feature selection aims to identify features that are crucial to…

机器学习 · 计算机科学 2023-04-18 Maximilian Stubbemann , Tobias Hille , Tom Hanika

Large language models (LLM) have emerged as a powerful tool for AI, with the key ability of in-context learning (ICL), where they can perform well on unseen tasks based on a brief series of task examples without necessitating any…

机器学习 · 计算机科学 2024-05-31 Zhenmei Shi , Junyi Wei , Zhuoyan Xu , Yingyu Liang

This study reviews the topic of big data management in the 21st-century. There are various developments that have facilitated the extensive use of that form of data in different organizations. The most prominent beneficiaries are internet…

计算机与社会 · 计算机科学 2015-09-08 Okal Christopher Otieno

We study the scaling properties of latent diffusion models (LDMs) with an emphasis on their sampling efficiency. While improved network architecture and inference algorithms have shown to effectively boost sampling efficiency of diffusion…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Kangfu Mei , Zhengzhong Tu , Mauricio Delbracio , Hossein Talebi , Vishal M. Patel , Peyman Milanfar

Although Perplexity is a widely used performance metric for language models, the values are highly dependent upon the number of words in the corpus and is useful to compare performance of the same corpus only. In this paper, we propose a…

计算与语言 · 计算机科学 2020-11-30 Jihyeon Roh , Sang-Hoon Oh , Soo-Young Lee

We surely enjoy the larger the better models for their superior performance in the last couple of years when both the hardware and software support the birth of such extremely huge models. The applied fields include text mining and others.…

计算与语言 · 计算机科学 2024-06-04 Hanjuan Huang , Hao-Jia Song , Hsing-Kuo Pao

Robust estimation is much more challenging in high dimensions than it is in one dimension: Most techniques either lead to intractable optimization problems or estimators that can tolerate only a tiny fraction of errors. Recent work in…

机器学习 · 计算机科学 2018-03-14 Ilias Diakonikolas , Gautam Kamath , Daniel M. Kane , Jerry Li , Ankur Moitra , Alistair Stewart

Algorithmic recourse provides counterfactual action plans that help people overturn unfavorable AI decisions. While diverse recourse sets may improve transparency and motivation, they may also impose cognitive load and negative emotions by…

人机交互 · 计算机科学 2026-05-13 Tomu Tominaga , Naomi Yamashita , Takeshi Kurashima

While many large infrastructure networks, such as power, water, and natural gas systems, have similar physical properties governing flows, these systems tend to have distinctly different sizes and topological structures. This paper seeks to…

物理与社会 · 物理学 2015-10-30 Paul D. H. Hines , Seth Blumsack , Markus Schläpfer

Big Data can mean different things to different people. The scale and challenges of Big Data are often described using three attributes, namely Volume, Velocity and Variety (3Vs), which only reflect some of the aspects of data. In this…

分布式、并行与集群计算 · 计算机科学 2016-01-14 Caesar Wu , Rajkumar Buyya , Kotagiri Ramamohanarao

The size-dependent nature of the so-called group or departmental h-index is reconsidered in this paper. While the influence of unit size on such collective measures was already demonstrated a decade ago, institutional ratings based on this…

数字图书馆 · 计算机科学 2022-02-02 Olesya Mryglod , Yurij Holovatch , Ralph Kenna

The claim that large-scale structure data independently prefers the Lambda Cold Dark Matter model is a myth. However, an updated compilation of large-scale structure observations cannot rule out Lambda CDM at 95% confidence. We explore the…

天体物理学 · 物理学 2007-05-23 Eric Gawiser

Rich clusters of galaxies are the most massive virialized systems known. Even though they contain only a small fraction of all galaxies, rich clusters provide a powerful tool for the study of galaxy formation, dark matter, large-scale…

天体物理学 · 物理学 2007-05-23 Neta A. Bahcall

A major aim of evolutionary biology is to explain the respective roles of adaptive versus non-adaptive changes in the evolution of complexity. While selection is certainly responsible for the spread and maintenance of complex phenotypes,…

种群与进化 · 定量生物学 2017-01-17 Thomas LaBar , Christoph Adami

Kernel density estimation is a convenient way to estimate the probability density of a distribution given the sample of data points. However, it has certain drawbacks: proper description of the density using narrow kernels needs large data…

数据分析、统计与概率 · 物理学 2015-02-27 Anton Poluektov