English
Related papers

Related papers: The Evaluation Gap in Astronomy -- Explained throu…

200 papers

Score reliability is necessary for establishing a validity argument for an instrument, and is therefore highly important to investigate. Depending on the proposed instrument use and score interpretations, differing degrees of precision in…

Physics Education · Physics 2017-02-23 Robert M. Talbot

There is an evident and rapid trend towards the adoption of evaluation exercises for national research systems for purposes, among others, of improving allocative efficiency in public funding of individual institutions. However the desired…

Digital Libraries · Computer Science 2018-11-06 Giovanni Abramo , Ciriaco Andrea D'Angelo

Tasks that require information about the world imply a trade-off between the time spent on observation and the variance of the response. In particular, fast decisions need to rely on uncertain information. However, standard estimates of…

Neurons and Cognition · Quantitative Biology 2023-07-18 Sahel Azizpour , Viola Priesemann , Johannes Zierenberg , Anna Levina

Statistical inference in observational science typically relies on a fundamental assumption: as sample size increases and uncertainties decrease, the inferred results should converge to the true physical quantities. This assumption…

Machine Learning · Computer Science 2026-05-01 Zhipeng Zhang

Evaluation is a crucial aspect of human existence and plays a vital role in various fields. However, it is often approached in an empirical and ad-hoc manner, lacking consensus on universal concepts, terminologies, theories, and…

Human-Computer Interaction · Computer Science 2024-04-02 Jianfeng Zhan , Lei Wang , Wanling Gao , Hongxiao Li , Chenxi Wang , Yunyou Huang , Yatao Li , Zhengxin Yang , Guoxin Kang , Chunjie Luo , Hainan Ye , Shaopeng Dai , Zhifei Zhang

Spatial embodied intelligence requires agents to act to acquire information under partial observability. While multimodal foundation models excel at passive perception, their capacity for active, self-directed exploration remains…

This article presents the first systematic comparative survey of how public bodies, international organisations, national regulators, and the private sector define agentic artificial intelligence, identifying the technical inaccuracies…

Computers and Society · Computer Science 2026-03-31 Marcel Osmond , Thomas Jego

Citations are essential for recognizing scientific contributions, yet citation behavior is shaped by more than just relevance or quality. We analyzed approximately 255,000 refereed astronomy articles published between 2000 and 2025 to…

Digital Libraries · Computer Science 2026-03-27 Vardan Adibekyan , Olivier Demangeon , Tiago Campante , Nuno Santos , Susana Barros , Artur Hakobyan

Citations in science are being studied from several perspectives, among which approaches such as scientometrics and science of science. In this chapter I briefly review some of the literature on citations, citation distributions and models…

Digital Libraries · Computer Science 2025-05-12 V. A. Traag

The demand for global university league tables has been high over the past two decades. However, significant criticism of their methodologies is accumulating without being addressed. I revisit global university league tables by normalizing…

Physics and Society · Physics 2023-08-22 Saulo Mendes

Considerable efforts to measure and mitigate gender bias in recent years have led to the introduction of an abundance of tasks, datasets, and metrics used in this vein. In this position paper, we assess the current paradigm of gender bias…

Computation and Language · Computer Science 2022-10-21 Hadas Orgad , Yonatan Belinkov

A well-balanced exploration-exploitation trade-off is crucial for successful acquisition functions in Bayesian optimization. However, there is a lack of quantitative measures for exploration, making it difficult to analyze and compare…

Machine Learning · Computer Science 2026-05-15 Leonard Papenmeier , Nuojin Cheng , Stephen Becker , Luigi Nardi

Cultural AI benchmarks often rely on implicit assumptions about measured constructs, leading to vague formulations with poor validity and unclear interrelations. We propose exposing these assumptions using explicit cognitive models…

Artificial Intelligence · Computer Science 2024-09-26 Jonathan H. Rystrøm , Kenneth C. Enevoldsen

There is growing interest in leveraging LLMs to aid in astronomy and other scientific research, but benchmarks for LLM evaluation in general have not kept pace with the increasingly diverse ways that real people evaluate and use these…

Computation and Language · Computer Science 2025-08-07 Alina Hyk , Kiera McCormick , Mian Zhong , Ioana Ciucă , Sanjib Sharma , John F Wu , J. E. G. Peek , Kartheik G. Iyer , Ziang Xiao , Anjalie Field

Multi-agent coordination dilemmas expose a fundamental tension between individual optimization and collective welfare, yet characterizing such coordination requires metrics sensitive to temporal structure and collective dynamics. As a…

Multiagent Systems · Computer Science 2026-03-24 Nikolaos Al. Papadopoulos , Konstantinos Psannis

How should researchers analyze randomized experiments in which the main outcome is latent and measured in multiple ways but each measure contains some degree of error? We first identify a critical study-specific noncomparability problem in…

Econometrics · Economics 2026-01-13 Jiawei Fu , Donald P. Green

This resource letter provides a guide to research-based assessment instruments (RBAIs) of physics and astronomy content. These are standardized assessments that were rigorously developed and revised using student ideas and interviews,…

Physics Education · Physics 2017-04-05 Adrian Madsen , Sam McKagan , Eleanor C Sayre

Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not because of poorly designed benchmarks, but from implicit…

Computers and Society · Computer Science 2026-05-22 Naveen Raman , Santiago Cortes-Gomez , Mateo Dulce Rubio , Fei Fang , Bryan Wilder

The problems caused by the gap between system- and software-level architecting practices, especially in the context of Systems of Systems where the two disciplines inexorably meet, is a well known issue with a disappointingly low amount of…

Software Engineering · Computer Science 2021-02-10 Héctor Cadavid , Vasilios Andrikopoulos , Paris Avgeriou , P. Chris Broekema

(Un)conscious bias affects every aspect of the astronomical profession, from scientific activities (e.g., invitations to join collaborations, proposal selections, grant allocations, publication review processes, and invitations to attend…

Instrumentation and Methods for Astrophysics · Physics 2022-07-13 Alessandra Aloisi , Neill Reid
‹ Prev 1 4 5 6 7 8 10 Next ›