English
Related papers

Related papers: The Mock LISA Data Challenges: An overview

200 papers

In some areas of computing, natural language processing and information science, progress is made by sharing datasets and challenging the community to design the best algorithm for an associated task. This article introduces a shared…

Digital Libraries · Computer Science 2026-01-27 Mike Thelwall

As the largest radio telescope in the world, the Square Kilometre Array (SKA) will lead the next generation of radio astronomy. The feats of engineering required to construct the telescope array will be matched only by the techniques…

As observational datasets become larger and more complex, so too are the questions being asked of these data. Data simulations, i.e., synthetic data with properties (pixelization, noise, PSF, artifacts, etc.) akin to real data, are…

Instrumentation and Methods for Astrophysics · Physics 2019-10-25 Molly S. Peeples , Bjorn Emonts , Mark Kyprianou , Matthew T. Penny , Gregory F. Snyder , Christopher C. Stark , Michael Troxel , Neil T. Zimmerman , John ZuHone

Considering electronic implications in the Information Society (IS) as a complex system, complexity science tools are used to describe the processes that are seen to be taking place. The sometimes troublesome relationship between the…

Physics and Society · Physics 2015-03-19 Noemi L. Olivera , Araceli N. Proto , Marcel Ausloos

Many software engineering research papers rely on time-based data (e.g., commit timestamps, issue report creation/update/close dates, release dates). Like most real-world data however, time-based data is often dirty. To date, there are no…

Software Engineering · Computer Science 2021-03-23 Samuel W. Flint , Jigyasa Chauhan , Robert Dyer

We investigate five English NLP benchmark datasets (on the superGLUE leaderboard) and two Swedish datasets for bias, along multiple axes. The datasets are the following: Boolean Question (Boolq), CommitmentBank (CB), Winograd Schema…

Computation and Language · Computer Science 2023-09-19 Tosin Adewumi , Isabella Södergren , Lama Alkhaled , Sana Sabah Sabry , Foteini Liwicki , Marcus Liwicki

In recent years, many studies have applied wearable devices to capture psychophysiological data from software developers. However, the current literature lacks investigations that classify the studies and point out gaps to be explored. This…

Software Engineering · Computer Science 2021-06-01 Roger Vieira , Kleinner Farias

We introduce LongDA, a data analysis benchmark for evaluating LLM-based agents under documentation-intensive analytical workflows. In contrast to existing benchmarks that assume well-specified schemas and inputs, LongDA targets real-world…

Digital Libraries · Computer Science 2026-01-13 Yiyang Li , Zheyuan Zhang , Tianyi Ma , Zehong Wang , Keerthiram Murugesan , Chuxu Zhang , Yanfang Ye

Leading large language models (LLMs) are trained on public data. However, most of the world's data is dark data that is not publicly accessible, mainly in the form of private organizational or enterprise data. We show that the performance…

Databases · Computer Science 2024-12-31 Moe Kayali , Fabian Wenz , Nesime Tatbul , Çağatay Demiralp

Many practical studies rely on hypothesis testing procedures applied to data sets with missing information. An important part of the analysis is to determine the impact of the missing data on the performance of the test, and this can be…

Methodology · Statistics 2011-02-15 Dan L. Nicolae , Xiao-Li Meng , Augustine Kong

Research papers in the biomedical field come with large and complex data sets that are shared with the scientific community as unstructured data files via public data repositories. Examples are sequencing, microarray, and mass spectroscopy…

Quantitative Methods · Quantitative Biology 2020-05-28 Michael Huttner , Claudio Lottaz , Christian Kohler , Rainer Spang

We study to what extend LISA can observe features of gravitational wave spectra originating from cosmological first-order phase transitions. We focus on spectra which are of the form of double-broken power laws. These spectra are predicted…

Cosmology and Nongalactic Astrophysics · Physics 2021-12-21 Felix Giese , Thomas Konstandin , Jorinde van de Vis

Application of models to data is fraught. Data-generating collaborators often only have a very basic understanding of the complications of collating, processing and curating data. Challenges include: poor data collection practices, missing…

Databases · Computer Science 2017-05-08 Neil D. Lawrence

Bias in AI systems, especially those relying on natural language data, raises ethical and practical concerns. Underrepresentation of certain groups often leads to uneven performance across demographics. Traditional fairness methods, such as…

Computation and Language · Computer Science 2025-10-16 Sai Suhruth Reddy Karri , Yashwanth Sai Nallapuneni , Laxmi Narasimha Reddy Mallireddy , Gopichand G

The Laser Interferometer Space Antenna (LISA) mission, scheduled for launch in the mid-2030s, is a gravitational wave observatory in space designed to detect sources emitting in the millihertz band. LISA is an ESA flagship mission,…

Simulations in information access (IA) have recently gained interest, as shown by various tutorials and workshops around that topic. Simulations can be key contributors to central IA research and evaluation questions, especially around…

Information Retrieval · Computer Science 2025-05-20 Philipp Schaer , Christin Katharina Kreutz , Krisztian Balog , Timo Breuer , Andreas Konstantin Kruff

The selection, development, or comparison of machine learning methods in data mining can be a difficult task based on the target problem and goals of a particular study. Numerous publicly available real-world and simulated benchmark…

Machine Learning · Computer Science 2017-03-03 Randal S. Olson , William La Cava , Patryk Orzechowski , Ryan J. Urbanowicz , Jason H. Moore

The proliferation of LLM bias probes introduces three significant challenges: (1) we lack principled criteria for choosing appropriate probes, (2) we lack a system for reconciling conflicting results across probes, and (3) we lack formal…

Computers and Society · Computer Science 2025-03-04 Kirsten N. Morehouse , Siddharth Swaroop , Weiwei Pan

Textual data used to train large language models (LLMs) exhibits multifaceted bias manifestations encompassing harmful language and skewed demographic distributions. Regulations such as the European AI Act require identifying and mitigating…

‹ Prev 1 4 5 6 7 8 10 Next ›