Related papers: Benchmark for two-dimensional large scale coherent…
We study an ultrarelativistic QED plasma in thermal equilibrium. Plasmons - photon collective excitations - are postulated to correspond not to poles of the retarded photon propagator but to poles of the propagator multiplied by the fine…
The Web is a rich source of structured data in the form of tables, from product catalogs and knowledge bases to scientific datasets. However, the heterogeneity of the structure and semantics of these tables makes it challenging to build a…
Large Language Models (LLMs) promise to streamline software code reviews, but their ability to produce consistent assessments remains an open question. In this study, we tested four leading LLMs -- GPT-4o mini, GPT-4o, Claude 3.5 Sonnet,…
We present Task Bench, a parameterized benchmark designed to explore the performance of parallel and distributed programming systems under a variety of application scenarios. Task Bench lowers the barrier to benchmarking multiple…
Long wavelength plasma non-uniformities rotating in the azimuthal direction ("rotating spokes") have been observed in a number of experiments on Hall thrusters or magnetron discharges. We use a two-dimensional (2D), axial-azimuthal…
Randomized benchmarking has emerged as a popular and easy-to-implement experimental technique for gauging the quality of gate operations in quantum computing devices. A typical randomized benchmarking procedure identifies the exponential…
Molecular dynamics simulations have emerged as a fundamental instrument for studying biomolecules. At the same time, it is desirable to perform simulations of a collection of particles under various conditions in which the molecules can…
A fundamental task in particle-in-cell (PIC) simulations of plasma physics is solving for charged particle motion in electromagnetic fields. This problem is especially challenging when the plasma is strongly magnetized due to numerical…
In space and astrophysical plasmas turbulence leads to the development of coherent structures characterized by a strong current density and important magnetic shears. Using hybrid-kinetic simulations of turbulence (3D with different energy…
Machine learning-based modeling of physical systems has experienced increased interest in recent years. Despite some impressive progress, there is still a lack of benchmarks for Scientific ML that are easy to use but still challenging and…
Symbolic execution is a powerful technique for bug finding and program testing. It is successful in finding bugs in real-world code. The core reasoning techniques use constraint solving, path exploration, and search, which are also the same…
Synthetic turbulence is a relevant tool to study complex astrophysical and space plasma environments inaccessible by direct simulation. However, conventional models lack intermittent coherent structures, which are essential in realistic…
A key hurdle is demonstrating compute resource capability with limited benchmarks. We propose workflow templates as a solution, offering adaptable designs for specific scientific applications. Our paper identifies common usage patterns for…
The concept of scalability analysis of numerical parallel applications has been revisited, with the specific goals defined for the performance estimation of research applications. A series of Community Climate Model System (CCSM) numerical…
Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant tasks. Existing suites have either saturated, heavily depend…
As quantum devices scale up, many-body quantum gates and algorithms begin to surpass what is possible to simulate classically. Validation methods which rely on such classical simulation, such as process tomography and randomized…
Understanding magnetized laser-plasma interactions is important for controlling magneto-inertial fusion experiments and developing magnetically assisted radiation and particle sources. In the long-pulse regime, interactions are dominated by…
Applying post selection in each step of an iterated protocol leads to sensitive quantum dynamics that may be utilized to test and benchmark current quantum computers. An example of this type of protocols was originally proposed for the task…
With hundreds of multilingual embedding models available, practitioners lack clear guidance on which provide genuine cross-lingual semantic alignment versus task performance through language-specific patterns. Task-driven benchmarks (MTEB)…
Large language models often achieve strong benchmark gains without corresponding improvements in broader capability. We hypothesize that this discrepancy arises from differences in training regimes induced by data distribution. To…