中文
相关论文

相关论文: Performance evaluation through DEA benchmarking ad…

200 篇论文

Evaluating models on large benchmarks is very resource-intensive, especially during the period of rapid model evolution. Existing efficient evaluation methods estimate the performance of target models by testing them only on a small and…

机器学习 · 计算机科学 2025-06-03 Peiwen Yuan , Yueqi Zhang , Shaoxiong Feng , Yiwei Li , Xinglin Wang , Jiayi Shi , Chuyi Tan , Boyuan Pan , Yao Hu , Kan Li

Machine learning (ML) is increasingly applied to optimize system performance in tasks such as resource management and network simulation. Unlike traditional ML tasks (e.g., image classification), networked systems often operate in…

机器学习 · 计算机科学 2026-05-15 Daiyang Yu , Xinyu Chen , Yihan Zhang , Yan Liang , Yaqi Qiao , Fan Lai

During the past sixty years, a lot of effort has been made regarding the productive efficiency. Such endeavours provided an extensive bibliography on this subject, culminating in two main methods, named the Stochastic Frontier Analysis…

最优化与控制 · 数学 2019-08-14 Anibal Galindro , Micael Santos , Delfim F. M. Torres , Ana Marta-Costa

Portfolio managers are typically constrained by turnover limits, minimum and maximum stock positions, cardinality, a target market capitalization and sometimes the need to hew to a style (such as growth or value). In addition, portfolio…

投资组合管理 · 定量金融 2012-01-04 Andrew Clark , Jeff Kenyon

In data envelopment model (DEA), while either the most or least distance based frameworks can be implemented for targeting, the latter is often more relevant than the former from a managerial point of view due to easy attainability of the…

最优化与控制 · 数学 2015-04-01 Mahmood Mehdiloozad , Mohammad Bagher Ahmadi

Research and development (R&D) of countries play a major role in a long-term development of the economy. We measure the R&D efficiency of all 28 member countries of the European Union in the years 2008--2014. Super-efficient data…

应用统计 · 统计学 2024-05-09 Vladimír Holý , Karel Šafr

The conventional evaluation protocols on machine learning models rely heavily on a labeled, i.i.d-assumed testing dataset, which is not often present in real world applications. The Automated Model Evaluation (AutoEval) shows an alternative…

机器学习 · 计算机科学 2024-03-18 Ru Peng , Heming Zou , Haobo Wang , Yawen Zeng , Zenan Huang , Junbo Zhao

Dynamic benchmarks interweave model fitting and data collection in an attempt to mitigate the limitations of static benchmarks. In contrast to an extensive theoretical and empirical study of the static setting, the dynamic counterpart lags…

机器学习 · 计算机科学 2023-03-03 Ali Shirali , Rediet Abebe , Moritz Hardt

We propose a method for test-time adaptation of pretrained depth completion models. Depth completion models, trained on some ``source'' data, often predict erroneous outputs when transferred to ``target'' data captured in novel…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Younjoon Chung , Hyoungseob Park , Patrick Rim , Xiaoran Zhang , Jihe He , Ziyao Zeng , Safa Cicek , Byung-Woo Hong , James S. Duncan , Alex Wong

Mixture-of-Experts (MoE) architectures have emerged as a promising direction, offering efficiency and scalability by activating only a subset of parameters during inference. However, current research remains largely performance-centric,…

机器学习 · 计算机科学 2025-09-30 Jiahao Ying , Mingbao Lin , Qianru Sun , Yixin Cao

High Performance Distributed Computing is essential to boost scientific progress in many areas of science and to efficiently deploy a number of complex scientific applications. These applications have different characteristics that require…

分布式、并行与集群计算 · 计算机科学 2014-12-04 Mariza Ferro , Antonio R. Mury , Laion F. Manfroi , Bruno Schlze

Large language models are increasingly deployed with test-time strategies: sample $N$ responses, score them with a reward model or verifier, and return the best. This deployment rule exposes a mismatch in post-training: standard objectives…

机器学习 · 计算机科学 2026-05-12 Muheng Li , Jian Qian , Wenlong Mou

Many unsupervised domain adaptation (UDA) methods have been proposed to bridge the domain gap by utilizing domain invariant information. Most approaches have chosen depth as such information and achieved remarkable success. Despite their…

计算机视觉与模式识别 · 计算机科学 2022-11-17 Ting-Hsuan Liao , Huang-Ru Liao , Shan-Ya Yang , Jie-En Yao , Li-Yuan Tsao , Hsu-Shen Liu , Bo-Wun Cheng , Chen-Hao Chao , Chia-Che Chang , Yi-Chen Lo , Chun-Yi Lee

We present a new tool, GPA, that can generate key performance measures for very large systems. Based on solving systems of ordinary differential equations (ODEs), this method of performance analysis is far more scalable than stochastic…

性能 · 计算机科学 2010-06-29 Anton Stefanek , Richard Hayden , Jeremy Bradley

Models that top leaderboards often perform unsatisfactorily when deployed in real world applications; this has necessitated rigorous and expensive pre-deployment model testing. A hitherto unexplored facet of model performance is: Are our…

计算与语言 · 计算机科学 2021-06-11 Swaroop Mishra , Anjana Arunkumar

Software Process Improvement (SPI) encompasses the analysis and modification of the processes within software development, aimed at improving key areas that contribute to the organizations' goals. The task of evaluating whether the selected…

When we are primarily interested in solving several problems jointly with a given prescribed high performance accuracy for each target application, then Foundation Models should for most cases be used rather than problem-specific models. We…

计算机视觉与模式识别 · 计算机科学 2024-06-27 Nikolaos Dionelis , Casper Fibaek , Luke Camilleri , Andreas Luyts , Jente Bosmans , Bertrand Le Saux

Comparing model performances on benchmark datasets is an integral part of measuring and driving progress in artificial intelligence. A model's performance on a benchmark dataset is commonly assessed based on a single or a small set of…

人工智能 · 计算机科学 2021-11-09 Kathrin Blagec , Georg Dorffner , Milad Moradi , Matthias Samwald

Every data selection method inherently has a target. In practice, these targets often emerge implicitly through benchmark-driven iteration: researchers develop selection strategies, train models, measure benchmark performance, then refine…

Recently, the Deep Learning community has become interested in evolutionary optimization (EO) as a means to address hard optimization problems, e.g. meta-learning through long inner loop unrolls or optimizing non-differentiable operators.…

神经与进化计算 · 计算机科学 2023-11-07 Robert Tjarko Lange , Yujin Tang , Yingtao Tian