English
Related papers

Related papers: Heterogeneous Judge-Aware Ranking with Sensitivity…

200 papers

Massive amounts of data are the foundation of data-driven recommendation models. As an inherent nature of big data, data heterogeneity widely exists in real-world recommendation systems. It reflects the differences in the properties among…

Information Retrieval · Computer Science 2023-05-26 Zimu Wang , Jiashuo Liu , Hao Zou , Xingxuan Zhang , Yue He , Dongxu Liang , Peng Cui

Decentralized large language model inference networks require lightweight mechanisms to reward high quality outputs under heterogeneous latency and cost. Proof of Quality provides scalable verification by sampling evaluator nodes that score…

Cryptography and Security · Computer Science 2026-01-30 Arther Tian , Alex Ding , Frank Chen , Simon Wu , Aaron Chan

In this paper, we define an underlying data generating process that allows for different magnitudes of cross-sectional dependence, along with time series autocorrelation. This is achieved via high-dimensional moving average processes of…

Econometrics · Economics 2025-07-22 Jiti Gao , Fei Liu , Bin Peng , Yayi Yan

The assessment of explainability in Legal Judgement Prediction (LJP) systems is of paramount importance in building trustworthy and transparent systems, particularly considering the reliance of these systems on factors that may lack legal…

Computation and Language · Computer Science 2024-02-28 Santosh T. Y. S. S , Nina Baumgartner , Matthias Stürmer , Matthias Grabmair , Joel Niklaus

Metaheuristic algorithms are essential for solving complex optimization problems in different fields. However, the difficulty in comparing and rating these algorithms remains due to the wide range of performance metrics and problem…

Neural and Evolutionary Computing · Computer Science 2024-11-28 Evgenia-Maria K. Goula , Dimitris G. Sotiropoulos

Recently, real-world recommendation systems need to deal with millions of candidates. It is extremely challenging to conduct sophisticated end-to-end algorithms on the entire corpus due to the tremendous computation costs. Therefore,…

Information Retrieval · Computer Science 2021-10-15 Ruobing Xie , Qi Liu , Shukai Liu , Ziwei Zhang , Peng Cui , Bo Zhang , Leyu Lin

Trustworthy multi-view classification (TMVC) addresses the challenge of achieving reliable decision-making in complex scenarios where multi-source information is heterogeneous, inconsistent, or even conflicting. Existing TMVC approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Haojian Huang , Jiahao Shi , Zhe Liu , Harold Haodong Chen , Han Fang , Hao Sun , Zhongjiang He

Understanding causal heterogeneity is essential for scientific discovery in domains such as biology and medicine. However, existing methods lack causal awareness, with insufficient modeling of heterogeneity, confounding, and observational…

Machine Learning · Computer Science 2025-10-29 Wenrui Li , Qinghao Zhang , Xiaowo Wang

The "LLM-as-a-Judge" paradigm, using Large Language Models (LLMs) as automated evaluators, is pivotal to LLM development, offering scalable feedback for complex tasks. However, the reliability of these judges is compromised by various…

Computation and Language · Computer Science 2026-05-22 Qingquan Li , Shaoyu Dou , Kailai Shao , Chao Chen , Haixiang Hu

Rigorous and reproducible evaluation is critical for assessing the state of the art and for guiding scientific advances in Artificial Intelligence. Evaluation is challenging in practice due to several reasons, including benchmark…

Multiple defendants in a criminal fact description generally exhibit complex interactions, and cannot be well handled by existing Legal Judgment Prediction (LJP) methods which focus on predicting judgment results (e.g., law articles,…

Computation and Language · Computer Science 2023-12-12 Yougang Lyu , Jitai Hao , Zihan Wang , Kai Zhao , Shen Gao , Pengjie Ren , Zhumin Chen , Fang Wang , Zhaochun Ren

Advances in object recognition flourished in part because of the availability of high-quality datasets and associated benchmarks. However, these benchmarks---such as ILSVRC---are relatively task-specific, focusing predominately on…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Brett D. Roads , Bradley C. Love

Large language models (LLMs) are evolving fast and are now frequently used as evaluators, in a process typically referred to as LLM-as-a-Judge, which provides quality assessments of model outputs. However, recent research points out…

Computation and Language · Computer Science 2026-01-27 Hugo Silva , Mateus Mendes , Hugo Gonçalo Oliveira

Preference mechanisms, such as human preference, LLM-as-a-Judge (LaaJ), and reward models, are central to aligning and evaluating large language models (LLMs). Yet, the underlying concepts that drive these preferences remain poorly…

Computation and Language · Computer Science 2025-05-30 Nitay Calderon , Liat Ein-Dor , Roi Reichart

LLM-as-Judge has emerged as a scalable alternative to human evaluation, enabling large language models (LLMs) to provide reward signals in trainings. While recent work has explored multi-agent extensions such as multi-agent debate and…

Artificial Intelligence · Computer Science 2025-09-19 Chiyu Ma , Enpei Zhang , Yilun Zhao , Wenjun Liu , Yaning Jia , Peijun Qing , Lin Shi , Arman Cohan , Yujun Yan , Soroush Vosoughi

Aligning structured data is a fundamental problem in computer vision and machine learning, underlying tasks such as time series analysis, human action recognition, and visual representation learning. Existing alignment methods, including…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Lei Wang , Syuan-Hao Li , Yongsheng Gao , Piotr Koniusz

Large multimodal models (LMMs) are increasingly adopted as judges in multimodal evaluation systems due to their strong instruction following and consistency with human preferences. However, their ability to follow diverse, fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Tianyi Xiong , Yi Ge , Ming Li , Zuolong Zhang , Pranav Kulkarni , Kaishen Wang , Qi He , Zeying Zhu , Chenxi Liu , Ruibo Chen , Tong Zheng , Yanshuo Chen , Xiyao Wang , Renrui Zhang , Wenhu Chen , Heng Huang

When groups of people are tasked with making a judgment, the issue of uncertainty often arises. Existing methods to reduce uncertainty typically focus on iteratively improving specificity in the overall task instruction. However,…

Human-Computer Interaction · Computer Science 2023-10-10 Quan Ze Chen , Amy X. Zhang

Rankings and scores are two common data types used by judges to express preferences and/or perceptions of quality in a collection of objects. Numerous models exist to study data of each type separately, but no unified statistical model…

Methodology · Statistics 2022-09-02 Michael Pearce , Elena A. Erosheva

Classical latent-score ranking models often fail to distinguish objects' intrinsic scores from contextual effects, which are typically nonlinear and can dominate the observed outcomes. To address this, we introduce a semiparametric ranking…

Methodology · Statistics 2026-04-22 Yuanhang Luo , Shuxing Fang , Ruijian Han , Yiming Xu