English
Related papers

Related papers: Evaluating Recabilities of Foundation Models: A Mu…

200 papers

AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attributes and for comparing model performance, tracking…

Artificial Intelligence · Computer Science 2024-11-21 Anka Reuel , Amelia Hardy , Chandler Smith , Max Lamparth , Malcolm Hardy , Mykel J. Kochenderfer

In today's digital landscape, Deep Recommender Systems (DRS) play a crucial role in navigating and customizing online content for individual preferences. However, conventional methods, which mainly depend on single recommendation task,…

Information Retrieval · Computer Science 2025-03-03 Xiangyu Zhao , Yichao Wang , Bo Chen , Jingtong Gao , Yuhao Wang , Xiaopeng Li , Pengyue Jia , Qidong Liu , Huifeng Guo , Ruiming Tang

For task-oriented dialog systems to be maximally useful, it must be able to process conversations in a way that is (1) generalizable with a small number of training examples for new task domains, and (2) robust to user input in various…

Computation and Language · Computer Science 2021-01-01 Baolin Peng , Chunyuan Li , Zhu Zhang , Chenguang Zhu , Jinchao Li , Jianfeng Gao

Deep research, in which an agent searches the open web, collects evidence, and derives an answer through extended reasoning, is a prominent use case for frontier language models. Frontier deep research products score high on existing…

Artificial Intelligence · Computer Science 2026-05-21 Sixiong Xie , Zhuofan Shi , Haiyang Shen , Jiuzheng Wang , Siqi Zhong , Mugeng Liu , Chongyang Pan , Peilun Jia , Baoqing Sun , Xiang Jing , Yun Ma

Rigorous and reproducible evaluation is critical for assessing the state of the art and for guiding scientific advances in Artificial Intelligence. Evaluation is challenging in practice due to several reasons, including benchmark…

In real-world vision-language applications, practitioners increasingly rely on large, pretrained foundation models rather than custom-built solutions, despite limited transparency regarding their training data and processes. While these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Mario Koddenbrock , Rudolf Hoffmann , David Brodmann , Erik Rodner

The impressive performance of ChatGPT and other foundation-model-based products in human language understanding has prompted both academia and industry to explore how these models can be tailored for specific industries and application…

Artificial Intelligence · Computer Science 2025-09-22 Haolong Chen , Hanzhi Chen , Zijian Zhao , Kaifeng Han , Guangxu Zhu , Yichen Zhao , Ying Du , Wei Xu , Qingjiang Shi

Multimodal foundation models have demonstrated impressive capabilities across diverse tasks. However, their potential as plug-and-play solutions for missing modality reconstruction remains underexplored. To bridge this gap, we identify and…

Multimedia · Computer Science 2026-05-25 Guanzhou Ke , Bo Wang , Guoqing Chao , Weiming Hu , Shengfeng He

Instruction-following is a foundational capability of large language models (LLMs), with its improvement hinging on scalable and accurate feedback from judge models. However, the reliability of current judge models in instruction-following…

Computation and Language · Computer Science 2026-04-17 Bosi Wen , Yilin Niu , Cunxiang Wang , Xiaoying Ling , Ying Zhang , Pei Ke , Hongning Wang , Minlie Huang

Frontier model progress is often measured by academic benchmarks, which offer a limited view of performance in real-world professional contexts. Existing evaluations often fail to assess open-ended, economically consequential tasks in…

Foundation models have demonstrated remarkable generalization, data efficiency, and robustness properties across various domains. In this paper, we explore the feasibility of foundation models for applications in the control domain. The…

Machine Learning · Computer Science 2024-12-18 Martin Ziegler , Andres Felipe Posada-Moreno , Friedrich Solowjow , Sebastian Trimpe

Multi-domain recommendation (MDR) aims to provide recommendations for different domains (e.g., types of products) with overlapping users/items and is common for platforms such as Amazon, Facebook, and LinkedIn that host multiple services.…

Information Retrieval · Computer Science 2023-08-15 Wentao Ning , Xiao Yan , Weiwen Liu , Reynold Cheng , Rui Zhang , Bo Tang

While the OneRec series has successfully unified the fragmented recommendation pipeline into an end-to-end generative framework, a significant gap remains between recommendation systems and general intelligence. Constrained by isolated…

In this work, we study the problem of learning a single model for multiple domains. Unlike the conventional machine learning scenario where each domain can have the corresponding model, multiple domains (i.e., applications/users) may share…

Machine Learning · Computer Science 2019-05-23 Qi Qian , Shenghuo Zhu , Jiasheng Tang , Rong Jin , Baigui Sun , Hao Li

We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates agents on multi-step…

Existing database benchmarks primarily focus on performance under ideal running environments. However, in real-world scenarios, databases probably face numerous adverse events. Quantifying the ability to cope with these events from a…

Databases · Computer Science 2025-11-17 Puyun Hu , Wei Pan , Xun Jian , Zeqi Ma , Tianjie Li , Yang Shen , Chengzhi Han , Yudong Zhao , Zhanhuai Li

Transferability scores aim to quantify how well a model trained on one domain generalizes to a target domain. Despite numerous methods proposed for measuring transferability, their reliability and practical usefulness remain inconclusive,…

Machine Learning · Computer Science 2025-04-30 Alireza Kazemi , Helia Rezvani , Mahsa Baktashmotlagh

Task load detection is essential for optimizing human performance across diverse applications, yet current models often lack generalizability beyond narrow experimental domains. While prior research has focused on individual tasks and…

Machine Learning · Computer Science 2025-09-03 Maximilian P. Oppelt , Andreas Foltyn , Nadine R. Lang-Richter , Bjoern M. Eskofier

Recommendations Systems allow users to identify trending items among a community while being timely and relevant to the user's expectations. When the purpose of various Recommendation Systems differs, the required type of recommendations…

Information Retrieval · Computer Science 2022-05-05 Dinuka Ravijaya Piyadigama , Guhanathan Poravi

We propose the Multimodal Clinical Benchmark for Emergency Care (MC-BEC), a comprehensive benchmark for evaluating foundation models in Emergency Medicine using a dataset of 100K+ continuously monitored Emergency Department visits from…

Machine Learning · Computer Science 2023-11-10 Emma Chen , Aman Kansal , Julie Chen , Boyang Tom Jin , Julia Rachel Reisler , David A Kim , Pranav Rajpurkar
‹ Prev 1 3 4 5 6 7 10 Next ›