中文
相关论文

相关论文: Benchmark It Yourself (BIY): Preparing a Dataset a…

200 篇论文

The Bhatt Conjectures framework introduces rigorous, hierarchical benchmarks for evaluating AI reasoning and understanding, moving beyond pattern matching to assess representation invariance, robustness, and metacognitive self-awareness.…

密码学与安全 · 计算机科学 2025-06-23 Manish Bhatt

Metacognition, the ability to monitor and regulate one's own reasoning, remains under-evaluated in AI benchmarking. We introduce MEDLEY-BENCH, a benchmark of behavioural metacognition that separates independent reasoning, private…

人工智能 · 计算机科学 2026-04-20 Farhad Abtahi , Abdolamir Karbalaie , Eduardo Illueca-Fernandez , Fernando Seoane

Conversation logs from AI platforms are increasingly used to measure occupational exposure to artificial intelligence, but the users observed in these logs are not the workforce. We show that platform-derived exposure scores combine…

人工智能 · 计算机科学 2026-05-28 Michelle Yin , Burhan Ogut

Programmers are turning to AI coding assistants to answer questions about their code. Benchmarks are needed to soundly evaluate these systems and understand their performance. To enable such a study, we curate a benchmark of real-world…

软件工程 · 计算机科学 2026-05-06 Ferida Mohammed , Fatma Ayad , Petros Maniatis , Satish Chandra , Elizabeth Dinella

Artificial intelligence (AI) tools are being incorporated into scientific research workflows with the potential to enhance efficiency in tasks such as document analysis, question answering (Q&A), and literature search. However, system…

人工智能 · 计算机科学 2026-05-13 Anthea Dathe , Kiran Hoffmann , Aline Mangold

In the rapidly evolving domain of Recommender Systems (RecSys), new algorithms frequently claim state-of-the-art performance based on evaluations over a limited set of arbitrarily selected datasets. However, this approach may fail to…

AI models have become extremely popular and accessible to the general public. However, they are continuously under the scanner due to their demonstrable biases toward various sections of the society like people of color and non-binary…

计算机与社会 · 计算机科学 2023-10-11 Siddharth D Jaiswal , Ankit Kumar Verma , Animesh Mukherjee

As demand drives systems to generalize to various domains and problems, the study of multitask, transfer and lifelong learning has become an increasingly important pursuit. In discrete domains, performance on the Atari game suite has…

人工智能 · 计算机科学 2017-08-16 Peter Henderson , Wei-Di Chang , Florian Shkurti , Johanna Hansen , David Meger , Gregory Dudek

Large language models (LLMs) have recently received considerable attention as alternative solutions for task planning. However, comparing the performance of language-oriented task planners becomes difficult, and there exists a dearth of…

人工智能 · 计算机科学 2024-02-14 Jae-Woo Choi , Youngwoo Yoon , Hyobin Ong , Jaehong Kim , Minsu Jang

The rising use of Artificial Intelligence (AI) in human detection on Edge camera systems has led to accurate but complex models, challenging to interpret and debug. Our research presents a diagnostic method using Explainable AI (XAI) for…

计算机视觉与模式识别 · 计算机科学 2024-01-19 Truong Thanh Hung Nguyen , Vo Thanh Khang Nguyen , Quoc Hung Cao , Van Binh Truong , Quoc Khanh Nguyen , Hung Cao

This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate whether frontier models sabotage safety research when…

人工智能 · 计算机科学 2026-04-02 Alexandra Souly , Robert Kirk , Jacob Merizian , Abby D'Cruz , Xander Davies

Memory systems for AI assistants were built for single-user dialogue and fail characteristically when applied to multi-party social group settings. This gap matters for the social assistants being built today: group-acting agents embedded…

计算与语言 · 计算机科学 2026-05-19 Olukunle Owolabi

Data-driven artificial intelligence models require explainability in intelligent manufacturing to streamline adoption and trust in modern industry. However, recently developed explainable artificial intelligence (XAI) techniques that…

机器学习 · 计算机科学 2025-02-04 Joseph Cohen , Xun Huan , Jun Ni

Current bias evaluation methods rarely engage with communities impacted by AI systems. Inspired by bug bounties, bias bounties have been proposed as a reward-based method that involves communities in AI bias detection by asking users of AI…

计算机与社会 · 计算机科学 2025-10-03 Sergej Kucenko , Nathaniel Dennler , Fengxiang He

The proliferation of the Internet of Things (IoT) and its cutting-edge AI-enabled applications (e.g., autonomous vehicles and smart industries) combine two paradigms: data-driven systems and their deployment on the edge. Usually, edge…

机器学习 · 计算机科学 2025-08-01 Ghazal Sobhani , Md. Monzurul Amin Ifath , Tushar Sharma , Israat Haque

Forecasting startup success is notoriously difficult, partly because meaningful outcomes, such as exits, large funding rounds, and sustained revenue growth, are rare and can take years to materialize. As a result, signals are sparse and…

机器学习 · 计算机科学 2026-04-08 Mostapha Benhenda

AI-generated faces have enriched human life, such as entertainment, education, and art. However, they also pose misuse risks. Therefore, detecting AI-generated faces becomes crucial, yet current detectors show biased performance across…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Li Lin , Santosh , Mingyang Wu , Xin Wang , Shu Hu

Motivated by the goals of dataset pruning and defect identification, a growing body of methods have been developed to score individual examples within a dataset. These methods, which we call "example difficulty scores", are typically used…

机器学习 · 计算机科学 2024-01-04 Devin Kwok , Nikhil Anand , Jonathan Frankle , Gintare Karolina Dziugaite , David Rolnick

Machine learning models that learn from dynamic graphs face nontrivial challenges in learning and inference as both nodes and edges change over time. The existing large-scale graph benchmark datasets that are widely used by the community…