中文
相关论文

相关论文: GAIA: a benchmark for General AI Assistants

200 篇论文

The advancement of large language models (LLMs) has significantly accelerated the development of search agents capable of autonomously gathering information through multi-turn web interactions. Various benchmarks have been proposed to…

Organizations are increasingly exploring delegation of screening and negotiation tasks to AI systems, yet deployment in high-stakes B2B settings is constrained by governance: preventing unauthorized commitments, ensuring sufficient…

人工智能 · 计算机科学 2025-11-11 Siming Zhao , Qi Li

Human Intelligence (HI) excels at combining basic skills to solve complex tasks. This capability is vital for Artificial Intelligence (AI) and should be embedded in comprehensive AI Agents, enabling them to harness expert models for complex…

人工智能 · 计算机科学 2023-11-06 Yingqiang Ge , Wenyue Hua , Kai Mei , Jianchao Ji , Juntao Tan , Shuyuan Xu , Zelong Li , Yongfeng Zhang

As machine intelligence evolves, the need to test and compare the problem-solving abilities of different AI models grows. However, current benchmarks are often simplistic, allowing models to perform uniformly well and making it difficult to…

The evolution of artificial intelligence (AI) has profoundly impacted human society, driving significant advancements in multiple sectors. AGI, distinguished by its ability to execute diverse real-world tasks with efficiency and…

人工智能 · 计算机科学 2024-11-26 Tao Feng , Chuanyang Jin , Jingyu Liu , Kunlun Zhu , Haoqin Tu , Zirui Cheng , Guanyu Lin , Jiaxuan You

Physical reasoning is a crucial aspect in the development of general AI systems, given that human learning starts with interacting with the physical world before progressing to more complex concepts. Although researchers have studied and…

人工智能 · 计算机科学 2023-12-19 Andrew Melnik , Robin Schiewer , Moritz Lange , Andrei Muresanu , Mozhgan Saeidi , Animesh Garg , Helge Ritter

The goal of achieving Artificial General Intelligence (AGI) is to imitate humans and surpass them. Models such as OpenAI's o1, o3, and DeepSeek's R1 have demonstrated that large language models (LLMs) with human-like reasoning capabilities…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Yansheng Qiu , Li Xiao , Zhaopan Xu , Pengfei Zhou , Zheng Wang , Kaipeng Zhang

Data-driven operations management often relies on parameters estimated from costly human-generated labels. Recent advances in large language models (LLMs) and other AI systems offer inexpensive auxiliary data, but introduce a new challenge:…

机器学习 · 计算机科学 2026-04-17 Cheng Lu , Mengxin Wang , Dennis J. Zhang , Heng Zhang

Human cognitive biases in software engineering can lead to costly errors. While general-purpose AI (GPAI) systems may help mitigate these biases due to their non-human nature, their training on human-generated data raises a critical…

Significant focus has been placed on integrating large language models (LLMs) with various tools in developing general-purpose agents. This poses a challenge to LLMs' tool-use capabilities. However, there are evident gaps between existing…

计算与语言 · 计算机科学 2024-11-25 Jize Wang , Zerun Ma , Yining Li , Songyang Zhang , Cailian Chen , Kai Chen , Xinyi Le

We introduce WebGames, a comprehensive benchmark suite designed to evaluate general-purpose web-browsing AI agents through a collection of 50+ interactive challenges. These challenges are specifically crafted to be straightforward for…

We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing…

Privacy policies of websites are often lengthy and intricate. Privacy assistants assist in simplifying policies and making them more accessible and user friendly. The emergence of generative AI (genAI) offers new opportunities to build…

密码学与安全 · 计算机科学 2023-12-20 Aamir Hamid , Hemanth Reddy Samidi , Tim Finin , Primal Pappachan , Roberto Yus

The booming development of AI agents presents unprecedented opportunities for automating complex tasks across various domains. However, their multi-step, multi-tool collaboration capabilities in the financial sector remain underexplored.…

OpenAI's o3 achieves a high score of 87.5 % on ARC-AGI, a benchmark proposed to measure intelligence. This raises the question whether systems based on Large Language Models (LLMs), particularly o3, demonstrate intelligence and progress…

人工智能 · 计算机科学 2025-01-14 Rolf Pfister , Hansueli Jud

Artificial General Intelligence (AGI) has been a long-standing goal of humanity, with the aim of creating machines capable of performing any intellectual task that humans can do. To achieve this, AGI researchers draw inspiration from the…

Artificial intelligence (AI) researchers have been developing and refining large language models (LLMs) that exhibit remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. The…

Numerous studies have assessed the proficiency of AI systems, particularly large language models (LLMs), in facilitating everyday tasks such as email writing, question answering, and creative content generation. However, researchers face…

This demo will present the Research Assistant (RA) tool developed to assist with six main types of research tasks defined as standardized instruction templates, instantiated with user input, applied finally as prompts to well-known--for…

计算与语言 · 计算机科学 2024-05-24 Mahsa Shamsabadi , Jennifer D'Souza

As general-purpose artificial intelligence systems become increasingly integrated into society and are used for information seeking, content generation, problem solving, textual analysis, coding, and running processes, it is crucial to…

计算机与社会 · 计算机科学 2025-08-28 Ljubisa Bojic , Dylan Seychell , Milan Cabarkapa
‹ 上一页 1 2 3 10 下一页 ›