中文
相关论文

相关论文: Multi-Agent Reasoning with Consistency Verificatio…

200 篇论文

In biomedical engineering, artificial intelligence has become a pivotal tool for enhancing medical diagnostics, particularly in medical image classification tasks such as detecting pneumonia from chest X-rays and breast cancer screening.…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Jingsong Xia , Siqi Wang

Retrieval-augmented generation (RAG) systems offer a promising approach to reduce hallucinations and improve answer accuracy in large language models (LLMs), a requirement for reliable, financial analysis where answers must be grounded in…

机器学习 · 计算机科学 2026-05-26 Magnus Samuelsen , Wilmer Nyström , Somnath Mazumdar , Mansoor Hussain , Mikkel Strange

For safety, medical AI systems undergo thorough evaluations before deployment, validating their predictions against a ground truth which is assumed to be fixed and certain. However, this ground truth is often curated in the form of…

Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, existing medical visual question answering (VQA) benchmarks collapse model capabilities…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Yixiong Chen , Wenjie Xiao , Pedro R. A. S. Bassi , Boyan Wang , Liang He , Xinze Zhou , Sezgin Er , Ibrahim Ethem Hamamci , Zongwei Zhou , Alan Yuille

Prevailing medical AI operates on an unrealistic ''one-shot'' model, diagnosing from a complete patient file. However, real-world diagnosis is an iterative inquiry where Clinicians sequentially ask questions and order tests to strategically…

人工智能 · 计算机科学 2026-02-02 Yufei He , Juncheng Liu , Zhiyuan Hu , Yulin Chen , Yue Liu , Yuan Sui , Yibo Li , Nuo Chen , Jun Hu , Bryan Hooi , Xinxing Xu , Jiang Bian

A crucial challenge in decentralized systems is state estimation in the presence of unknown inputs, particularly within heterogeneous sensor networks with dynamic topologies. While numerous consensus algorithms have been introduced, they…

系统与控制 · 电气工程与系统科学 2024-12-13 Zida Wu , Ankur Mehta

We decompose physician disagreement in the HealthBench medical AI evaluation dataset to understand where variance resides and what observable features can explain it. Rubric identity accounts for 15.8% of met/not-met label variance but only…

人工智能 · 计算机科学 2026-03-10 Satya Borgohain , Roy Mariathas

Large language models can generate scientific simulation code, but the generated code silently fails on most non-textbook problems. We show that classical mathematical validation -- well-posedness, convergence, and error certification --…

软件工程 · 计算机科学 2026-03-30 Chengshuai Yang

Ensuring factual consistency and reliable reasoning remains a critical challenge for medical vision-language models. We introduce MEDFACT-R1, a two-stage framework that integrates external knowledge grounding with reinforcement learning to…

计算机视觉与模式识别 · 计算机科学 2025-09-19 Gengliang Li , Rongyu Chen , Bin Li , Linlin Yang , Guodong Ding

Although AI agents have demonstrated impressive capabilities in long-horizon reasoning, their reliability is severely hampered by the ``Spiral of Hallucination,'' where early epistemic errors propagate irreversibly. Existing methods face a…

人工智能 · 计算机科学 2026-01-23 Jiaxin Zhang , Prafulla Kumar Choubey , Kung-Hsiang Huang , Caiming Xiong , Chien-Sheng Wu

We present a novel approach for claim verification from tabular data documents. Recent LLM-based approaches either employ complex pretraining/fine-tuning or decompose verification into subtasks, often lacking comprehensive explanations and…

计算与语言 · 计算机科学 2026-04-21 Rudra Ranajee Saha , Laks V. S. Lakshmanan , Raymond T. Ng

Multi-Agent Reinforcement Learning (MARL) has gained significant interest in recent years, enabling sequential decision-making across multiple agents in various domains. However, most existing explanation methods focus on centralized MARL,…

人工智能 · 计算机科学 2025-11-14 Kayla Boggess , Sarit Kraus , Lu Feng

We present a visual symptom checker that combines a pre-trained Convolutional Neural Network (CNN) with a Reinforcement Learning (RL) agent as a Question Answering (QA) model. This method increases the classification confidence and accuracy…

人工智能 · 计算机科学 2019-08-09 Mohamed Akrout , Amir-massoud Farahmand , Tory Jarmain , Latif Abid

Building reliable retrieval-augmented generation (RAG) systems requires more than adding powerful components; it requires understanding how they interact. Using ablation studies on 50 queries (15 answerable, 10 edge cases, and 25…

计算与语言 · 计算机科学 2025-12-01 Jithin Krishnan

Diagnosing hepatic diseases accurately and interpretably is critical, yet it remains challenging in real-world clinical settings. Existing AI approaches for clinical diagnosis often lack transparency, structured reasoning, and…

人工智能 · 计算机科学 2026-03-06 Zheng Li , Jiayi Xu , Zhikai Hu , Hechang Chen , Lele Cong , Yunyun Wang , Shuchao Pang

Providing well-calibrated AI confidence can help promote users' appropriate trust in and reliance on AI, which are essential for AI-assisted decision-making. However, calibrating AI confidence -- providing confidence score that accurately…

人工智能 · 计算机科学 2025-09-30 Jingshu Li , Yitian Yang , Renwen Zhang , Q. Vera Liao , Tianqi Song , Zhengtao Xu , Yi-chieh Lee

Edge computing breaks with traditional autoscaling due to strict resource constraints, thus, motivating more flexible scaling behaviors using multiple elasticity dimensions. This work introduces an agent-based autoscaling framework that…

人工智能 · 计算机科学 2026-01-13 Boris Sedlak , Alireza Furutanpey , Zihang Wang , Víctor Casamayor Pujol , Schahram Dustdar

The rapid evolution of sophisticated cyberattacks has strained modern Security Operations Centers (SOC), which traditionally rely on rule-based or signature-driven detection systems. These legacy frameworks often generate high volumes of…

密码学与安全 · 计算机科学 2026-03-03 Chuanming Tang , Ling Qing , Shifeng Chen

We introduce a novel approach for calibrating uncertainty quantification (UQ) tailored for multi-modal large language models (LLMs). Existing state-of-the-art UQ methods rely on consistency among multiple responses generated by the LLM on…

Continual memory augmentation lets computer-using agents (CUAs) learn from prior interactions, but unvetted memories can encode domain-inappropriate or unsafe heuristics--spurious rules that drift from user intent and safety constraints. We…

‹ 上一页 1 8 9 10 下一页 ›