English
Related papers

Related papers: SoccerRef-Agents: Multi-Agent System for Automated…

200 papers

With the rapid development of mobile intelligent assistant technologies, multi-modal AI assistants have become essential interfaces for daily user interactions. However, current evaluation methods face challenges including high manual…

Artificial Intelligence · Computer Science 2025-10-22 Meiping Wang , Jian Zhong , Rongduo Han , Liming Kang , Zhengkun Shi , Xiao Liang , Xing Lin , Nan Gao , Haining Zhang

Penalties are fraught and game-changing moments in soccer games that teams explicitly prepare for. Consequently, there has been substantial interest in analyzing them in order to provide advice to practitioners. From a data science…

Machine Learning · Computer Science 2025-06-02 Lotte Bransen , Tim Janssen , Jesse Davis

Incident response (IR) requires fast, coordinated, and well-informed decision-making to contain and mitigate cyber threats. While large language models (LLMs) have shown promise as autonomous agents in simulated IR settings, their reasoning…

Computation and Language · Computer Science 2025-10-07 Zefang Liu , Arman Anwar

Modern vision-language models (VLMs) deliver impressive predictive accuracy yet offer little insight into 'why' a decision is reached, frequently hallucinating facts, particularly when encountering out-of-distribution data. Neurosymbolic…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Sanchit Sinha , Guangzhi Xiong , Zhenghao He , Aidong Zhang

Proactive agents that anticipate user intentions without explicit prompts represent a significant evolution in human-AI interaction, promising to reduce cognitive load and streamline workflows. However, existing datasets suffer from two…

Human-Computer Interaction · Computer Science 2026-02-11 Yuanbo Tang , Huaze Tang , Tingyu Cao , Lam Nguyen , Anping Zhang , Xinwen Cao , Chunkang Liu , Wenbo Ding , Yang Li

While foundation models (FMs), such as diffusion models and large vision-language models (LVLMs), have been widely applied in educational contexts, their ability to generate pedagogically effective visual explanations remains limited. Most…

Artificial Intelligence · Computer Science 2025-05-29 Haonian Ji , Shi Qiu , Siyang Xin , Siwei Han , Zhaorun Chen , Dake Zhang , Hongyi Wang , Huaxiu Yao

Multimodal large language models (MLLMs) have shown strong capabilities but remain limited to fixed modality pairs and require costly fine-tuning with large aligned datasets. Building fully omni-capable models that can integrate text,…

Artificial Intelligence · Computer Science 2025-11-06 Huawei Lin , Yunzhi Shi , Tong Geng , Weijie Zhao , Wei Wang , Ravender Pal Singh

Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. These applications require agents to perceive rapid state changes, attribute actions to the correct…

Computation and Language · Computer Science 2026-04-14 Yunzhe Wang , Runhui Xu , Kexin Zheng , Tianyi Zhang , Jayavibhav Niranjan Kogundi , Soham Hans , Volkan Ustun

Large language models (LLMs) have enabled remarkable advances in automated task-solving with multi-agent systems. However, most existing LLM-based multi-agent approaches rely on predefined agents to handle simple tasks, limiting the…

Artificial Intelligence · Computer Science 2024-05-01 Guangyao Chen , Siwei Dong , Yu Shu , Ge Zhang , Jaward Sesay , Börje F. Karlsson , Jie Fu , Yemin Shi

Multimodal Retrieval-Augmented Generation (mRAG) has emerged as a promising solution to address the temporal limitations of Multimodal Large Language Models (MLLMs) in real-world scenarios like news analysis and trending topics. However,…

Artificial Intelligence · Computer Science 2025-08-13 Yuechen Wang , Yuming Qiao , Dan Meng , Jun Yang , Haonan Lu , Zhenyu Yang , Xudong Zhang

Evaluating the performance of human is a common need across many applications, such as in engineering and sports. When evaluating human performance in completing complex and interactive tasks, the most common way is to use a metric having…

Machine Learning · Statistics 2023-03-24 Chaoyi Gu , Varuna De Silva

Agricultural visual question answering is essential for providing farmers and researchers with accurate and timely knowledge. However, many existing approaches are predominantly developed for evidence-constrained settings such as text-only…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yan Ke , Xin Yu , Heming Du , Scott Chapman , Helen Huang

Language-guided segmentation transcends the scope limitations of traditional semantic segmentation, enabling models to segment arbitrary target regions based on natural language instructions. Existing approaches typically adopt a two-stage…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Chao Hao , Jun Xu , Ji Du , Shuo Ye , Ziyue Qiao , Xiaodong Cun , Guangcong Wang , Xubin Zheng , Zitong Yu

This paper proposes a multi-agent artificial intelligence system that generates response-oriented media content in real time based on audio-derived emotional signals. Unlike conventional speech emotion recognition studies that focus…

Artificial Intelligence · Computer Science 2026-01-21 HyeYoung Lee

Despite significant advancements in Large Language Models (LLMs) and Large Vision-Language Models (LVLMs), current models still face substantial challenges in handling complex, multi-turn, and visually-grounded tasks that demand deep…

Computation and Language · Computer Science 2025-08-22 Seungmin Han , Haeun Kwon , Ji-jun Park , Taeyang Yoon

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feeding frame-level…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Noriyuki Kugo , Xiang Li , Zixin Li , Ashish Gupta , Arpandeep Khatua , Nidhish Jain , Chaitanya Patel , Yuta Kyuragi , Yasunori Ishii , Masamoto Tanabiki , Kazuki Kozuka , Ehsan Adeli

Evaluating multimodal large language models (MLLMs) is increasingly expensive, as the growing size and cross-modality complexity of benchmarks demand significant scoring efforts. To tackle with this difficulty, we introduce AutoJudger, an…

Computation and Language · Computer Science 2025-05-28 Xuanwen Ding , Chengjun Pan , Zejun Li , Jiwen Zhang , Siyuan Wang , Zhongyu Wei

The SoccerNet 2025 Challenges mark the fifth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in football video understanding. This year's challenges span four vision-based tasks: (1)…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Silvio Giancola , Anthony Cioppa , Marc Gutiérrez-Pérez , Jan Held , Carlos Hinojosa , Victor Joos , Arnaud Leduc , Floriane Magera , Karen Sanchez , Vladimir Somers , Artur Xarles , Antonio Agudo , Alexandre Alahi , Olivier Barnich , Albert Clapés , Christophe De Vleeschouwer , Sergio Escalera , Bernard Ghanem , Thomas B. Moeslund , Marc Van Droogenbroeck , Tomoki Abe , Saad Alotaibi , Faisal Altawijri , Steven Araujo , Xiang Bai , Xiaoyang Bi , Jiawang Cao , Vanyi Chao , Kamil Czarnogórski , Fabian Deuser , Mingyang Du , Tianrui Feng , Patrick Frenzel , Mirco Fuchs , Jorge García , Konrad Habel , Takaya Hashiguchi , Sadao Hirose , Xinting Hu , Yewon Hwang , Ririko Inoue , Riku Itsuji , Kazuto Iwai , Hongwei Ji , Yangguang Ji , Licheng Jiao , Yuto Kageyama , Yuta Kamikawa , Yuuki Kanasugi , Hyungjung Kim , Jinwook Kim , Takuya Kurihara , Bozheng Li , Lingling Li , Xian Li , Youxing Lian , Dingkang Liang , Hongkai Lin , Jiadong Lin , Jian Liu , Liang Liu , Shuaikun Liu , Zhaohong Liu , Yi Lu , Federico Méndez , Huadong Ma , Wenping Ma , Jacek Maksymiuk , Henry Mantilla , Ismail Mathkour , Daniel Matthes , Ayaha Motomochi , Amrulloh Robbani Muhammad , Haruto Nakayama , Joohyung Oh , Yin May Oo , Marcelo Ortega , Norbert Oswald , Rintaro Otsubo , Fabian Perez , Mengshi Qi , Cristian Rey , Abel Reyes-Angulo , Oliver Rose , Hoover Rueda-Chacón , Hideo Saito , Jose Sarmiento , Kanta Sawafuji , Atom Scott , Xi Shen , Pragyan Shrestha , Jae-Young Sim , Long Sun , Yuyang Sun , Tomohiro Suzuki , Licheng Tang , Masato Tonouchi , Ikuma Uchida , Henry O. Velesaca , Tiancheng Wang , Rio Watanabe , Jay Wu , Yongliang Wu , Shunzo Yamagishi , Di Yang , Xu Yang , Yuxin Yang , Hao Ye , Xinyu Ye , Calvin Yeung , Xuanlong Yu , Chao Zhang , Dingyuan Zhang , Kexing Zhang , Zhe Zhao , Xin Zhou , Wenbo Zhu , Julian Ziegler

Multimodal large language models (MLLMs) hold significant potential in medical applications, including disease diagnosis and clinical decision-making. However, these tasks require highly accurate, context-sensitive, and professionally…

Computation and Language · Computer Science 2025-09-01 Meidan Ding , Jipeng Zhang , Wenxuan Wang , Cheng-Yi Li , Wei-Chieh Fang , Hsin-Yu Wu , Haiqin Zhong , Wenting Chen , Linlin Shen

The rapid evolution of Retrieval-Augmented Generation (RAG) toward multimodal, high-stakes enterprise applications has outpaced the development of domain specific evaluation benchmarks. Existing datasets often rely on general-domain corpora…

Artificial Intelligence · Computer Science 2026-01-23 Chandan Kumar Sahu , Premith Kumar Chilukuri , Matthew Hetrich
‹ Prev 1 3 4 5 6 7 10 Next ›