English
Related papers

Related papers: AION: Next-Generation Tasks and Practical Harness …

200 papers

AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of how well agents really work. We introduce the Holistic…

The Radio Access Network (RAN) is evolving into a programmable and disaggregated infrastructure that increasingly relies on AI-native algorithms for optimization and closed-loop control. However, current RAN intelligence is still largely…

Networking and Internet Architecture · Computer Science 2026-04-07 Ioannis Panitsas , Leandros Tassiulas

As organizations move toward production deployments of AI agents, which execute non-deterministic workflows, maintain stateful sessions, and often operate with privileged access to internal services, the engineering challenge shifts from…

Artificial Intelligence · Computer Science 2026-05-28 Nikita Benkovich , Vitalii Valkov

This report describes ARIS (Auto-Research-in-sleep), an open-source research harness for autonomous research, including its architecture, assurance mechanisms, and early deployment experience. The performance of agent systems built on LLMs…

Software Engineering · Computer Science 2026-05-06 Ruofeng Yang , Yongcan Li , Shuai Li

Time series anomaly detection is a vital task in many domains, including patient monitoring in healthcare, forecasting in finance, and predictive maintenance in energy industries. This has led to a proliferation of anomaly detection…

Machine Learning · Computer Science 2024-11-26 Sarah Alnegheimish , Laure Berti-Equille , Kalyan Veeramachaneni

Modern IT system operation demands the integration of system software and hardware metrics. As a result, it generates a massive amount of data, which can be potentially used to make data-driven operational decisions. In the basic form, the…

Machine Learning · Computer Science 2022-11-16 Jiajia Li , Feng Tan , Cheng He , Zikai Wang , Haitao Song , Lingfei Wu , Pengwei Hu

Recent breakthroughs in natural language processing and computer vision, driven by efficient pre-training on large datasets, have enabled foundation models to excel on a wide range of tasks. However, this potential has not yet been fully…

Machine Learning · Computer Science 2025-02-03 Özgün Turgut , Philip Müller , Martin J. Menten , Daniel Rueckert

User behavior in the real world is diverse, cross-domain, and spans long time horizons. Existing user modeling benchmarks however remain narrow, focusing mainly on short sessions and next-item prediction within a single domain. Such…

Information Retrieval · Computer Science 2026-04-21 Arnav Goel , Pranjal A Chitale , Bhawna Paliwal , Bishal Santra , Amit Sharma

Understanding why some sequential planning problems are harder than others requires models that go beyond average performance. They should capture the specific pattern of which problems are hard, and ideally fail in the same way people do…

Robotics · Computer Science 2026-05-19 Michael Migacev , Vito Mengers , Antonia Köngeter , Oliver Brock

Understanding time series data is fundamental to many real-world applications. Recent work explores multimodal large language models (MLLMs) to enhance time series understanding with contextual information beyond numerical signals. This…

Machine Learning · Computer Science 2026-02-03 Yaxuan Kong , Yiyuan Yang , Shiyu Wang , Chenghao Liu , Yuxuan Liang , Ming Jin , Stefan Zohren , Dan Pei , Yan Liu , Qingsong Wen

A growing trend in modern data analysis is the integration of data management with learning, guided by accuracy, latency, and cost requirements. In practice, applications draw data of different formats from many sources. In the meanwhile,…

Databases · Computer Science 2025-10-15 Meihui Zhang , Liming Wang , Chi Zhang , Zhaojing Luo

Large Language Models are increasingly deployed as autonomous agents for complex real-world tasks, yet existing systems often focus on isolated improvements without a unifying design for robustness and adaptability. We propose a generalist…

AI agents are entering high-risk production settings, where they use tools, retain context, follow policies, handle private data, and interact with users over multiple turns. Yet many evaluation methods still judge isolated outputs or…

Multiagent Systems · Computer Science 2026-05-26 Fouad Bousetouane

Evidence on AI in software engineering still leans heavily toward individual task completion, while evidence on team-level delivery remains scarce. We report a retrospective longitudinal field study of Chiron, an industrial platform that…

Software Engineering · Computer Science 2026-03-23 Maximiliano Armesto , Christophe Kolb

The concept of "task" is at the core of artificial intelligence (AI): Tasks are used for training and evaluating AI systems, which are built in order to perform and automatize tasks we deem useful. In other fields of engineering theoretical…

Artificial Intelligence · Computer Science 2016-05-13 Kristinn R. Thórisson , Jordi Bieger , Thröstur Thorarensen , Jóna S. Sigurðardóttir , Bas R. Steunebrink

This paper presents the rAIson platform, a high-level technological environment for the development of automated, reliable and explainable decision-making agents. The research underlying the platform and its technological progress has now…

Multiagent Systems · Computer Science 2026-05-05 Pavlos Moraitis , Nikolaos Spanoudakis , Antonis Kakas

Cellular research and development (R&D) is throttled by six structural processes that each consume months of manual engineering work per iteration: (i) synthesizing new features from standards or research papers into production code; (ii)…

Agentic AI systems plan, use tools, maintain state, and produce multi-step trajectories with external effects. Those properties create a governance problem that differs materially from single-turn generative AI: important risks emerge dur-…

Artificial Intelligence · Computer Science 2026-04-08 Christopher Koch

LLM agents are increasingly deployed as executable systems that use tools, modify workspaces, and produce concrete artifacts. In such workflows, performance depends not only on the base model, but also on the harness: the system layer that…

Artificial Intelligence · Computer Science 2026-05-28 Yilun Yao , Xinyu Tan , Chao-Hsuan Liu , Yaoming Li , Zhengyang Wang , Wenhan Yu , Zhewen Tan , Yuxuan Tian , Guangxiang Zhao , Lin Sun , Xiangzheng Zhang , Tong Yang

We introduce TimeSeriesGym, a scalable benchmarking framework for evaluating Artificial Intelligence (AI) agents on time series machine learning engineering challenges. Existing benchmarks lack scalability, focus narrowly on model building…

Machine Learning · Computer Science 2025-05-20 Yifu Cai , Xinyu Li , Mononito Goswami , Michał Wiliński , Gus Welter , Artur Dubrawski