English
Related papers

Related papers: Taskmaster Deconstructed: A Quantitative Look at T…

200 papers

We introduce WorkBench: a benchmark dataset for evaluating agents' ability to execute tasks in a workplace setting. WorkBench contains a sandbox environment with five databases, 26 tools, and 690 tasks. These tasks represent common business…

Computation and Language · Computer Science 2024-08-06 Olly Styles , Sam Miller , Patricio Cerda-Mardini , Tanaya Guha , Victor Sanchez , Bertie Vidgen

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

This paper studies a novel stochastic compartmental model that describes the dynamics of trust in society. The population is split into three compartments representing levels of trust in society: trusters, skeptics and doubters. The focus…

Physics and Society · Physics 2024-09-20 Benedikt Valentin Meylahn , Koen De Turck , Michel Mandjes

Creating agents that can interact naturally with humans is a common goal in artificial intelligence (AI) research. However, evaluating these interactions is challenging: collecting online human-agent interactions is slow and expensive, yet…

Cooperative multi-agent reinforcement learning (MARL) benchmarks commonly emphasize aggregate outcomes such as return, success rate, or completion time. While essential, these metrics often fail to reveal how agents coordinate, particularly…

Multiagent Systems · Computer Science 2026-05-08 Maria Ana Cardei , Matthew Landers , Afsaneh Doryab

The Mixmaster or Bianchi IX cosmological model has become one of the archetypal settings for studying gravitational dynamics. The past decade has seen a vigourous debate about whether or not the Mixmaster's dynamics is chaotic. In this talk…

General Relativity and Quantum Cosmology · Physics 2007-05-23 Neil J. Cornish , Janna J. Levin

Large language models perform text generation through high-dimensional internal dynamics, yet the temporal organisation of these dynamics remains poorly understood. Most interpretability approaches emphasise static representations or causal…

Artificial Intelligence · Computer Science 2026-01-21 Hassan Ugail , Newton Howard

Suppose some cleverness score parameter is sufficiently interesting to be defined and then measured, perhaps for different strata of specialists or for the broader population. Such phenomena could have Gaussian distributions, when it comes…

Other Statistics · Statistics 2026-05-19 Nils Lid Hjort

This work focuses on the nature of visibility in societies where the behaviours of humans and algorithms influence each other - termed algorithmically infused societies. We propose a quantitative measure of visibility, with implications and…

Social and Information Networks · Computer Science 2024-12-09 Shaojing Sun , Zhiyuan Liu , David Waxman

In recent years, the remarkable progress of large language models (LLMs) has sparked interest in task automation, which involves decomposing complex tasks described by user instructions into sub-tasks and invoking external tools to execute…

Computation and Language · Computer Science 2024-11-04 Yongliang Shen , Kaitao Song , Xu Tan , Wenqi Zhang , Kan Ren , Siyu Yuan , Weiming Lu , Dongsheng Li , Yueting Zhuang

Benchmarks are often used as a standard to understand LLM capabilities in different domains. However, aggregate benchmark scores provide limited insight into compositional skill gaps of LLMs and how to improve them. To make these weaknesses…

Computation and Language · Computer Science 2026-04-22 Sungeun An , Swanand Ravindra Kadhe , Shailja Thakur , Chad DeLuca , Hima Patel

In this work, we present a new dataset for computational humor, specifically comparative humor ranking, which attempts to eschew the ubiquitous binary approach to humor detection. The dataset consists of tweets that are humorous responses…

Computation and Language · Computer Science 2017-04-18 Peter Potash , Alexey Romanov , Anna Rumshisky

Large language model-based web agents have demonstrated strong performance on realistic web interaction tasks. However, existing evaluations are predominantly conducted under relatively stable and well-behaved interaction conditions, which…

Software Engineering · Computer Science 2026-04-21 Haoyue Bai , Dong Wang , Long Chen , Bingguang Hao , Pengyang Shao , Yonghui Yang , Yicheng He , Chenyi Zhuang

Physiological signals can potentially be applied as objective measures to understand the behavior and engagement of users interacting with information access systems. However, the signals are highly sensitive, and many controls are required…

Information Retrieval · Computer Science 2023-04-27 Kaixin Ji , Damiano Spina , Danula Hettiachchi , Flora Dilys Salim , Falk Scholer

Beyond possessing large enough size to feed data hungry machines (eg, transformers), what attributes measure the quality of a dataset? Assuming that the definitions of such attributes do exist, how do we quantify among their relative…

Computer Vision and Pattern Recognition · Computer Science 2022-04-19 Rajat Modi , Aayush Jung Rana , Akash Kumar , Praveen Tirupattur , Shruti Vyas , Yogesh Singh Rawat , Mubarak Shah

This letter examines the controllability of consensus dynamics on matrix-weighed networks from a graph-theoretic perspective. Unlike the scalar-weighted networks, the rank of weight matrix introduces additional intricacies into…

Systems and Control · Electrical Eng. & Systems 2020-01-14 Lulu Pan , Haibin Shao , Mehran Mesbahi , Yugeng Xi , Dewei Li

Today's popular TV series tend to develop continuous, complex plots spanning several seasons, but are often viewed in controlled and discontinuous conditions. Consequently, most viewers need to be re-immersed in the story before watching a…

Recent interest in self-supervised dense tracking has yielded rapid progress, but performance still remains far from supervised methods. We propose a dense tracking model trained on videos without any annotations that surpasses previous…

Computer Vision and Pattern Recognition · Computer Science 2020-02-27 Zihang Lai , Erika Lu , Weidi Xie

Analyzing memes on the internet has emerged as a crucial endeavor due to the impact this multi-modal form of content wields in shaping online discourse. Memes have become a powerful tool for expressing emotions and sentiments, possibly even…

Although transformer-based models have shown exceptional empirical performance, the fundamental principles governing their training dynamics are inadequately characterized beyond configuration-specific studies. Inspired by empirical…

Machine Learning · Computer Science 2025-10-09 Zheng-An Chen , Tao Luo
‹ Prev 1 8 9 10 Next ›