English
Related papers

Related papers: OPOR-Bench: Evaluating Large Language Models on On…

200 papers

Recent advances in task-oriented dialogue (TOD) systems, driven by large language models (LLMs) with extensive API and tool integration, have enabled conversational agents to coordinate interleaved goals, maintain long-horizon context, and…

Computation and Language · Computer Science 2026-02-02 Yifei Zhang , Hooshang Nayyeri , Rinat Khaziev , Emine Yilmaz , Gokhan Tur , Dilek Hakkani-Tür , Hari Thadakamalla

As model context lengths continue to grow, concerns about whether models effectively use the full context length have persisted. While several carefully designed long-context evaluations have recently been released, these evaluations tend…

Computation and Language · Computer Science 2025-11-05 Amanda Bertsch , Adithya Pratapa , Teruko Mitamura , Graham Neubig , Matthew R. Gormley

Existing AI benchmarks for software automation rarely combine cross-application coordination, autonomous API discovery, and policy adherence. Real business workflows demand all three: a single task may span a CRM, inbox, calendar, and…

Artificial Intelligence · Computer Science 2026-04-22 Daniel Shepard , Robin Salimans

The automatic content analysis of mass media in the social sciences has become necessary and possible with the raise of social media and computational power. One particularly promising avenue of research concerns the use of opinion mining.…

Social and Information Networks · Computer Science 2015-12-01 Pedro Saleiro , Sílvio Amir , Mário J. Silva , Carlos Soares

We present AutoBench, a fully automated and self-sustaining framework for evaluating Large Language Models (LLMs) through reciprocal peer assessment. This paper provides a rigorous scientific validation of the AutoBench methodology,…

Computation and Language · Computer Science 2025-10-28 Dario Loi , Elena Maria Muià , Federico Siciliano , Giovanni Trappolini , Vincenzo Crisà , Peter Kruger , Fabrizio Silvestri

Understanding research papers remains challenging for foundation models due to specialized scientific discourse and complex figures and tables, yet existing benchmarks offer limited fine-grained evaluation at scale. To address this gap, we…

Computation and Language · Computer Science 2026-05-01 Yelin Chen , Fanjin Zhang , Suping Sun , Yunhe Pang , Yuanchun Wang , Jian Song , Xiaoyan Li , Lei Hou , Shu Zhao , Jie Tang , Juanzi Li

Large Language Models (LLMs) with agentic web search capabilities show strong potential for tasks requiring real-time information access and complex fact retrieval, yet evaluating such systems remains challenging. We introduce \bench, a…

Information Retrieval · Computer Science 2026-02-17 Yunfan Zhang , Kathleen McKeown , Smaranda Muresan

Realistic practice and tailored feedback are key processes for training peer counselors with clinical skills. However, existing mechanisms of providing feedback largely rely on human supervision. Peer counselors often lack mechanisms to…

Computation and Language · Computer Science 2024-03-26 Alicja Chaszczewicz , Raj Sanjay Shah , Ryan Louie , Bruce A Arnow , Robert Kraut , Diyi Yang

The ability to research and synthesize knowledge is central to human expertise and progress. A new class of AI systems--designed for generative research synthesis--aims to automate this process by retrieving information from the live web…

Computation and Language · Computer Science 2026-02-10 Liana Patel , Negar Arabzadeh , Harshit Gupta , Ankita Sundar , Ion Stoica , Matei Zaharia , Carlos Guestrin

AI agents are beginning to complete valuable, long-horizon business operations tasks, but training and evaluation environments for enterprise work still struggle to balance realism, verifiability, and scale. Environment and task creation…

Artificial Intelligence · Computer Science 2026-05-27 Maksim Ivanov , Abhijay Rana

Benchmarks play a significant role in how technology companies communicate about model capabilities and how researchers and the public understand generative AI systems. However, existing benchmarks have been criticized for their failure to…

Human-Computer Interaction · Computer Science 2026-04-29 Charlotte Li , Nick Hagar , Sachita Nishal , Jeremy Gilbert , Nick Diakopoulos

We propose a scalable method for constructing a temporal opinion knowledge base with large language models (LLMs) as automated annotators. Despite the demonstrated utility of time-series opinion analysis of text for downstream applications…

Computation and Language · Computer Science 2025-09-03 Gaurav Negi , Atul Kr. Ojha , Omnia Zayed , Paul Buitelaar

While online conversations can cover a vast amount of information in many different formats, abstractive text summarization has primarily focused on modeling solely news articles. This research gap is due, in part, to the lack of…

Computation and Language · Computer Science 2021-06-03 Alexander R. Fabbri , Faiaz Rahman , Imad Rizvi , Borui Wang , Haoran Li , Yashar Mehdad , Dragomir Radev

Off-policy evaluation (OPE) holds the promise of being able to leverage large, offline datasets for both evaluating and selecting complex policies for decision making. The ability to learn offline is particularly important in many…

Recently, instruction-following audio-language models have received broad attention for human-audio interaction. However, the absence of benchmarks capable of evaluating audio-centric interaction capabilities has impeded advancements in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-29 Qian Yang , Jin Xu , Wenrui Liu , Yunfei Chu , Ziyue Jiang , Xiaohuan Zhou , Yichong Leng , Yuanjun Lv , Zhou Zhao , Chang Zhou , Jingren Zhou

Video generation has achieved remarkable progress, with generated videos increasingly resembling real ones. However, the rapid advance in generation has outpaced the development of adequate evaluation metrics. Currently, the assessment of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Nabyl Quignon , Baptiste Chopin , Yaohui Wang , Antitza Dantcheva

While Retrieval Augmented Generation (RAG) is now widely adopted to enhance LLMs, evaluating its true performance benefits in a reproducible and interpretable way remains a major hurdle. Existing methods often fall short: they lack domain…

Information Retrieval · Computer Science 2025-08-11 Jiaxuan Liang , Shide Zhou , Kailong Wang

While large language models (LLMs) have become the de facto framework for literature-related tasks, they still struggle to function as domain-specific literature agents due to their inability to connect pieces of knowledge and reason across…

Digital Libraries · Computer Science 2026-03-03 Andreas Varvarigos , Ali Maatouk , Jiasheng Zhang , Ngoc Bui , Jialin Chen , Leandros Tassiulas , Rex Ying

Opinions in scientific research papers can be divergent, leading to controversies among reviewers. However, most existing datasets for opinion summarization are centered around product reviews and assume that the analyzed opinions are…

Computation and Language · Computer Science 2024-06-18 Qi Zeng , Mankeerat Sidhu , Ansel Blume , Hou Pong Chan , Lu Wang , Heng Ji

Large language models (LLMs) are increasingly capable of carrying out long-running, real-world tasks. However, as the amount of context grows, their reliability often deteriorates, a phenomenon known as "context rot". Existing long-context…

Artificial Intelligence · Computer Science 2026-02-10 Weihao Zeng , Yuzhen Huang , Junxian He
‹ Prev 1 4 5 6 7 8 10 Next ›