English
Related papers

Related papers: EchoChain: A Full-Duplex Benchmark for State-Updat…

200 papers

We introduce seqBench, a parametrized benchmark for probing sequential reasoning limits in Large Language Models (LLMs) through precise, multi-dimensional control over several key complexity dimensions. seqBench allows systematic variation…

Artificial Intelligence · Computer Science 2025-09-23 Mohammad Ramezanali , Mo Vazifeh , Paolo Santi

The maturation of Large Audio Language Models (LALMs) has raised growing expectations for them to comprehend complex audio much like humans. Current efforts primarily replicate text-based reasoning by contextualizing audio content through a…

The primary purpose of dialogue state tracking (DST), a critical component of an end-to-end conversational system, is to build a model that responds well to real-world situations. Although we often change our minds from time to time during…

Computation and Language · Computer Science 2022-10-13 Takyoung Kim , Yukyung Lee , Hoonsang Yoon , Pilsung Kang , Junseong Bang , Misuk Kim

Chain-of-Thought (CoT) reasoning has become a foundation for eliciting multi-step reasoning in large language models, but recent studies show that its benefits do not scale monotonically with chain length: while longer CoT generally enables…

Artificial Intelligence · Computer Science 2026-05-19 Bin Lei , Caiwen Ding , Jiachen Yang , Ang Li , Xin Eric Wang

The problem of synthetic speech detection has enjoyed considerable attention, with recent methods achieving low error rates across several established benchmarks. However, to what extent can low error rates on academic benchmarks translate…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-23 Ashi Garg , Zexin Cai , Lin Zhang , Henry Li Xinyuan , Leibny Paola García-Perera , Kevin Duh , Sanjeev Khudanpur , Matthew Wiesner , Nicholas Andrews

Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongside spoken words, and respond with empathy that aligns with…

Computation and Language · Computer Science 2026-03-06 Li Zhou , Lutong Yu , You Lyu , Yihang Lin , Zefeng Zhao , Junyi Ao , Yuhao Zhang , Benyou Wang , Haizhou Li

While Large Language Models and their underlying Transformer architecture are remarkably efficient, they do not reflect how our brain processes and learns a diversity of cognitive tasks such as language, nor how it leverages working memory.…

Machine Learning · Computer Science 2026-02-09 Yannis Bendi-Ouis , Xavier Hinaut

With the rapid integration of advanced reasoning capabilities into spoken dialogue models, the field urgently demands benchmarks that transcend simple interactions to address real-world complexity. However, current evaluations predominantly…

Computation and Language · Computer Science 2026-02-16 Yangzhuo Li , Shengpeng Ji , Yifu Chen , Tianle Liang , Haorong Ying , Yule Wang , Junbo Li , Jun Fang , Zhou Zhao

Conversational understanding is an integral part of modern intelligent devices. In a large fraction of the global traffic from customers using smart digital assistants, frictions in dialogues may be attributed to incorrect understanding of…

Machine Learning · Computer Science 2022-10-25 Niranjan Uma Naresh , Ziyan Jiang , Ankit , Sungjin Lee , Jie Hao , Xing Fan , Chenlei Guo

The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabilities. We introduce…

Computation and Language · Computer Science 2025-09-29 Ke Wang , Houxing Ren , Zimu Lu , Mingjie Zhan , Hongsheng Li

As multimodal content continues to expand at a rapid pace, audio retrieval has emerged as a key enabling technology for media search, content organization, and intelligent assistants. However, most existing benchmarks concentrate on…

Artificial Intelligence · Computer Science 2026-05-07 Honglei Zhang , Yuting Chen , Chenpeng Hu , Siyue Zhang , Yilei Shi

Speech-to-speech language models have recently emerged to enhance the naturalness of conversational AI. In particular, full-duplex models are distinguished by their real-time interactivity, including handling of pauses, interruptions, and…

Computation and Language · Computer Science 2026-05-13 Chung-Ming Chien , Manu Orsini , Eugene Kharitonov , Neil Zeghidour , Karen Livescu , Alexandre Défossez

We introduce Full-Duplex-Bench-v3 (FDB-v3), a benchmark for evaluating spoken language models under naturalistic speech conditions and multi-step tool use. Unlike prior work, our dataset consists entirely of real human audio annotated for…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-07 Guan-Ting Lin , Chen Chen , Zhehuai Chen , Hung-yi Lee

Tracking dialogue states to better interpret user goals and feed downstream policy learning is a bottleneck in dialogue management. Common practice has been to treat it as a problem of classifying dialogue content into a set of pre-defined…

Artificial Intelligence · Computer Science 2020-06-04 Lizi Liao , Yunshan Ma , Wenqiang Lei , Tat-Seng Chua

Full-Duplex Speech Dialogue Systems (Full-Duplex SDS) have significantly enhanced the naturalness of human-machine interaction by enabling real-time bidirectional communication. However, existing approaches face challenges such as…

Computation and Language · Computer Science 2025-05-30 Borui Liao , Yulong Xu , Jiao Ou , Kaiyuan Yang , Weihua Jian , Pengfei Wan , Di Zhang

End-to-end speech large language models ((LLMs)) extend the capabilities of text-based models to directly process and generate audio tokens. However, this often leads to a decline in reasoning and generation performance compared to text…

Sound · Computer Science 2025-05-21 Yuanbo Fang , Haoze Sun , Jun Liu , Tao Zhang , Zenan Zhou , Weipeng Chen , Xiaofen Xing , Xiangmin Xu

Replay speech attacks pose a significant threat to voice-controlled systems, especially in smart environments where voice assistants are widely deployed. While multi-channel audio offers spatial cues that can enhance replay detection…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-19 Michael Neri , Tuomas Virtanen

We introduce ReXTime, a benchmark designed to rigorously test AI models' ability to perform temporal reasoning within video events. Specifically, ReXTime focuses on reasoning across time, i.e. human-like understanding when the question and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Jr-Jen Chen , Yu-Chien Liao , Hsi-Che Lin , Yu-Chu Yu , Yen-Chun Chen , Yu-Chiang Frank Wang

Building user trust in dialogue agents requires smooth and consistent dialogue exchanges. However, agents can easily lose conversational context and generate irrelevant utterances. These situations are called dialogue breakdown, where agent…

Computation and Language · Computer Science 2023-01-23 Nathan Ng , Marzyeh Ghassemi , Narendran Thangarajan , Jiacheng Pan , Qi Guo

Currently used semantic parsing systems deployed in voice assistants can require weeks to train. Datasets for these models often receive small and frequent updates, data patches. Each patch requires training a new model. To reduce training…

Computation and Language · Computer Science 2021-03-23 Vladislav Lialin , Rahul Goel , Andrey Simanovsky , Anna Rumshisky , Rushin Shah