English
Related papers

Related papers: FlexDuo: A Pluggable System for Enabling Full-Dupl…

200 papers

While Large Language Models (LLMs) provide semantic flexibility for robotic task planning, their susceptibility to hallucination and logical inconsistency limits their reliability in long-horizon domains. To bridge the gap between…

Artificial Intelligence · Computer Science 2026-03-26 Keru Hua , Ding Wang , Yaoying Gu , Xiaoguang Ma

This paper proposes a dual-stage, low complexity, and reconfigurable technique to enhance the speech contaminated by various types of noise sources. Driven by input data and audio contents, the proposed dual-stage speech enhancement…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-18 Jun Yang , Nico Brailovsky

Efforts towards endowing robots with the ability to speak have benefited from recent advancements in natural language processing, in particular large language models. However, current language models are not fully incremental, as their…

Computation and Language · Computer Science 2025-04-03 Casey Kennington , Pierre Lison , David Schlangen

The Fluid Antenna System (FAS), which enables flexible Multiple-Input Multiple-Output (MIMO) communications, introduces new spatial degrees of freedom for next-generation wireless networks. Unlike traditional MIMO, FAS involves joint port…

Information Theory · Computer Science 2025-06-18 Chao Wang , Kai-Kit Wong , Zan Li , Liang Jin , Chan-Byoung Chae

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. To improve robustness of speaker recognition system performance in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Yanpei Shi , Qiang Huang , Thomas Hain

Robust task-oriented spoken dialogue agents require exposure to the full diversity of how people interact through speech. Building spoken user simulators that address this requires large-scale spoken task-oriented dialogue (TOD) data…

Computation and Language · Computer Science 2026-03-18 Jonggeun Lee , Junseong Pyo , Jeongmin Park , Yohan Jo

We consider a dynamic time division duplex (DTDD) enabled cell-free massive multiple-input multiple-output (CF-mMIMO) system, where each half-duplex (HD) access point (AP) is scheduled to operate in the uplink (UL) or downlink (DL) mode…

Signal Processing · Electrical Eng. & Systems 2022-05-24 Anubhab Chowdhury , Ribhu Chopra , Chandra R. Murthy

The goal of dialogue state tracking (DST) is to predict the current dialogue state given all previous dialogue contexts. Existing approaches generally predict the dialogue state at every turn from scratch. However, the overwhelming majority…

Computation and Language · Computer Science 2021-07-28 Jinyu Guo , Kai Shuang , Jijie Li , Zihan Wang

Instruction-tuned language models increasingly rely on large multi-turn dialogue corpora, but these datasets are often noisy and structurally inconsistent, with topic drift, repetitive chitchat, and mismatched answer formats across turns.…

Computation and Language · Computer Science 2026-04-21 Bo Li , Shikun Zhang , Wei Ye

The traditional vocoders have the advantages of high synthesis efficiency, strong interpretability, and speech editability, while the neural vocoders have the advantage of high synthesis quality. To combine the advantages of two vocoders,…

Sound · Computer Science 2022-03-08 Tao Wang , Ruibo Fu , Jiangyan Yi , Jianhua Tao , Zhengqi Wen

Dialogue state tracking (DST) is a key component of task-oriented dialogue systems. DST estimates the user's goal at each user turn given the interaction until then. State of the art approaches for state tracking rely on deep learning…

Computation and Language · Computer Science 2018-01-03 Abhinav Rastogi , Dilek Hakkani-Tur , Larry Heck

Modules in all existing speech separation networks can be categorized into single-input-multi-output (SIMO) modules and single-input-single-output (SISO) modules. SIMO modules generate more outputs than input, and SISO modules keep the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-18 Yi Luo , Zhuo Chen , Cong Han , Chenda Li , Tianyan Zhou , Nima Mesgarani

We introduce Moshi, a speech-text foundation model and full-duplex spoken dialogue framework. Current systems for spoken dialogue rely on pipelines of independent components, namely voice activity detection, speech recognition, textual…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-03 Alexandre Défossez , Laurent Mazaré , Manu Orsini , Amélie Royer , Patrick Pérez , Hervé Jégou , Edouard Grave , Neil Zeghidour

The capacity for highly complex, evidence-based, and strategically adaptive persuasion remains a formidable great challenge for artificial intelligence. Previous work, like IBM Project Debater, focused on generating persuasive speeches in…

Computation and Language · Computer Science 2025-11-25 Allen Roush , Devin Gonier , John Hines , Judah Goldfeder , Philippe Martin Wyder , Sanjay Basu , Ravid Shwartz Ziv

Echo and noise suppression is an integral part of a full-duplex communication system. Many recent acoustic echo cancellation (AEC) systems rely on a separate adaptive filtering module for linear echo suppression and a neural module for…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-07 Karn N. Watcharasupat , Thi Ngoc Tho Nguyen , Woon-Seng Gan , Shengkui Zhao , Bin Ma

Modern virtual personal assistants provide a convenient interface for completing daily tasks via voice commands. An important consideration for these assistants is the ability to recover from automatic speech recognition (ASR) and natural…

Computation and Language · Computer Science 2017-12-13 Maryam Fazel-Zarandi , Shang-Wen Li , Jin Cao , Jared Casale , Peter Henderson , David Whitney , Alborz Geramifard

Simultaneous speech-to-speech translation (SimulS2S) is essential for real-time multilingual communication, with increasing integration into meeting and streaming platforms. Despite this, SimulS2S remains underexplored in research, where…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-19 Amirbek Djanibekov , Luisa Bentivogli , Matteo Negri , Sara Papi

This paper considers the implementation and application possibilities of a compact full duplex multiple-input multiple-output (MIMO) architecture where direct communication exists between users, e.g., device-to-device (D2D) and cellular…

Information Theory · Computer Science 2017-03-27 MinKeun Chung , Min Soo Sim , Dong Ku Kim , Chan-Byoung Chae

Although in cellular networks full-duplex and dynamic time-division duplexing promise increased spectrum efficiency, their potential is so far challenged by increased interference. While previous studies have shown that self-interference…

Information Theory · Computer Science 2020-10-23 José Mairton B. da Silva , Gustav Wikström , Ratheesh K. Mungara , Carlo Fischione

This paper considers a cellular system with a full-duplex base station and half-duplex users. The base station can activate one user in uplink or downlink (half-duplex mode), or two different users one in each direction simultaneously…

Information Theory · Computer Science 2018-05-01 Shahram Shahsavari , David Ramirez , Elza Erkip