English
Related papers

Related papers: Who Spoke What? A Latent Variable Framework for th…

200 papers

This paper presents a unified model to perform language and speaker recognition simultaneously and altogether. The model is based on a multi-task recurrent neural network where the output of one task is fed as the input of the other,…

Sound · Computer Science 2017-05-24 Lantian Li , Zhiyuan Tang , Dong Wang , Andrew Abel , Yang Feng , Shiyue Zhang

While signal conversion and disentangled representation learning have shown promise for manipulating data attributes across domains such as audio, image, and multimodal generation, existing approaches, especially for speech style…

Sound · Computer Science 2025-10-10 Jonathan Svirsky , Ofir Lindenbaum , Uri Shaham

Sequential data often possesses a hierarchical structure with complex dependencies between subsequences, such as found between the utterances in a dialogue. In an effort to model this kind of generative process, we propose a neural…

Computation and Language · Computer Science 2016-06-15 Iulian Vlad Serban , Alessandro Sordoni , Ryan Lowe , Laurent Charlin , Joelle Pineau , Aaron Courville , Yoshua Bengio

Turn-taking prediction is the task of anticipating when the speaker in a conversation will yield their turn to another speaker to begin speaking. This project expands on existing strategies for turn-taking prediction by employing a…

Sound · Computer Science 2024-12-25 Hyunbae Jeon , Frederic Guintu , Rayvant Sahni

Large Language Models (LLMs) have demonstrated remarkable capabilities in understanding and generating human-like text, yet they largely operate as reactive agents, responding only when directly prompted. This passivity creates an…

Computation and Language · Computer Science 2026-05-18 Deep Anil Patel , Iain Melvin , Christopher Malon , Martin Renqiang Min

Numerous online conversations are produced on a daily basis, resulting in a pressing need to conversation understanding. As a basis to structure a discussion, we identify the responding relations in the conversation discourse, which link…

Computation and Language · Computer Science 2021-04-20 Lu Ji , Jing Li , Zhongyu Wei , Qi Zhang , Xuanjing Huang

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-09 Shaojin Ding , Quan Wang , Shuo-yiin Chang , Li Wan , Ignacio Lopez Moreno

Despite achieving satisfactory performance in speaker verification using deep neural networks, variable-duration utterances remain a challenge that threatens the robustness of systems. To deal with this issue, we propose a speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Ju-ho Kim , Hye-jin Shim , Jungwoo Heo , Ha-Jin Yu

Text-independent speaker verification is an important artificial intelligence problem that has a wide spectrum of applications, such as criminal investigation, payment certification, and interest-based customer services. The purpose of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-22 Jiwei Xu , Xinggang Wang , Bin Feng , Wenyu Liu

In this paper, we address the problem of enhancing the speech of a speaker of interest in a cocktail party scenario when visual information of the speaker of interest is available. Contrary to most previous studies, we do not learn visual…

Computation and Language · Computer Science 2021-02-04 Giovanni Morrone , Luca Pasa , Vadim Tikhanoff , Sonia Bergamaschi , Luciano Fadiga , Leonardo Badino

Multi-speaker automatic speech recognition (MASR) aims to predict ''who spoke when and what'' from multi-speaker speech, a key technology for multi-party dialogue understanding. However, most existing approaches decouple temporal modeling…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Yifan Hu , Peiji Yang , Zhisheng Wang , Yicheng Zhong , Rui Liu

Cognitive modeling commonly relies on asking participants to complete a battery of varied tests in order to estimate attention, working memory, and other latent variables. In many cases, these tests result in highly variable observation…

This paper proposes attentive statistics pooling for deep speaker embedding in text-independent speaker verification. In conventional speaker embedding, frame-level features are averaged over all the frames of a single utterance to form an…

Audio and Speech Processing · Electrical Eng. & Systems 2019-02-27 Koji Okabe , Takafumi Koshinaka , Koichi Shinoda

Probabilistic Latent Semantic Analysis is a novel statistical technique for the analysis of two-mode and co-occurrence data, which has applications in information retrieval and filtering, natural language processing, machine learning from…

Machine Learning · Computer Science 2013-01-30 Thomas Hofmann

Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning of Large Language Models (LLMs), and recent Mixture-of-Experts (MoE) extensions further enhance flexibility by dynamically combining multiple LoRA experts. However, existing…

Machine Learning · Computer Science 2026-04-14 Lin Mu , Haiyang Wang , Li Ni , Lei Sang , Zhize Wu , Peiquan Jin , Yiwen Zhang

With latent variables, stochastic recurrent models have achieved state-of-the-art performance in modeling sound-wave sequence. However, opposite results are also observed in other domains, where standard recurrent networks often outperform…

Machine Learning · Computer Science 2019-09-17 Zihang Dai , Guokun Lai , Yiming Yang , Shinjae Yoo

Large Language Models (LLMs) have demonstrated remarkable performance across various disciplines and tasks. However, benchmarking their capabilities with multilingual spoken queries remains largely unexplored. In this study, we introduce…

Computation and Language · Computer Science 2025-05-27 Firoj Alam , Md Arid Hasan , Shammur Absar Chowdhury

This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker embeddings to enhance the performance of the phrasing model. We…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Dong Yang , Yuki Saito , Takaaki Saeki , Tomoki Koriyama , Wataru Nakata , Detai Xin , Hiroshi Saruwatari

We present a generative dialogue system capable of operating in a full-duplex manner, allowing for seamless interaction. It is based on a large language model (LLM) carefully aligned to be aware of a perception module, a motor function…

Computation and Language · Computer Science 2024-10-30 Peng Wang , Songshuo Lu , Yaohua Tang , Sijie Yan , Wei Xia , Yuanjun Xiong

We propose an approach to extract speaker embeddings that are robust to speaking style variations in text-independent speaker verification. Typically, speaker embedding extraction includes training a DNN for speaker classification and using…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-29 Amber Afshan , Abeer Alwan