中文
相关论文

相关论文: Gamified Speaker Comparison by Listening

200 篇论文

Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based…

Voice assistants have become an essential tool for people with various disabilities because they enable complex phone- or tablet-based interactions without the need for fine-grained motor control, such as with touchscreens. However, these…

音频与语音处理 · 电气工程与系统科学 2022-02-17 Colin Lea , Zifang Huang , Dhruv Jain , Lauren Tooley , Zeinab Liaghat , Shrinath Thelapurath , Leah Findlater , Jeffrey P. Bigham

This technical report describes the EML submission to the first VoxCeleb speaker diarization challenge. Although the aim of the challenge has been the offline processing of the signals, the submitted system is basically the EML online…

声音 · 计算机科学 2020-10-26 Omid Ghahabi , Volker Fischer

End-to-end spoken dialogue models such as GPT-4o-audio have recently garnered significant attention in the speech domain. However, the evaluation of spoken dialogue models' conversational performance has largely been overlooked. This is…

音频与语音处理 · 电气工程与系统科学 2025-09-24 Shengpeng Ji , Tianle Liang , Yangzhuo Li , Jialong Zuo , Minghui Fang , Jinzheng He , Yifu Chen , Zhengqing Liu , Ziyue Jiang , Xize Cheng , Siqi Zheng , Jin Xu , Junyang Lin , Zhou Zhao

Text-based games present a unique challenge for autonomous agents to operate in natural language and handle enormous action spaces. In this paper, we propose the Contextual Action Language Model (CALM) to generate a compact set of action…

计算与语言 · 计算机科学 2020-10-07 Shunyu Yao , Rohan Rao , Matthew Hausknecht , Karthik Narasimhan

In this article we propose a novel approach for adapting speaker embeddings to new domains based on adversarial training of neural networks. We apply our embeddings to the task of text-independent speaker verification, a challenging,…

音频与语音处理 · 电气工程与系统科学 2018-11-08 Gautam Bhattacharya , Jahangir Alam , Patrick Kenny

We present "Human or Not?", an online game inspired by the Turing test, that measures the capability of AI chatbots to mimic humans in dialog, and of humans to tell bots from other humans. Over the course of a month, the game was played by…

人工智能 · 计算机科学 2023-06-01 Daniel Jannai , Amos Meron , Barak Lenz , Yoav Levine , Yoav Shoham

The most pressing challenge in the field of voice biometrics is selecting the most efficient technique of speaker recognition. Every individual's voice is peculiar, factors like physical differences in vocal organs, accent and pronunciation…

声音 · 计算机科学 2017-12-05 Rishi Charan , Manisha. A , Karthik. R , Rajesh Kumar M

In an era of human-computer interaction with increasingly agentic AI systems capable of connecting with users conversationally, speech is an important modality for commanding agents. By recognizing and using speech emotions (i.e., how a…

人机交互 · 计算机科学 2025-04-14 Ilhan Aslan , Timothy Merritt , Stine S. Johansen , Niels van Berkel

While there has been significant progress towards modelling coherence in written discourse, the work in modelling spoken discourse coherence has been quite limited. Unlike the coherence in text, coherence in spoken discourse is also…

计算与语言 · 计算机科学 2021-01-05 Rajaswa Patil , Yaman Kumar Singla , Rajiv Ratn Shah , Mika Hama , Roger Zimmermann

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. To improve robustness of speaker recognition system performance in…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Yanpei Shi , Qiang Huang , Thomas Hain

Despite its broad practical applications such as in fraud prevention, open-set speaker identification (OSI) has received less attention in the speaker recognition community compared to speaker verification (SV). OSI deals with determining…

音频与语音处理 · 电气工程与系统科学 2023-07-04 Raghuveer Peri , Seyed Omid Sadjadi , Daniel Garcia-Romero

As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-specific understanding to support user authorization,…

声音 · 计算机科学 2026-05-15 KiHyun Nam , Jungwoo Heo , Siu Bae , Ha-Jin Yu , Joon Son Chung

We present a new listening head generation benchmark, for synthesizing responsive feedbacks of a listener (e.g., nod, smile) during a face-to-face conversation. As the indispensable complement to talking heads generation, listening head…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Mohan Zhou , Yalong Bai , Wei Zhang , Ting Yao , Tiejun Zhao , Tao Mei

The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabilities. We introduce…

计算与语言 · 计算机科学 2025-09-29 Ke Wang , Houxing Ren , Zimu Lu , Mingjie Zhan , Hongsheng Li

Study Objective: Machine learning models have advanced medical image processing and can yield faster, more accurate diagnoses. Despite a wealth of available medical imaging data, high-quality labeled data for model training is lacking. We…

We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding…

声音 · 计算机科学 2021-02-11 Zeqian Li , Jacob Whitehill

Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent…

声音 · 计算机科学 2020-11-05 Arsha Nagrani , Joon Son Chung , Andrew Zisserman

Concerns that interacting with generative AI homogenizes human cognition are largely based on evidence from text-based interactions, potentially conflating the effects of AI systems with those of written communication. This study examines…

人机交互 · 计算机科学 2026-03-31 Neelam Modi Jain , Dan J. Wang

Facial recognition system is one of the major successes of Artificial intelligence and has been used a lot over the last years. But, images are not the only biometric present: audio is another possible biometric that can be used as an…

音频与语音处理 · 电气工程与系统科学 2020-08-26 Prateek Mishra