中文
相关论文

相关论文: The MIT Voice Name System

200 篇论文

Smart home voice assistants enable users to conveniently interact with IoT devices and perform Internet searches; however, they also collect the voice input that can carry sensitive personal information about users. Previous papers…

人机交互 · 计算机科学 2024-03-12 Xinhang Ma , Sirui Chen

Confidence scores of automatic speech recognition (ASR) outputs are often inadequately communicated, preventing its seamless integration into analytical workflows. In this paper, we introduce ConFides, a visual analytic system developed in…

人机交互 · 计算机科学 2024-07-26 Sunwoo Ha , Chaehun Lim , R. Jordan Crouser , Alvitta Ottley

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvements over audio-only…

Virtual assistants, also known as intelligent conversational systems such as Google's Virtual Assistant and Apple's Siri, interact with human-like responses to users' queries and finish specific tasks. Meanwhile, existing recommendation…

信息检索 · 计算机科学 2019-01-08 Dimitrios Rafailidis , Yannis Manolopoulos

In the area of Internet of Things (IoT) voice assistants have become an important interface to operate smart speakers, smartphones, and even automobiles. To save power and protect user privacy, voice assistants send commands to the cloud…

机器学习 · 计算机科学 2021-09-22 Yanjiao Chen , Yijie Bai , Richard Mitev , Kaibo Wang , Ahmad-Reza Sadeghi , Wenyuan Xu

Audio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information. However, most existing efforts still focus on audio modality to improve…

音频与语音处理 · 电气工程与系统科学 2023-06-21 Yuchen Hu , Ruizhe Li , Chen Chen , Chengwei Qin , Qiushi Zhu , Eng Siong Chng

A vishing attack is a form of social engineering where attackers use phone calls to deceive individuals into disclosing sensitive information, such as personal data, financial information, or security credentials. Attackers exploit the…

密码学与安全 · 计算机科学 2025-06-17 João Figueiredo , Afonso Carvalho , Daniel Castro , Daniel Gonçalves , Nuno Santos

The performance of speaker-related systems usually degrades heavily in practical applications largely due to the presence of background noise. To improve the robustness of such systems in unknown noisy environments, this paper proposes a…

声音 · 计算机科学 2018-05-04 Siyang Song , Shuimei Zhang , Björn Schuller , Linlin Shen , Michel Valstar

The rapid evolution of Artificial Intelligence (AI)-based Virtual Assistants (VAs) e.g., Google Gemini, ChatGPT, Microsoft Copilot, and High-Flyer Deepseek has turned them into convenient interfaces for managing emerging technologies such…

人工智能 · 计算机科学 2025-05-13 Jennifer Mondragon , Carlos Rubio-Medrano , Gael Cruz , Dvijesh Shastri

This paper describes noisy speech recognition for an augmented reality headset that helps verbal communication within real multiparty conversational environments. A major approach that has actively been studied in simulated environments is…

音频与语音处理 · 电气工程与系统科学 2022-07-18 Yicheng Du , Aditya Arie Nugraha , Kouhei Sekiguchi , Yoshiaki Bando , Mathieu Fontaine , Kazuyoshi Yoshii

Healthcare alert systems (HAS) are undergoing rapid evolution, propelled by advancements in artificial intelligence (AI), Internet of Things (IoT) technologies, and increasing health consciousness. Despite significant progress, a…

计算机与社会 · 计算机科学 2024-08-26 Yulan Gao , Ziqiang Ye , Ming Xiao , Yue Xiao , Dong In Kim

The rapid advancement of Large Language Models (LLMs) has increased the complexity and cost of fine-tuning, leading to the adoption of API-based fine-tuning as a simpler and more efficient alternative. While this method is popular among…

计算与语言 · 计算机科学 2025-03-17 Hongyu Su , Yifeng Gao , Yifan Ding , Xingjun Ma

Voice assistants are now widely available, and to activate them a keyword spotting (KWS) algorithm is used. Modern KWS systems are mainly trained using supervised learning methods and require a large amount of labelled data to achieve a…

音频与语音处理 · 电气工程与系统科学 2024-03-28 Jacob Mørk , Holger Severin Bovbjerg , Gergely Kiss , Zheng-Hua Tan

Personalization of on-device speech recognition (ASR) has seen explosive growth in recent years, largely due to the increasing popularity of personal assistant features on mobile devices and smart home speakers. In this work, we present…

音频与语音处理 · 电气工程与系统科学 2022-06-28 Shaojin Ding , Rajeev Rikhye , Qiao Liang , Yanzhang He , Quan Wang , Arun Narayanan , Tom O'Malley , Ian McGraw

Accurate prediction of the user intent to interact with a voice assistant (VA) on a device (e.g. on the phone) is critical for achieving naturalistic, engaging, and privacy-centric interactions with the VA. To this end, we present a novel…

计算与语言 · 计算机科学 2022-10-24 Pranay Dighe , Prateeth Nayak , Oggi Rudovic , Erik Marchi , Xiaochuan Niu , Ahmed Tewfik

Recent breakthroughs in intelligent speech and digital human technologies have primarily targeted mainstream adult users, often overlooking the distinct vocal patterns and interaction styles of seniors and children. These demographics…

声音 · 计算机科学 2025-07-22 Haiying Xu , Haoze Liu , Mingshi Li , Siyu Cai , Guangxuan Zheng , Yuhuang Jia , Jinghua Zhao , Yong Qin

Voice recognition technology enables the execution of real-world operations through a single voice command. This paper introduces a voice recognition system that involves converting input voice signals into corresponding text using an…

机器人学 · 计算机科学 2023-12-08 Lochan Basyal

Visual Speech Recognition (VSR) aims to infer speech into text depending on lip movements alone. As it focuses on visual information to model the speech, its performance is inherently sensitive to personal lip appearances and movements, and…

计算与语言 · 计算机科学 2024-10-21 Minsu Kim , Hyung-Il Kim , Yong Man Ro

The primary objective of speech enhancement is to reduce background noise while preserving the target's speech. A common dilemma occurs when a speaker is confined to a noisy environment and receives a call with high background and…

声音 · 计算机科学 2023-01-24 Amanda Shu , Hamza Khalid , Haohui Liu , Shikhar Agnihotri , Joseph Konan , Ojas Bhargave

This paper describes the system developed by the NPU team for the 2020 personalized voice trigger challenge. Our submitted system consists of two independently trained subsystems: a small footprint keyword spotting (KWS) system and a…

声音 · 计算机科学 2021-03-01 Jingyong Hou , Li Zhang , Yihui Fu , Qing Wang , Zhanheng Yang , Qijie Shao , Lei Xie