English
Related papers

Related papers: Synesthesia: Detecting Screen Content via Remote A…

200 papers

Speechreading or lipreading is the technique of understanding and getting phonetic features from a speaker's visual features such as movement of lips, face, teeth and tongue. It has a wide range of multimedia applications such as in…

We provide a state-of-the-art analysis of acoustic side channels, cover all the significant academic research in the area, discuss their security implications and countermeasures, and identify areas for future research. We also make an…

Cryptography and Security · Computer Science 2023-08-09 Ping Wang , Shishir Nagaraja , Aurélien Bourquard , Haichang Gao , Jeff Yan

One of the current principal defenses against weaponized synthetic media continues to be the ability of the targeted individual to visually or auditorily recognize AI-generated content when they encounter it. However, as the realism of…

Human-Computer Interaction · Computer Science 2026-04-06 Di Cooke , Abigail Edwards , Sophia Barkoff , Kathryn Kelly

Communication barriers pose significant challenges for individuals with hearing and speech impairments, often limiting their ability to effectively interact in everyday environments. This project introduces a real-time assistive technology…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Brandone Fonya , Clarence Worrell

The increasing prevalence of microphones in everyday devices and the growing reliance on online services have amplified the risk of acoustic side-channel attacks (ASCAs) targeting keyboards. This study explores deep learning techniques,…

Machine Learning · Computer Science 2025-02-20 Jin Hyun Park , Seyyed Ali Ayati , Yichen Cai

Devices capable of detecting and categorizing acoustic scenes have numerous applications such as providing context-aware user experiences. In this paper, we address the task of characterizing acoustic scenes in a workplace setting from…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-12 Arindam Jati , Amrutha Nadarajan , Karel Mundnich , Shrikanth Narayanan

There is strong interest in the generation of synthetic video imagery of people talking for various purposes, including entertainment, communication, training, and advertisement. With the development of deep fake generation models,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Qiaomu Miao , Sinhwa Kang , Stacy Marsella , Steve DiPaola , Chao Wang , Ari Shapiro

Cough monitoring can enable new individual pulmonary health applications. Subject cough event detection is the foundation for continuous cough monitoring. Recently, the rapid growth in smart hearables has opened new opportunities for such…

Sound · Computer Science 2023-03-21 Xiyuxing Zhang , Yuntao Wang , Jingru Zhang , Yaqing Yang , Shwetak Patel , Yuanchun Shi

Acoustic eavesdropping is a privacy risk, but existing attacks rarely work in real outdoor situations where people make phone calls on the move. We present SuperEar, the first portable system that uses acoustic metamaterials to reliably…

Sound · Computer Science 2026-01-19 Zhiyuan Ning , Zhanyong Tang , Juan He , Weizhi Meng , Yuntian Chen , Ji Zhang , Zheng Wang

Recent advances in artificial speech and audio technologies have improved the abilities of deep-fake operators to falsify media and spread malicious misinformation. Anyone with limited coding skills can use freely available speech synthesis…

Most modern smartphones are equipped with a barometer to sample air pressure. Accessing these samples is deemed harmless, hence does not require any permission. In this work, we show, however, that these samples can reveal sensitive…

Cryptography and Security · Computer Science 2020-08-11 Alireza Hafez , Dorsa Nahid , Majid Khabbazian

It has always been a rather tough task to communicate with someone possessing a hearing impairment. One of the most tested ways to establish such a communication is through the use of sign based languages. However, not many people are aware…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Sharanya Mukherjee , Md Hishaam Akhtar , Kannadasan R

Audio scene cartography for real or simulated stereo recordings is presented. This audio scene analysis is performed doing successively: a perceptive 10-subbands analysis, calculation of temporal laws for relative delays and gains between…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-10 Laurent Millot , Gérard Pelé , Mohammed Elliq

When interacting with smart devices such as mobile phones or wearables, the user typically invokes a virtual assistant (VA) by saying a keyword or by pressing a button on the device. However, in many cases, the VA can accidentally be…

Sound · Computer Science 2021-10-12 Ognjen Rudovic , Akanksha Bindal , Vineet Garg , Pramod Simha , Pranay Dighe , Sachin Kajarekar

With the advent of the pandemic, the use of video conferencing platforms as a means of communication has greatly increased and with it, so have the remote opportunities. The deaf and dumb have traditionally faced several issues in…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Kshitij Deshpande , Varad Mashalkar , Kaustubh Mhaisekar , Amaan Naikwadi , Archana Ghotkar

Pharyngeal health plays a vital role in essential human functions such as breathing, swallowing, and vocalization. Early detection of swallowing abnormalities, also known as dysphagia, is crucial for timely intervention. However, current…

Machine Learning · Computer Science 2026-02-04 Jade Chng , Rong Xing , Yunfei Luo , Kristen Linnemeyer-Risser , Tauhidur Rahman , Andrew Yousef , Philip A Weissbrod

The problem of identifying voice commands has always been a challenge due to the presence of noise and variability in speed, pitch, etc. We will compare the efficacies of several neural network architectures for the speech recognition…

Machine Learning · Statistics 2020-11-25 Sanjay Krishna Gouda , Salil Kanetkar , David Harrison , Manfred K Warmuth

Automatic speech recognition systems have created exciting possibilities for applications, however they also enable opportunities for systematic eavesdropping. We propose a method to camouflage a person's voice over-the-air from these…

Sound · Computer Science 2022-02-18 Mia Chiquier , Chengzhi Mao , Carl Vondrick

This paper presents a novel soft tactile skin (STS) technology operating with sound waves. In this innovative approach, the sound waves generated by a speaker travel in channels embedded in a soft membrane and get modulated due to a…

Robotics · Computer Science 2024-03-01 Vishnu Rajendran S , Willow Mandil , Simon Parsons , Amir Ghalamzan E

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…