English
Related papers

Related papers: Real time spectrogram inversion on mobile phone

200 papers

Recent advances in text-to-speech (TTS) synthesis, such as Tacotron and WaveRNN, have made it possible to construct a fully neural network based TTS system, by coupling the two components together. Such a system is conceptually simple as it…

We propose a novel approach for time-scale modification of audio signals. Unlike traditional methods that rely on the framing technique or the short-time Fourier transform to preserve the frequency during temporal stretching, our neural…

Sound · Computer Science 2023-10-09 Ernie Chu , Ju-Ting Chen , Chia-Ping Chen

Modern wake word detection systems usually rely on neural networks for acoustic modeling. Transformers has recently shown superior performance over LSTM and convolutional networks in various sequence modeling tasks with their better…

Computation and Language · Computer Science 2021-02-10 Yiming Wang , Hang Lv , Daniel Povey , Lei Xie , Sanjeev Khudanpur

Streaming reinforcement learning has emerged as an online learning paradigm that conforms to the restrictions of natural learning agents that process data incrementally, i.e. with a batch size of 1 and no replay buffer. While streaming RL…

Machine Learning · Computer Science 2026-05-26 Noah Farr , Aryaman Reddi , Carlo D'Eramo , Jan Peters

Recently Transformer-based hyperspectral image (HSI) change detection methods have shown remarkable performance. Nevertheless, existing attention mechanisms in Transformers have limitations in local feature representation. To address this…

Image and Video Processing · Electrical Eng. & Systems 2024-11-22 Ziyi Wang , Feng Gao , Junyu Dong , Qian Du

We introduce Slam, a recipe for training high-quality Speech Language Models (SLMs) on a single academic GPU in 24 hours. We do so through empirical analysis of model initialisation and architecture, synthetic training data, preference…

Machine Learning · Computer Science 2025-05-23 Gallil Maimon , Avishai Elmakies , Yossi Adi

We introduce Glinthawk, an architecture for offline Large Language Model (LLM) inference. By leveraging a two-tiered structure, Glinthawk optimizes the utilization of the high-end accelerators ("Tier 1") by offloading the attention…

Machine Learning · Computer Science 2025-02-12 Pouya Hamadanian , Sadjad Fouladi

A digital video is a collection of individual frames, while streaming the video the scene utilized the time slice for each frame. High refresh rate and high frame rate is the demand of all high technology applications. The action tracking…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Rishik Mishra , Neeraj Gupta , Nitya Shukla

Generating an image from a given text description has two goals: visual realism and semantic consistency. Although significant progress has been made in generating high-quality and visually realistic images using generative adversarial…

Computation and Language · Computer Science 2019-03-15 Tingting Qiao , Jing Zhang , Duanqing Xu , Dacheng Tao

Smartphone sensing offers an unobtrusive and scalable way to track daily behaviors linked to mental health, capturing changes in sleep, mobility, and phone use that often precede symptoms of stress, anxiety, or depression. While most prior…

Machine Learning · Computer Science 2026-01-14 Kaidong Feng , Zhu Sun , Roy Ka-Wei Lee , Xun Jiang , Yin-Leng Theng , Yi Ding

Real time sensor based applications in pervasive computing require edge deployable models to ensure low latency privacy and efficient interaction. A prime example is sensor based human activity recognition where models must balance accuracy…

Machine Learning · Computer Science 2026-03-30 Deepika Gurung , Lala Shakti Swarup Ray , Mengxi Liu , Bo Zhou , Paul Lukowicz

We present FlashLips, a two-stage, mask-free lip-sync system that decouples lips control from rendering and achieves real-time performance, with our U-Net variant running at over 100 FPS on a single GPU, while matching the visual quality of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Andreas Zinonos , Michał Stypułkowski , Antoni Bigata , Stavros Petridis , Maja Pantic , Nikita Drobyshev

The exploration of the latent space in StyleGANs and GAN inversion exemplify impressive real-world image editing, yet the trade-off between reconstruction quality and editing quality remains an open problem. In this study, we revisit…

Computer Vision and Pattern Recognition · Computer Science 2023-06-02 Kai Katsumata , Duc Minh Vo , Bei Liu , Hideki Nakayama

We show that pre-trained Generative Adversarial Networks (GANs), e.g., StyleGAN, can be used as a latent bank to improve the restoration quality of large-factor image super-resolution (SR). While most existing SR approaches attempt to…

Computer Vision and Pattern Recognition · Computer Science 2020-12-02 Kelvin C. K. Chan , Xintao Wang , Xiangyu Xu , Jinwei Gu , Chen Change Loy

Moment retrieval (MR) and highlight detection (HD) aim to identify relevant moments and highlights in video from corresponding natural language query. Large language models (LLMs) have demonstrated proficiency in various computer vision…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Yunzhuo Sun , Yifang Xu , Zien Xie , Yukun Shu , Sidan Du

Signal reconstruction from its mel-spectrogram is known as mel-spectrogram inversion and has many applications, including speech and foley sound synthesis. In this paper, we propose a mel-spectrogram inversion method based on a rigorous…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Yoshiki Masuyama , Natsuki Ueno , Nobutaka Ono

We propose SLARM, a feed-forward model that unifies dynamic scene reconstruction, semantic understanding, and real-time streaming inference. SLARM captures complex, non-uniform motion through higher-order motion modeling, trained solely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Zhicheng Qiu , Jiarui Meng , Tong-an Luo , Yican Huang , Xuan Feng , Xuanfu Li , ZHan Xu

Self-Supervised Learning (SSL) has demonstrated strong performance in speech processing, particularly in automatic speech recognition. In this paper, we explore an SSL pretraining framework that leverages masked language modeling with…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Aleksandr Kutsakov , Alexandr Maximenko , Georgii Gospodinov , Pavel Bogomolov , Fyodor Minkin

Building Free-Viewpoint Videos in a streaming manner offers the advantage of rapid responsiveness compared to offline training methods, greatly enhancing user experience. However, current streaming approaches face challenges of high…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Jinbo Yan , Rui Peng , Zhiyan Wang , Luyang Tang , Jiayu Yang , Jie Liang , Jiahao Wu , Ronggang Wang

Modern Gaussian Splatting methods have proven highly effective for real-time photorealistic rendering of 3D scenes. However, integrating semantic information into this representation remains a significant challenge, especially in…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Roman Titkov , Egor Zubkov , Dmitry Yudin , Jaafar Mahmoud , Malik Mohrat , Gennady Sidorov