English
Related papers

Related papers: Hybrid noise shaping for audio coding using perfec…

200 papers

In wideband sub-Terahertz (sub-THz) massive multiple-input multiple-output (MIMO) communication systems, the beam squint effect manifests as a substantial degradation in array gain. To mitigate the aforementioned beam squint effect, a…

Signal Processing · Electrical Eng. & Systems 2024-10-30 Dang Qua Nguyen , Taejoon Kim

Recent parallel neural text-to-speech (TTS) synthesis methods are able to generate speech with high fidelity while maintaining high performance. However, these systems often lack control over the output prosody, thus restricting the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-30 Shreyas Seshadri , Tuomo Raitio , Dan Castellani , Jiangchuan Li

As the recently proposed voice cloning system, NAUTILUS, is capable of cloning unseen voices using untranscribed speech, we investigate the feasibility of using it to develop a unified cross-lingual TTS/VC system. Cross-lingual speech…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-09 Hieu-Thi Luong , Junichi Yamagishi

Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-01 Yuanyuan Wang , Hangting Chen , Dongchao Yang , Zhiyong Wu , Xixin Wu

Multivariate time series (MTS) classification is foundational to pervasive computing and financial analysis, yet existing multi-scale paradigms are often constrained by suboptimal representation fidelity. We identify two critical…

Machine Learning · Computer Science 2026-05-22 Fan Zhang , Yating Cui , Hua Wang

This paper proposes robust nonlinear transform coding (Robust-NTC), a generalizable digital joint source-channel coding (JSCC) framework that couples variational latent modeling with channel-adaptive transmission. Unlike learning-based JSCC…

Signal Processing · Electrical Eng. & Systems 2026-04-24 Jihun Park , Junyong Shin , Jinsung Park , Yo-Seb Jeon

Accurate Multivariate Time Series (MTS) forecasting is crucial for collaborative design of complex systems, Digital Twin building, and maintenance ahead of time. However, the collaborative industrial environment presents new challenges for…

Machine Learning · Computer Science 2025-12-02 Shaoxun Wang , Xingjun Zhang , Kun Xia , Qianyang Li , Jiawei Cao , Zhendong Tan

We propose a novel regularizer for supervised learning called Conditioning on Noisy Targets (CNT). This approach consists in conditioning the model on a noisy version of the target(s) (e.g., actions in imitation learning or labels in…

Machine Learning · Computer Science 2022-10-28 Alexia Jolicoeur-Martineau , Alex Lamb , Vikas Verma , Aniket Didolkar

We propose a new digital-to-analog converter (DAC) for realizing a synapse circuit of mixed-signal spiking neural networks. We named this circuit "time-domain DAC (TDAC)". This produces weights for converting a digital input code into…

Signal Processing · Electrical Eng. & Systems 2020-01-22 Seiji Uenohara , Kazuyuki Aihara

Speaker embedding is an important front-end module to explore discriminative speaker features for many speech applications where speaker information is needed. Current SOTA backbone networks for speaker embedding are designed to aggregate…

Sound · Computer Science 2022-03-18 Ruiteng Zhang , Jianguo Wei , Xugang Lu , Wenhuan Lu , Di Jin , Junhai Xu , Lin Zhang , Yantao Ji , Jianwu Dang

We consider $N$-way data arrays and low-rank tensor factorizations where the time mode is coded as a sparse linear combination of temporal elements from an over-complete library. Our method, Shape Constrained Tensor Decomposition (SCTD) is…

Machine Learning · Statistics 2016-08-17 Bethany Lusch , Eric C. Chi , J. Nathan Kutz

Time-encoding of continuous-time signals is an alternative sampling paradigm to conventional methods such as Shannon's sampling. In time-encoding, the signal is encoded using a sequence of time instants where an event occurs, and hence fall…

Signal Processing · Electrical Eng. & Systems 2021-09-06 Abijith Jagannath Kamath , Sunil Rudresh , Chandra Sekhar Seelamantula

Deep learning approaches have emerged that aim to transform an audio signal so that it sounds as if it was recorded in the same room as a reference recording, with applications both in audio post-production and augmented reality. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-16 Christian J. Steinmetz , Vamsi Krishna Ithapu , Paul Calamia

Data efficient voice cloning aims at synthesizing target speaker's voice with only a few enrollment samples at hand. To this end, speaker adaptation and speaker encoding are two typical methods based on base model trained from multiple…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-12 Jian Cong , Shan Yang , Lei Xie , Guoqiao Yu , Guanglu Wan

The introduction of 5G has changed the wireless communication industry. Whereas previous generations of cellular technology are mainly based on communication for people, the wireless industry is discovering that 5G may be an era of…

Information Theory · Computer Science 2023-03-09 Bohang Zhang , Zhaoujun Nan , Sheng Zhou , Zhisheng Niu

We present a transformer-based speech-declipping model that effectively recovers clipped signals across a wide range of input signal-to-distortion ratios (SDRs). While recent time-domain deep neural network (DNN)-based declippers have…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-20 Younghoo Kwon , Jung-Woo Choi

Deep learning-based speech enhancement models achieve remarkable performance when test distributions match training conditions, but often degrade when deployed in unpredictable real-world environments with domain shifts. To address this…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-09 Tobias Raichle , Niels Edinger , Bin Yang

Audio classifiers frequently face domain shift, when models trained on one dataset lose accuracy on data recorded in acoustically different conditions. Previous Test-Time Adaptation (TTA) research in speech and sound analysis often…

Sound · Computer Science 2025-11-25 Weichuang Shao , Iman Yi Liao , Tomas Henrique Bode Maul , Tissa Chandesa

Anomaly detection in multi-variate time series (MVTS) data is a huge challenge as it requires simultaneous representation of long term temporal dependencies and correlations across multiple variables. More often, this is solved by breaking…

Machine Learning · Computer Science 2022-02-09 Theivendiram Pranavan , Terence Sim , Arulmurugan Ambikapathi , Savitha Ramasamy

Dynamic mode decomposition (DMD) provides a principled approach to extract physically interpretable spatial modes from time-resolved flow field data, along with a linear model for how the amplitudes of these modes evolve in time. Recently,…

Fluid Dynamics · Physics 2020-07-29 Aditya G. Nair , Benjamin Strom , Bingni W. Brunton , Steven L. Brunton
‹ Prev 1 4 5 6 7 8 10 Next ›