English
Related papers

Related papers: Bayesian Semiparametric Markov Renewal Mixed Model…

200 papers

In this study, we propose the global context guided channel and time-frequency transformations to model the long-range, non-local time-frequency dependencies and channel variances in speaker representations. We use the global context…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-10 Wei Xia , John H. L. Hansen

A mixed sample data augmentation strategy is proposed to enhance the performance of models on audio scene classification, sound event classification, and speech enhancement tasks. While there have been several augmentation methods shown to…

Sound · Computer Science 2021-08-09 Gwantae Kim , David K. Han , Hanseok Ko

With read-aloud speech synthesis achieving high naturalness scores, there is a growing research interest in synthesising spontaneous speech. However, human spontaneous face-to-face conversation has both spoken and non-verbal aspects (here,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-15 Shivam Mehta , Siyang Wang , Simon Alexanderson , Jonas Beskow , Éva Székely , Gustav Eje Henter

We propose a deep generative Markov State Model (DeepGenMSM) learning framework for inference of metastable dynamical systems and prediction of trajectories. After unsupervised training on time series data, the model contains (i) a…

Machine Learning · Statistics 2019-01-14 Hao Wu , Andreas Mardt , Luca Pasquali , Frank Noe

Speaker diarization accuracy can be affected by both acoustics and conversation characteristics. Determining the cause of diarization errors is difficult because speaker voice acoustics and conversation structure co-vary, and the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-11 Scott Seyfarth , Sundararajan Srinivasan , Katrin Kirchhoff

This paper proposes a hierarchical spatial-temporal model for modelling the spectrograms of animal calls. The motivation stems from analyzing recordings of the so-called grunt calls emitted by various lemur species. Our goal is to identify…

The replicator-mutator dynamic was originally derived to model the evolution of language, and since the model was derived in such a general manner, it has been applied to the dynamics of social behavior and decision making in multi-agent…

Probability · Mathematics 2020-05-05 Andrew Vlasic

This paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-16 Yukiya Hono , Kei Hashimoto , Yoshihiko Nankaku , Keiichi Tokuda

In complex ecosystems such as microbial communities, there is constant ecological and evolutionary feedback between the residing species and the environment occurring on concurrent timescales. Species respond and adapt to their surroundings…

Populations and Evolution · Quantitative Biology 2023-10-17 Jim Wu , David J. Schwab , Trevor GrandPre

This paper proposes a new approach to duration modelling for statistical parametric speech synthesis in which a recurrent statistical model is trained to output a phone transition probability at each timestep (acoustic frame). Unlike…

Computation and Language · Computer Science 2020-07-28 Srikanth Ronanki , Oliver Watts , Simon King , Gustav Eje Henter

We study the role played by noise on the QW introduced in [1], a 1D model that is inspired by a two particle interacting QW. The noise is introduced by a random change in the value of the phase during the evolution, from a constant…

Quantum Physics · Physics 2017-03-29 Ivan Marquez-Martin , Giuseppe Di Molfetta , Armando Perez

The processes leading to change in languages are manifold. In order to reduce ambiguity in the transmission of information, agreement on a set of conventions for recurring problems is favored. In addition to that, speakers tend to use…

Physics and Society · Physics 2015-06-17 Cristina-Maria Pop , Erwin Frey

Linguistic knowledge plays a crucial role in spoken language comprehension. It provides essential semantic and syntactic context for speech perception in noisy environments. However, most speech enhancement (SE) methods predominantly rely…

Computation and Language · Computer Science 2025-03-11 Kuo-Hsuan Hung , Xugang Lu , Szu-Wei Fu , Huan-Hsin Tseng , Hsin-Yi Lin , Chii-Wann Lin , Yu Tsao

Markov community models have been applied to sessile organisms because such models facilitate estimation of transition probabilities by tracking species occupancy at many fixed observation points over multiple periods of time. Estimation of…

Applications · Statistics 2017-06-12 Keiichi Fukaya , J. Andrew Royle , Takehiro Okuda , Masahiro Nakaoka , Takashi Noda

Large language models have the ability to generate text that mimics patterns in their inputs. We introduce a simple Markov Chain sequence modeling task in order to study how this in-context learning (ICL) capability emerges. In our setting,…

Machine Learning · Computer Science 2024-02-20 Benjamin L. Edelman , Ezra Edelman , Surbhi Goel , Eran Malach , Nikolaos Tsilivis

The recently-proposed mixture invariant training (MixIT) is an unsupervised method for training single-channel sound separation models in the sense that it does not require ground-truth isolated reference sources. In this paper, we…

Sound · Computer Science 2021-10-22 Aswin Sivaraman , Scott Wisdom , Hakan Erdogan , John R. Hershey

Speech enhancement in hearing aids remains a difficult task in nonstationary acoustic environments, mainly because current signal processing algorithms rely on fixed, manually tuned parameters that cannot adapt in situ to different users or…

Speech Emotion Conversion aims to modify the emotion expressed in input speech while preserving lexical content and speaker identity. Recently, generative modeling approaches have shown promising results in changing local acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-18 Navin Raj Prabhu , Danilo de Oliveira , Nale Lehmann-Willenbrock , Timo Gerkmann

Marmoset monkeys exhibit complex vocal communication, challenging the view that nonhuman primates vocal communication is entirely innate, and show similar features of human speech, such as vocal labeling of others and turn-taking. Studying…

Computation and Language · Computer Science 2025-09-16 Talia Sternberg , Michael London , David Omer , Yossi Adi

The neutral theory of genetic and linguistic evolution holds that the relative frequencies of variants evolve by random drift. Neutral evolution remains a plausible null model of language change. In this paper we provide evidence against…

Physics and Society · Physics 2020-05-18 James Burridge , Tamsin Blaxter