Related papers: Constructing spoke subfactors using the jellyfish …
In this work, we introduce a multi-task transformer for speech deepfake detection, capable of predicting formant trajectories and voicing patterns over time, ultimately classifying speech as real or fake, and highlighting whether its…
We derive a compact analytic formula for a complete basis of conformally invariant tensor structures for three-point functions of conserved operators in arbitrary 4D Lorentz representations. The construction follows directly from a novel…
Subtraction is a powerful technique for creating new bijections from old. Let's reinvent it! While we're at it, let's reinvent division as well.
In this paper we present a simple re-ranking method for Automatic Sentence Simplification based on the noisy channel scheme. Instead of directly computing the best simplification given a complex text, the re-ranking method also considers…
Attractor-based end-to-end diarization is achieving comparable accuracy to the carefully tuned conventional clustering-based methods on challenging datasets. However, the main drawback is that it cannot deal with the case where the number…
PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command-line interface and a simple code structure. This paper…
Parameter-efficient fine-tuning (PEFT) using labeled task data can significantly improve the performance of large language models (LLMs) on the downstream task. However, there are 7000 languages in the world and many of these languages lack…
We propose a novel training algorithm for a multi-speaker neural text-to-speech (TTS) model based on multi-task adversarial training. A conventional generative adversarial network (GAN)-based training algorithm significantly improves the…
This article describes a system for analyzing acoustic data to assist in the diagnosis and classification of children's speech sound disorders (SSDs) using a computer. The analysis concentrated on identifying and categorizing four distinct…
This paper presents the SJTU system for both text-dependent and text-independent tasks in short-duration speaker verification (SdSV) challenge 2021. In this challenge, we explored different strong embedding extractors to extract robust…
Speech signals are complex intermingling of various informative factors, and this information blending makes decoding any of the individual factors extremely difficult. A natural idea is to factorize each speech frame into independent…
Morphological analysis involves predicting the syntactic traits of a word (e.g. {POS: Noun, Case: Acc, Gender: Fem}). Previous work in morphological tagging improves performance for low-resource languages (LRLs) through cross-lingual…
In this paper we further develop the method of quaternion typification of Clifford algebra elements suggested by the author in the previous papers. On the basis of new classification of Clifford algebra elements it is possible to find out…
There is a natural construction which associates to a finitely generated, countable, discrete group $G$ and a 3-cocycle $\omega$ of $G$ an inclusion of II$_1$ factors, the so-called diagonal subfactors (with cocycle). In the case when the…
We investigate the computational efficiency of two stochastic based alternatives to the Sequential Propagator Method used in Lattice QCD calculations of heavy-light semileptonic form factors. In the first method, we replace the sequential…
The triplication method for constructing strong starters in $Z_{3m}$ from starters in $Z_{m}$ (say, a starter of order 21 from a starter of order 7) was proposed by the authors in 2025. The method reduced construction of the particular…
Large Language Models (LLMs) frequently lack domain-specific knowledge and even fine-tuned models tend to hallucinate. Hence, more reliable models that can include external knowledge are needed. We present a pipeline, 4StepFocus, and…
Speech signals are complex composites of various information, including phonetic content, speaker traits, channel effect, etc. Decomposing this complicated mixture into independent factors, i.e., speech factorization, is fundamentally…
Speech information can be roughly decomposed into four components: language content, timbre, pitch, and rhythm. Obtaining disentangled representations of these components is useful in many speech analysis and generation applications.…
It is known that each of the successive quotient groups of the grope and solvable filtrations of the knot concordance group has an infinite rank subgroup. The generating knots of these subgroups are constructed using iterated doubling…