中文
相关论文

相关论文: Glottal Source Estimation using an Automatic Chirp…

200 篇论文

For audio source separation applications, it is common to estimate the magnitude of the short-time Fourier transform (STFT) of each source. In order to further synthesizing time-domain signals, it is necessary to recover the phase of the…

声音 · 计算机科学 2018-02-28 Paul Magron , Roland Badeau , Bertrand David

Recent studies have introduced various approaches for prompt-tuning black-box vision-language models, referred to as black-box prompt-tuning (BBPT). While BBPT has demonstrated considerable potential, it is often found that many existing…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Seonghwan Park , Jaehyeon Jeong , Yongjun Kim , Jaeho Lee , Namhoon Lee

Spiral acquisitions are preferred in real-time MRI because of their efficiency, which has made it possible to capture vocal tract dynamics during natural speech. A fundamental limitation of spirals is blurring and signal loss due to…

图像与视频处理 · 电气工程与系统科学 2021-02-16 Yongwan Lim , Shrikanth S. Narayanan , Krishna S. Nayak

Jointly Gaussian memoryless sources are observed at N distinct terminals. The goal is to efficiently encode the observations in a distributed fashion so as to enable reconstruction of any one of the observations, say the first one, at the…

信息论 · 计算机科学 2008-05-14 Saurabha Tavildar , Pramod Viswanath , Aaron B. Wagner

Recently, denoising diffusion probabilistic models and generative score matching have shown high potential in modelling complex data distributions while stochastic calculus has provided a unified point of view on these techniques allowing…

机器学习 · 计算机科学 2021-08-06 Vadim Popov , Ivan Vovk , Vladimir Gogoryan , Tasnima Sadekova , Mikhail Kudinov

Pre-aspiration is defined as the period of glottal friction occurring in sequences of vocalic/consonantal sonorants and phonetically voiceless obstruents. We propose two machine learning methods for automatic measurement of pre-aspiration…

计算与语言 · 计算机科学 2017-06-16 Yaniv Sheena , Míša Hejná , Yossi Adi , Joseph Keshet

Synthesized speech is common today due to the prevalence of virtual assistants, easy-to-use tools for generating and modifying speech signals, and remote work practices. Synthesized speech can also be used for nefarious purposes, including…

声音 · 计算机科学 2022-05-05 Emily R. Bartusiak , Edward J. Delp

The center of gravity $x_{g}= \sum_{i}E_{i} x_{i}/\sum_{i} E_{i}$ as an algorithm for position measurements is carefully analyzed. Many mathematical consequences of discretization are extracted. The origin of the systematic error of the…

仪器与探测器 · 物理学 2019-08-14 Gregorio Landi

The extended source effect on microlensing magnification is non-negligible and must be taken into account for in an analysis of microlensing. However, the evaluation of the extended source magnification is numerically expensive because it…

天体物理仪器与方法 · 物理学 2022-10-05 Sunao Sugiyama

The Regge-Wheeler-Zerilli (RWZ) wave-equation describes Schwarzschild-Droste black hole perturbations. The source term contains a Dirac distribution and its derivative. We have previously designed a method of integration in time domain. It…

广义相对论与量子宇宙学 · 物理学 2016-03-22 P. Ritter , S. Aoudia , A. Spallicci , S. Cordier

Various XAI attribution methods have been recently proposed for the transformer architecture, allowing for insights into the decision-making process of large language models by assigning importance scores to input tokens and intermediate…

计算与语言 · 计算机科学 2025-02-25 Leila Arras , Bruno Puri , Patrick Kahardipraja , Sebastian Lapuschkin , Wojciech Samek

A blind source separation method is described to extract sources from data mixtures where the underlying sources are assumed to be sparse and uncorrelated. The approach used is to detect and analyse segments of time where one source exists…

信号处理 · 电气工程与系统科学 2018-02-06 Malcolm Woolfson

We present various improvements to the deformation method for computing the zeta function of smooth projective hypersurfaces over finite fields using $p$-adic cohomology. This includes new bounds for the $p$-adic and $t$-adic precisions…

数论 · 数学 2014-09-11 Sebastian Pancratz , Jan Tuitman

In this work, we investigate the partition function of 2d CFT under root-$T\bar{T}$ deformation. We demonstrate that the deformed partition function satisfies a flow equation. At large central charge sector, the deformed partition function…

高能物理 - 理论 · 物理学 2025-12-03 Miao He

As a challenging vision-language task, Zero-Shot Composed Image Retrieval (ZS-CIR) is designed to retrieve target images using bi-modal (image+text) queries. Typical ZS-CIR methods employ an inversion network to generate pseudo-word tokens…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Haiwen Li , Fei Su , Zhicheng Zhao

We study an efficient dynamic blind source separation algorithm of convolutive sound mixtures based on updating statistical information in the frequency domain, andminimizing the support of time domain demixing filters by a weighted least…

统计理论 · 数学 2007-05-23 Jie Liu , Jack Xin , Yingyong Qi

Although large language models can be prompted for both zero- and few-shot learning, performance drops significantly when no demonstrations are available. In this paper, we introduce Z-ICL, a new zero-shot method that closes the gap by…

计算与语言 · 计算机科学 2023-06-06 Xinxi Lyu , Sewon Min , Iz Beltagy , Luke Zettlemoyer , Hannaneh Hajishirzi

Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-art end-to-end text-to-speech synthesis system, to uncover…

计算与语言 · 计算机科学 2018-08-07 Daisy Stanton , Yuxuan Wang , RJ Skerry-Ryan

Audio tagging aims to assign predefined tags to audio clips to indicate the class information of audio events. Sequential audio tagging (SAT) means detecting both the class information of audio events, and the order in which they occur…

声音 · 计算机科学 2022-10-25 Yuanbo Hou , Yun Wang , Wenwu Wang , Dick Botteldooren

A geometrically-motivated method for primary-ambient decomposition is proposed and evaluated in an up-mixing application. The method consists of two steps, accommodating a particularly intuitive explanation. The first step consists of…

音频与语音处理 · 电气工程与系统科学 2022-06-07 Jouni Paulus , Matteo Torcoli