English
Related papers

Related papers: The Concatenator: A Bayesian Approach To Real Time…

200 papers

This presentation introduces a self-supervised learning approach to the synthesis of new video clips from old ones, with several new key elements for improved spatial resolution and realism: It conditions the synthesis process on contextual…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Guillaume Le Moing , Jean Ponce , Cordelia Schmid

This paper studies deep network architectures to address the problem of video classification. A multi-stream framework is proposed to fully utilize the rich multimodal information in videos. Specifically, we first train three Convolutional…

Computer Vision and Pattern Recognition · Computer Science 2015-11-12 Zuxuan Wu , Yu-Gang Jiang , Xi Wang , Hao Ye , Xiangyang Xue , Jun Wang

Bayesian filtering approximates the true underlying behavior of a time-varying system by inverting an explicit generative model to convert noisy measurements into state estimates. This process typically requires either storage, inversion,…

Machine Learning · Computer Science 2023-11-20 Gianluca M. Bencomo , Jake C. Snell , Thomas L. Griffiths

Performance-score synchronization is an integral task in signal processing, which entails generating an accurate mapping between an audio recording of a performance and the corresponding musical score. Traditional synchronization methods…

Sound · Computer Science 2022-04-20 Ruchit Agrawal , Daniel Wolff , Simon Dixon

While Transformer has become the de-facto standard for speech, modeling upon the fine-grained frame-level features remains an open challenge of capturing long-distance dependencies and distributing the attention weights. We propose…

Computation and Language · Computer Science 2023-05-30 Chen Xu , Yuhao Zhang , Chengbo Jiao , Xiaoqian Liu , Chi Hu , Xin Zeng , Tong Xiao , Anxiang Ma , Huizhen Wang , JingBo Zhu

Real-time satellite imaging has a central role in monitoring, detecting and estimating the intensity of key natural phenomena such as floods, earthquakes, etc. One important constraint of satellite imaging is the trade-off between…

Image and Video Processing · Electrical Eng. & Systems 2023-01-09 Haoqing Li , Bhavya Duvvuri , Ricardo Borsoi , Tales Imbiriba , Edward Beighley , Deniz Erdogmus , Pau Closas

The remarkable capabilities of pretrained image diffusion models have been utilized not only for generating fixed-size images but also for creating panoramas. However, naive stitching of multiple images often results in visible seams.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Yuseung Lee , Kunho Kim , Hyunjin Kim , Minhyuk Sung

Video editing and generation methods often rely on pre-trained image-based diffusion models. During the diffusion process, however, the reliance on rudimentary noise sampling techniques that do not preserve correlations present in…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Pascal Chang , Jingwei Tang , Markus Gross , Vinicius C. Azevedo

Recent advances in music generation produce impressive samples, however, practical creation still lacks two key capabilities: composer-style structural editing and minute-scale coherence. We present MusicWeaver, a framework for generating…

Sound · Computer Science 2026-01-30 Xuanchen Wang , Heng Wang , Weidong Cai

Recent advancements in video generation have seen a shift towards unified, transformer-based foundation models that can handle multiple conditional inputs in-context. However, these models have primarily focused on modalities like text,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Wenze Liu , Weicai Ye , Minghong Cai , Quande Liu , Xintao Wang , Xiangyu Yue

High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and contextual accuracy, primarily due to their reliance on limited…

Sound · Computer Science 2025-06-03 Shunian Chen , Xinyuan Xie , Zheshu Chen , Liyan Zhao , Owen Lee , Zhan Su , Qilin Sun , Benyou Wang

This paper introduces Chord Colourizer, a near real-time system that detects the musical key of an audio signal and visually represents it through a novel graphical user interface (GUI). The system assigns colours to musical notes based on…

Human-Computer Interaction · Computer Science 2025-10-14 Paul Haimes

In this paper, we present a novel audio synthesizer, CAESynth, based on a conditional autoencoder. CAESynth synthesizes timbre in real-time by interpolating the reference sounds in their shared latent feature space, while controlling a…

Sound · Computer Science 2021-11-10 Aaron Valero Puche , Sukhan Lee

Efficiently compressing high-dimensional audio signals into a compact and informative latent space is crucial for various tasks, including generative modeling and music information retrieval (MIR). Existing audio autoencoders, however,…

Sound · Computer Science 2025-01-30 Marco Pasini , Stefan Lattner , George Fazekas

Auditory perception involves cues in the monaural auditory pathways as well as binaural cues based on differences between the ears. So far auditory models have often focused on either monaural or binaural experiments in isolation. Although…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-24 Thomas Biberger , Stephan D. Ewert

In view of the current availability and variety of measured data, there is an increasing demand for powerful signal processing tools that can cope successfully with the associated problems that often arise when data are being analysed. In…

Data Analysis, Statistics and Probability · Physics 2014-12-16 Tomislav Stankovski , Andrea Duggento , Peter V. E. McClintock , Aneta Stefanovska

This paper proposes a simple and effective approach for automatic recognition of Cued Speech (CS), a visual communication tool that helps people with hearing impairment to understand spoken language with the help of hand gestures that can…

Computation and Language · Computer Science 2022-04-12 Sanjana Sankar , Denis Beautemps , Thomas Hueber

Identifying musical instruments in polyphonic music recordings is a challenging but important problem in the field of music information retrieval. It enables music search by instrument, helps recognize musical genres, or can make music…

Sound · Computer Science 2016-12-28 Yoonchang Han , Jaehun Kim , Kyogu Lee

In this paper, we propose an effective two-stage approach named Grounded-Dreamer to generate 3D assets that can accurately follow complex, compositional text prompts while achieving high fidelity by using a pre-trained multi-view diffusion…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Xiaolong Li , Jiawei Mo , Ying Wang , Chethan Parameshwara , Xiaohan Fei , Ashwin Swaminathan , CJ Taylor , Zhuowen Tu , Paolo Favaro , Stefano Soatto

In the realm of music AI, arranging rich and structured multi-track accompaniments from a simple lead sheet presents significant challenges. Such challenges include maintaining track cohesion, ensuring long-term coherence, and optimizing…

Sound · Computer Science 2024-11-26 Jingwei Zhao , Gus Xia , Ziyu Wang , Ye Wang
‹ Prev 1 8 9 10 Next ›