中文
相关论文

相关论文: NONOTO: A Model-agnostic Web Interface for Interac…

200 篇论文

Recent advancements in video generation have been remarkable, yet many existing methods struggle with issues of consistency and poor text-video alignment. Moreover, the field lacks effective techniques for text-guided video inpainting, a…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Bojia Zi , Shihao Zhao , Xianbiao Qi , Jianan Wang , Yukai Shi , Qianyu Chen , Bin Liang , Kam-Fai Wong , Lei Zhang

Collaborative filtering is one of the most common scenarios and popular research topics in recommender systems. Among existing methods, latent factor models, i.e., learning a specific embedding for each user/item by reconstructing the…

信息检索 · 计算机科学 2022-04-27 Yunfan Wu , Qi Cao , Huawei Shen , Shuchang Tao , Xueqi Cheng

We propose a system that learns from artistic pairings of music and corresponding album cover art. The goal is to 'translate' paintings into music and, in further stages of development, the converse. We aim to deploy this system as an…

声音 · 计算机科学 2020-08-25 Prateek Verma , Constantin Basica , Pamela Davis Kivelson

Generating musical audio directly with neural networks is notoriously difficult because it requires coherently modeling structure at many different timescales. Fortunately, most music is also highly structured and can be represented as…

Deep Learning models have shown very promising results in automatically composing polyphonic music pieces. However, it is very hard to control such models in order to guide the compositions towards a desired goal. We are interested in…

机器学习 · 计算机科学 2021-03-11 Lucas N. Ferreira , Jim Whitehead

We present ASSIST, an object-wise neural radiance field as a panoptic representation for compositional and realistic simulation. Central to our approach is a novel scene node data structure that stores the information of each object in a…

计算机视觉与模式识别 · 计算机科学 2023-11-13 Zhide Zhong , Jiakai Cao , Songen Gu , Sirui Xie , Weibo Gao , Liyi Luo , Zike Yan , Hao Zhao , Guyue Zhou

This paper introduces LIMITER, a gamified digital musical instrument for harnessing and performing microtonal and justly intonated sounds. While microtonality in Western music remains a niche and esoteric system that can be difficult both…

人机交互 · 计算机科学 2025-07-14 Antonis Christou

We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a shared diffusion…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Junyi Chen , Tong He , Zhoujie Fu , Pengfei Wan , Kun Gai , Weicai Ye

Supporting voice commands in applications presents significant benefits to users. However, adding such support to existing GUI-based web apps is effort-consuming with a high learning barrier, as shown in our formative study, due to the lack…

Autoregressive generative transformers are key in music generation, producing coherent compositions but facing challenges in human-machine collaboration. We propose RefinPaint, an iterative technique that improves the sampling process. It…

声音 · 计算机科学 2024-11-12 Pedro Ramoneda , Martin Rocamora , Taketo Akama

This paper is a survey and an analysis of different ways of using deep learning (deep artificial neural networks) to generate musical content. We propose a methodology based on five dimensions for our analysis: Objective - What musical…

声音 · 计算机科学 2019-08-09 Jean-Pierre Briot , Gaëtan Hadjeres , François-David Pachet

Deep generative models have shown success in automatically synthesizing missing image regions using surrounding context. However, users cannot directly decide what content to synthesize with such approaches. We propose an end-to-end network…

计算机视觉与模式识别 · 计算机科学 2018-03-23 Yinan Zhao , Brian Price , Scott Cohen , Danna Gurari

Recurrent Neural Networks (RNNS) are now widely used on sequence generation tasks due to their ability to learn long-range dependencies and to generate sequences of arbitrary length. However, their left-to-right generation procedure only…

人工智能 · 计算机科学 2017-09-20 Gaëtan Hadjeres , Frank Nielsen

With the advancements in denoising diffusion probabilistic models (DDPMs), image inpainting has significantly evolved from merely filling information based on nearby regions to generating content conditioned on various prompts such as text,…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Lingzhi Pan , Tong Zhang , Bingyuan Chen , Qi Zhou , Wei Ke , Sabine Süsstrunk , Mathieu Salzmann

We propose Polyffusion, a diffusion model that generates polyphonic music scores by regarding music as image-like piano roll representations. The model is capable of controllable music generation with two paradigms: internal control and…

声音 · 计算机科学 2023-07-21 Lejun Min , Junyan Jiang , Gus Xia , Jingwei Zhao

Text-guided image inpainting endeavors to generate new content within specified regions of images using textual prompts from users. The primary challenge is to accurately align the inpainted areas with the user-provided prompts while…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Chao Gong , Dong Li , Yingwei Pan , Jingjing Chen , Ting Yao , Tao Mei

We describe a proof-of-principle implementation of a system for drawing melodies that abstracts away from a note-level input representation via melodic contours. The aim is to allow users to express their musical intentions without…

声音 · 计算机科学 2023-05-22 Tashi Namgyal , Peter Flach , Raul Santos-Rodriguez

This paper presents NetWorks (NW), an interactive music generation system that uses a hierarchically clustered scale free network to generate music that ranges from orderly to chaotic. NW was inspired by the Honing Theory of creativity,…

声音 · 计算机科学 2019-07-17 Shawn Bell , Liane Gabora

Generating music is an interesting and challenging problem in the field of machine learning. Mimicking human creativity has been popular in recent years, especially in the field of computer vision and image processing. With the advent of…

声音 · 计算机科学 2020-11-03 Ashish Ranjan , Varun Nagesh Jolly Behera , Motahar Reza

Performance RNN is a machine-learning system designed primarily for the generation of solo piano performances using an event-based (rather than audio) representation. More specifically, Performance RNN is a long short-term memory (LSTM)…

声音 · 计算机科学 2022-02-22 Nicholas Meade , Nicholas Barreyre , Scott C. Lowe , Sageev Oore