中文
相关论文

相关论文: Soloist: Generating Mixed-Initiative Tutorials fro…

200 篇论文

The success of Large Language Models (LLMs) has significantly propelled the research of video understanding. To harvest the benefits of well-trained expert models (i.e., tools), video LLMs prioritize the exploration of tool usage…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Yuyang Liu , Meng Cao , Xinyuan Shi , Xiaondan Liang

Humans have been developing and playing musical instruments for millennia. With technological advancements, instruments were becoming ever more sophisticated. In recent decades computer-supported innovations have also been introduced in…

人机交互 · 计算机科学 2022-11-04 Jordan Aiko Deja , Sven Mayer , Klen Čopič Pucihar , Matjaž Kljun

Supervised deep learning approaches to underdetermined audio source separation achieve state-of-the-art performance but require a dataset of mixtures along with their corresponding isolated source signals. Such datasets can be extremely…

Tutorial videos are a valuable resource for people looking to learn new tasks. People often learn these skills by viewing multiple tutorial videos to get an overall understanding of a task by looking at different approaches to achieve the…

人机交互 · 计算机科学 2025-03-28 Saelyne Yang , Anh Truong , Juho Kim , Dingzeyu Li

We present SingSong, a system that generates instrumental music to accompany input vocals, potentially offering musicians and non-musicians alike an intuitive new way to create music featuring their own voice. To accomplish this, we build…

Fully-supervised models for source separation are trained on parallel mixture-source data and are currently state-of-the-art. However, such parallel data is often difficult to obtain, and it is cumbersome to adapt trained models to mixtures…

音频与语音处理 · 电气工程与系统科学 2022-11-30 Ge Zhu , Jordan Darefsky , Fei Jiang , Anton Selitskiy , Zhiyao Duan

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce MultiFoley, a model…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Ziyang Chen , Prem Seetharaman , Bryan Russell , Oriol Nieto , David Bourgin , Andrew Owens , Justin Salamon

In this paper, we introduce a novel semi-supervised learning framework for end-to-end speech separation. The proposed method first uses mixtures of unseparated sources and the mixture invariant training (MixIT) criterion to train a teacher…

声音 · 计算机科学 2021-09-10 Jisi Zhang , Catalin Zorila , Rama Doddipatla , Jon Barker

In recent years, the guitar has received increased attention from the music information retrieval (MIR) community driven by the challenges posed by its diverse playing techniques and sonic characteristics. Mainly fueled by deep learning…

声音 · 计算机科学 2025-09-30 Jackson Loth , Pedro Sarmento , Saurjya Sarkar , Zixun Guo , Mathieu Barthet , Mark Sandler

We introduce Generative Lecture, a concept that makes existing lecture videos interactive through generative AI and AI clone instructors. By leveraging interactive avatars powered by HeyGen, ElevenLabs, and GPT-5, we embed an AI instructor…

人机交互 · 计算机科学 2025-12-30 Hye-Young Jo , Ada Yi Zhao , Xiaoan Liu , Ryo Suzuki

Multimodal Federated Learning frequently encounters challenges of client modality heterogeneity, leading to undesired performances for secondary modality in multimodal learning. It is particularly prevalent in audiovisual learning, with…

音频与语音处理 · 电气工程与系统科学 2024-08-29 Tiantian Feng , Tuo Zhang , Salman Avestimehr , Shrikanth S. Narayanan

Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work that studies the interaction between audio and visual…

音频与语音处理 · 电气工程与系统科学 2021-03-19 Abhinav Shukla , Stavros Petridis , Maja Pantic

In this paper, we introduce Foley Music, a system that can synthesize plausible music for a silent video clip about people playing musical instruments. We first identify two key intermediate representations for a successful video to music…

计算机视觉与模式识别 · 计算机科学 2020-07-22 Chuang Gan , Deng Huang , Peihao Chen , Joshua B. Tenenbaum , Antonio Torralba

Finetuning large language models with a variety of instruction-response pairs has enhanced their capability to understand and follow instructions. Current instruction tuning primarily relies on teacher models or human intervention to…

计算与语言 · 计算机科学 2025-06-06 Ming Li , Pei Chen , Chenguang Wang , Hongyu Zhao , Yijun Liang , Yupeng Hou , Fuxiao Liu , Tianyi Zhou

Creation of images using generative adversarial networks has been widely adapted into multi-modal regime with the advent of multi-modal representation models pre-trained on large corpus. Various modalities sharing a common representation…

声音 · 计算机科学 2022-06-10 Yoonjeon Kim , Joel Jang , Sumin Shin

YouTube users looking for instructions for a specific task may spend a long time browsing content trying to find the right video that matches their needs. Creating a visual summary (abridged version of a video) provides viewers with a quick…

计算机视觉与模式识别 · 计算机科学 2022-08-16 Medhini Narasimhan , Arsha Nagrani , Chen Sun , Michael Rubinstein , Trevor Darrell , Anna Rohrbach , Cordelia Schmid

Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance,…

多媒体 · 计算机科学 2026-03-03 Jiadong Wang , Ke Zhang , Xinyuan Qian , Ruijie Tao , Haizhou Li , Björn Schuller

A flexible recommendation and retrieval system requires music similarity in terms of multiple partial elements of musical pieces to allow users to select the element they want to focus on. A method for music similarity learning using…

声音 · 计算机科学 2025-07-18 Yuka Hashizume , Li Li , Atsushi Miyashita , Tomoki Toda

We propose CoLyricist, an AI-assisted lyric writing tool designed to support the typical workflows of experienced lyricists and enhance their creative efficiency. While lyricists have unique processes, many follow common stages. Tools that…

人机交互 · 计算机科学 2026-02-27 Masahiro Yoshida , Bingxuan Li , Songyan Zhao , Qinyi Zhou , Shiwei Hu , Xiang Anthony Chen , Nanyun Peng

Multimedia recommendation, which incorporates various modalities (e.g., images, texts, etc.) into user or item representation to improve recommendation quality, and self-supervised learning carries multimedia recommendation to a plateau of…

信息检索 · 计算机科学 2024-12-18 Zhuangzhuang He , Zihan Wang , Yonghui Yang , Haoyue Bai , Le Wu