中文
相关论文

相关论文: Soloist: Generating Mixed-Initiative Tutorials fro…

200 篇论文

In this paper, we study the problem of procedure planning in instructional videos. Here, an agent must produce a plausible sequence of actions that can transform the environment from a given start to a desired goal state. When learning…

计算机视觉与模式识别 · 计算机科学 2022-05-06 He Zhao , Isma Hadji , Nikita Dvornik , Konstantinos G. Derpanis , Richard P. Wildes , Allan D. Jepson

Users often take notes for instructional videos to access key knowledge later without revisiting long videos. Automated note generation tools enable users to obtain informative notes efficiently. However, notes generated by existing…

Digital art portfolios serve as impactful mediums for artists to convey their visions, weaving together visuals, audio, interactions, and narratives. However, without technical backgrounds, design students often find it challenging to…

人机交互 · 计算机科学 2023-11-27 Tao Long , Weirui Peng

Music recommendation for videos attracts growing interest in multi-modal research. However, existing systems focus primarily on content compatibility, often ignoring the users' preferences. Their inability to interact with users for further…

机器学习 · 计算机科学 2024-03-12 Zhikang Dong , Bin Chen , Xiulong Liu , Pawel Polak , Peng Zhang

We present a self-supervised learning method to learn audio and video representations. Prior work uses the natural correspondence between audio and video to define a standard cross-modal instance discrimination task, where a model is…

计算机视觉与模式识别 · 计算机科学 2021-03-31 Pedro Morgado , Ishan Misra , Nuno Vasconcelos

This paper presents a deep reinforcement learning algorithm for online accompaniment generation, with potential for real-time interactive human-machine duet improvisation. Different from offline music generation and harmonization, online…

机器学习 · 计算机科学 2020-02-11 Nan Jiang , Sheng Jin , Zhiyao Duan , Changshui Zhang

Drawing inspiration from the notion of cognitive incongruence associated with Stroop's famous experiment, from musical principles, and from the observation that music consumption on an individual basis is becoming increasingly ubiquitous,…

音频与语音处理 · 电气工程与系统科学 2018-11-19 Ishwarya Ananthabhotla , Joseph A. Paradiso

Audio representations for music information retrieval are typically learned via supervised learning in a task-specific fashion. Although effective at producing state-of-the-art results, this scheme lacks flexibility with respect to the…

声音 · 计算机科学 2022-02-18 Ilaria Manco , Emmanouil Benetos , Elio Quinton , Gyorgy Fazekas

Recently deep learning based recommendation systems have been actively explored to solve the cold-start problem using a hybrid approach. However, the majority of previous studies proposed a hybrid model where collaborative filtering and…

信息检索 · 计算机科学 2018-07-19 Jongpil Lee , Kyungyun Lee , Jiyoung Park , Jangyeon Park , Juhan Nam

Text-to-image generative models have made significant advancements in recent years; however, accurately capturing intricate details in textual prompts-such as entity missing, attribute binding errors, and incorrect relationships remains a…

Generating lifelike conversational avatars requires modeling not just isolated speakers, but the dynamic, reciprocal interaction of speaking and listening. However, modeling the listener is exceptionally challenging: direct audio-driven…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Xuangeng Chu , Ruicong Liu , Yifei Huang , Yun Liu , Yichen Peng , Bo Zheng

Recent vision-language models outperform vision-only models on many image classification tasks. However, because of the absence of paired text/image descriptions, it remains difficult to fine-tune these models for fine-grained image…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Kathleen M. Lewis , Emily Mu , Adrian V. Dalca , John Guttag

In this work, we propose Mutual Forcing, a framework for fast autoregressive audio-video generation with long-horizon audio-video synchronization. Our approach addresses two key challenges: joint audio-video modeling and fast autoregressive…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Yupeng Zhou , Lianghua Huang , Zhifan Wu , Jiabao Wang , Yupeng Shi , Biao Jiang , Daquan Zhou , Yu Liu , Ming-Ming Cheng , Qibin Hou

Independent learners often struggle with sustaining focus and emotional regulation in unstructured or distracting settings. Although some rely on ambient aids such as music, ASMR, or visual backgrounds to support concentration, these tools…

人工智能 · 计算机科学 2025-05-07 George Xi Wang , Jingying Deng , Safinah Ali

Deep learning based visual to sound generation systems essentially need to be developed particularly considering the synchronicity aspects of visual and audio features with time. In this research we introduce a novel task of guiding a class…

机器学习 · 计算机科学 2021-07-21 Sanchita Ghose , John J. Prevost

Generative AI is revolutionizing content creation and has the potential to enable real-time, personalized educational experiences. We investigated the effectiveness of converting textbook chapters into AI-generated podcasts and explored the…

人机交互 · 计算机科学 2024-10-30 Tiffany D. Do , Usama Bin Shafqat , Elsie Ling , Nikhil Sarda

We propose a unified model for three inter-related tasks: 1) to \textit{separate} individual sound sources from a mixed music audio, 2) to \textit{transcribe} each sound source to MIDI notes, and 3) to\textit{ synthesize} new pieces based…

声音 · 计算机科学 2021-08-10 Liwei Lin , Qiuqiang Kong , Junyan Jiang , Gus Xia

Recent developments in deep learning have resulted in code-generation models that produce source code from natural language and code-based prompts with high accuracy. This is likely to have profound effects in the classroom, where novices…

Automatic transcription of guitar strumming is an underrepresented and challenging task in Music Information Retrieval (MIR), particularly for extracting both strumming directions and chord progressions from audio signals. While existing…

声音 · 计算机科学 2025-08-12 Sebastian Murgul , Johannes Schimper , Michael Heizmann

Video description entails automatically generating coherent natural language sentences that narrate the content of a given video. We introduce CLearViD, a transformer-based model for video description generation that leverages curriculum…

计算机视觉与模式识别 · 计算机科学 2023-11-09 Cheng-Yu Chuang , Pooyan Fazli