English
Related papers

Related papers: Music to Dance as Language Translation using Seque…

200 papers

The paper describes results on two components of a research program focused on motion-based communication mediated by the dynamics of a control system. Specifically we are interested in how mobile agents engaged in a shared activity such as…

Systems and Control · Computer Science 2011-09-29 J. Baillieul , K. Özcimder

Recent progress in text-based Large Language Models (LLMs) and their extended ability to process multi-modal sensory data have led us to explore their applicability in addressing music information retrieval (MIR) challenges. In this paper,…

Information Retrieval · Computer Science 2025-01-24 Kun Fang , Ziyu Wang , Gus Xia , Ichiro Fujinaga

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains challenging and untouched so far. Fortunately, human motion…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Biao Jiang , Xin Chen , Wen Liu , Jingyi Yu , Gang Yu , Tao Chen

Automatic music transcription (AMT) aims to convert raw audio to symbolic music representation. As a fundamental problem of music information retrieval (MIR), AMT is considered a difficult task even for trained human experts due to overlap…

Sound · Computer Science 2023-02-28 Shenli Yuan , Lingjie Kong , Jiushuang Guo

Interaction between humanoids involves bidirectional and nonverbal reactivity, coordination and synchrony. Toward socially aware robots and interactive virtual agents, we present SalsaAgent, a language model that generates expressive,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Payam Jome Yazdian , Zoe Stanley , Angelica Lim

Audio-to-score alignment (A2SA) is a multimodal task consisting in the alignment of audio signals to music scores. Recent literature confirms the benefits of Automatic Music Transcription (AMT) for A2SA at the frame-level. In this work, we…

Sound · Computer Science 2022-01-03 Federico Simonetta , Stavros Ntalampiras , Federico Avanzini

Generating realistic dyadic human motion from text descriptions presents significant challenges, particularly for extended interactions that exceed typical training sequence lengths. While recent transformer-based approaches have shown…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Julian Tanke , Takashi Shibuya , Kengo Uchida , Koichi Saito , Yuki Mitsufuji

Multimodal Large Language Models (MLLMs) have achieved great success in Speech-to-Text Translation (S2TT) tasks. However, current research is constrained by two key challenges: language coverage and efficiency. Most of the popular S2TT…

Computation and Language · Computer Science 2026-04-14 Yexing Du , Kaiyuan Liu , Youcheng Pan , Bo Yang , Keqi Deng , Xie Chen , Yang Xiang , Ming Liu , Bing Qin , YaoWei Wang

Dance and music are closely related forms of expression, with mutual retrieval between dance videos and music being a fundamental task in various fields like education, art, and sports. However, existing methods often suffer from unnatural…

Sound · Computer Science 2023-10-17 Kaixing Yang , Xukun Zhou , Xulong Tang , Ran Diao , Hongyan Liu , Jun He , Zhaoxin Fan

The ability of generative large language models (LLMs) to perform in-context learning has given rise to a large body of research into how best to prompt models for various natural language processing tasks. In this paper, we focus on…

Computation and Language · Computer Science 2024-08-02 Armel Zebaze , Benoît Sagot , Rachel Bawden

Automatic choreography generation is a challenging task because it often requires an understanding of two abstract concepts - music and dance - which are realized in the two different modalities, namely audio and video, respectively. In…

Multimedia · Computer Science 2018-11-05 Juheon Lee , Seohyun Kim , Kyogu Lee

The task of music-driven dance generation involves creating coherent dance movements that correspond to the given music. While existing methods can produce physically plausible dances, they often struggle to generalize to out-of-set data.…

Sound · Computer Science 2024-11-12 Bo Han , Teng Zhang , Zeyu Ling , Yi Ren , Xiang Yin , Feilin Han

Sign language translation (SLT) addresses the problem of translating information from a sign language in video to a spoken language in text. Existing studies, while showing progress, are often limited to narrow domains and/or few sign…

Computation and Language · Computer Science 2024-07-17 Biao Zhang , Garrett Tanzer , Orhan Firat

The primary objective of simultaneous machine translation (SiMT) is to minimize latency while preserving the quality of the final translation. Drawing inspiration from CPU branch prediction techniques, we propose incorporating branch…

Computation and Language · Computer Science 2023-12-25 Aoxiong Yin , Tianyun Zhong , Haoyuan Li , Siliang Tang , Zhou Zhao

We propose a novel task for generating 3D dance movements that simultaneously incorporate both text and music modalities. Unlike existing works that generate dance movements using a single modality such as music, our goal is to produce…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Kehong Gong , Dongze Lian , Heng Chang , Chuan Guo , Zihang Jiang , Xinxin Zuo , Michael Bi Mi , Xinchao Wang

Multimodal Large Language Models (MLLMs) have achieved notable success in enhancing translation performance by integrating multimodal information. However, existing research primarily focuses on image-guided methods, whose applicability is…

Computation and Language · Computer Science 2026-03-04 Yexing Du , Youcheng Pan , Zekun Wang , Zheng Chu , Yichong Huang , Kaiyuan Liu , Bo Yang , Yang Xiang , Ming Liu , Bing Qin

Our paper aims to generate diverse and realistic animal motion sequences from textual descriptions, without a large-scale animal text-motion dataset. While the task of text-driven human motion synthesis is already extensively studied and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Zhangsihao Yang , Mingyuan Zhou , Mengyi Shan , Bingbing Wen , Ziwei Xuan , Mitch Hill , Junjie Bai , Guo-Jun Qi , Yalin Wang

Diffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited…

Sound · Computer Science 2023-08-04 Ke Chen , Yusong Wu , Haohe Liu , Marianna Nezhurina , Taylor Berg-Kirkpatrick , Shlomo Dubnov

Full integration of robots into real-life applications necessitates their ability to interpret and execute natural language directives from untrained users. Given the inherent variability in human language, equivalent directives may be…

Robotics · Computer Science 2025-04-08 Eran Beeri Bamani , Eden Nissinman , Rotem Atari , Nevo Heimann Saadon , Avishai Sintov

Driving 3D characters to dance following a piece of music is highly challenging due to the spatial constraints applied to poses by choreography norms. In addition, the generated dance sequence also needs to maintain temporal coherency with…

Sound · Computer Science 2022-03-28 Li Siyao , Weijiang Yu , Tianpei Gu , Chunze Lin , Quan Wang , Chen Qian , Chen Change Loy , Ziwei Liu