中文
相关论文

相关论文: Towards Multimodal MIR: Predicting individual diff…

200 篇论文

Basic personality traits are typically assessed through questionnaires. Here we consider phone-based metrics as a way to asses personality traits. We use data from smartphones with custom data-collection software distributed to 730…

社会与信息网络 · 计算机科学 2018-03-06 Bjarke Mønsted , Anders Mollgaard , Joachim Mathiesen

In this study, we aim to determine if generalized sounds and music can share a common emotional space, improving predictions of emotion in terms of arousal and valence. We propose the use of multiple datasets as a multi-domain learning…

声音 · 计算机科学 2024-08-15 Federico Simonetta , Francesca Certo , Stavros Ntalampiras

The rapid development of musical AI technologies has expanded the creative potential of various musical activities, ranging from music style transformation to music generation. However, little research has investigated how musical AIs can…

人机交互 · 计算机科学 2024-04-16 Jingjing Sun , Jingyi Yang , Guyue Zhou , Yucheng Jin , Jiangtao Gong

Integrating physiological signals such as electroencephalogram (EEG), with other data such as interview audio, may offer valuable multimodal insights into psychological states or neurological disorders. Recent advancements with Large…

人机交互 · 计算机科学 2024-08-15 Yongquan Hu , Shuning Zhang , Ting Dang , Hong Jia , Flora D. Salim , Wen Hu , Aaron J. Quigley

Assessment of job performance, personalized health and psychometric measures are domains where data-driven and ubiquitous computing exhibits the potential of a profound impact in the future. Existing techniques use data extracted from…

Automated personality and soft skill assessment from multimodal behavioral data remains challenging due to limited datasets and methods that fail to capture geometric structure inherent in human traits. We introduce RecruitView, a dataset…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Amit Kumar Gupta , Farhan Sheth , Hammad Shaikh , Dheeraj Kumar , Angkul Puniya , Deepak Panwar , Sandeep Chaurasia , Priya Mathur

Multi-modality is an important feature of sensor based activity recognition. In this work, we consider two inherent characteristics of human activities, the spatially-temporally varying salience of features and the relations between…

人机交互 · 计算机科学 2019-05-23 Kaixuan Chen , Lina Yao , Dalin Zhang , Bin Guo , Zhiwen Yu

Quantitative estimation of human joint motion in daily living spaces is essential for early detection and rehabilitation tracking of neuromusculoskeletal disorders (e.g., Parkinson's) and mitigating trip and fall risks for older adults.…

人机交互 · 计算机科学 2025-03-24 Yiwen Dong , Jessica Rose , Hae Young Noh

Research on group activity recognition mostly leans on the standard two-stream approach (RGB and Optical Flow) as their input features. Few have explored explicit pose information, with none using it directly to reason about the persons…

计算机视觉与模式识别 · 计算机科学 2021-10-11 Mauricio Perez , Jun Liu , Alex C. Kot

Musical performance requires prediction to operate instruments, to perform in groups and to improvise. In this paper, we investigate how a number of digital musical instruments (DMIs), including two of our own, have applied predictive…

声音 · 计算机科学 2018-12-21 Charles P. Martin , Kai Olav Ellefsen , Jim Torresen

This study analyzes the possible relationship between personality traits, in terms of Big Five (extraversion, agreeableness, responsibility, emotional stability and openness to experience), and social interactions mediated by digital…

计算机与社会 · 计算机科学 2023-06-12 Andrea Mercado , Alethia Hume , Ivano Bison , Fausto Giunchiglia , Amarsanaa Ganbold , Luca Cernuzzi

Multimodal emotion recognition (MMER) systems typically outperform unimodal systems by leveraging the inter- and intra-modal relationships between, e.g., visual, textual, physiological, and auditory modalities. This paper proposes an MMER…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Paul Waligora , Haseeb Aslam , Osama Zeeshan , Soufiane Belharbi , Alessandro Lameiras Koerich , Marco Pedersoli , Simon Bacon , Eric Granger

Estimating the fundamental frequency, or melody, is a core task in Music Information Retrieval (MIR). Various studies have explored signal processing, machine learning, and deep-learning-based approaches, with a very recent focus on…

音频与语音处理 · 电气工程与系统科学 2025-09-23 Aayush Jaiswal , Parampreet Singh , Vipul Arora

Mean Opinion Score (MOS) prediction for text to music systems requires evaluating both overall musical quality and text prompt alignment. This paper introduces WhisQ, a multimodal architecture that addresses this dual-assessment challenge…

声音 · 计算机科学 2025-06-09 Jakaria Islam Emon , Kazi Tamanna Alam , Md. Abu Salek

While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music necessitate tailored approaches. Music Audio-Visual Question Answering (Music AVQA) particularly…

A music piece is both comprehended hierarchically, from sonic events to melodies, and sequentially, in the form of repetition and variation. Music from different cultures establish different aesthetics by having different style conventions…

声音 · 计算机科学 2021-11-25 Shlomo Dubnov , Kevin Huang , Cheng-i Wang

We present a technique to search for the presence of crucial events in music, based on the analysis of the music volume. Earlier work on this issue was based on the assumption that crucial events correspond to the change of music notes,…

物理与社会 · 物理学 2018-03-14 April Pease , Korosh Mahmoodi , Bruce J. West

This research aims to quantify human walking patterns through depth cameras to (1) detect walking pattern changes of a person with and without a motion-restricting device or a walking aid, and to (2) identify distinct walking patterns from…

人机交互 · 计算机科学 2019-03-22 Behnam Malmir

Depression is a widespread mental health issue affecting diverse age groups, with notable prevalence among college students and the elderly. However, existing datasets and detection methods primarily focus on young adults, neglecting the…

We focus on multi-modal fusion for egocentric action recognition, and propose a novel architecture for multi-modal temporal-binding, i.e. the combination of modalities within a range of temporal offsets. We train the architecture with three…

计算机视觉与模式识别 · 计算机科学 2019-08-23 Evangelos Kazakos , Arsha Nagrani , Andrew Zisserman , Dima Damen