English
Related papers

Related papers: Multimodal Banking Dataset: Understanding Client N…

200 papers

Multimodal sentiment analysis aims to recognize people's attitudes from multiple communication channels such as verbal content (i.e., text), voice, and facial expressions. It has become a vibrant and important research topic in natural…

Machine Learning · Computer Science 2022-02-23 Xingbo Wang , Jianben He , Zhihua Jin , Muqiao Yang , Yong Wang , Huamin Qu

Large language models (LLMs) have rapidly evolved from general-purpose systems to multimodal models capable of processing text, images, and audio. As both general-purpose LLMs (GLLMs) and multimodal LLMs (MLLMs) gain widespread adoption,…

Software Engineering · Computer Science 2026-04-08 Yujian Liu , Xiao Yu , Jacky Keung , Xing Hu , Xin Xia , Xiaoxue Ma

Multimodal machine learning is a vibrant multi-disciplinary research field that aims to design computer agents with intelligent capabilities such as understanding, reasoning, and learning through integrating multiple communicative…

Machine Learning · Computer Science 2023-02-21 Paul Pu Liang , Amir Zadeh , Louis-Philippe Morency

Multi-modal time series analysis has recently emerged as a prominent research area in data mining, driven by the increasing availability of diverse data modalities, such as text, images, and structured tabular data from real-world sources.…

Autism spectrum disorder (ASD) is characterized by significant challenges in social interaction and comprehending communication signals. Recently, therapeutic interventions for ASD have increasingly utilized Deep learning powered-computer…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Pavan Uttej Ravva , Behdokht Kiafar , Pinar Kullu , Jicheng Li , Anjana Bhat , Roghayeh Leila Barmaki

This study presents a new high-fidelity multi-modal dataset containing 16000+ geometric variants of automotive hoods useful for machine learning (ML) applications such as engineering component design and process optimization, and…

Machine Learning · Computer Science 2025-11-11 Vansh Sharma , Harish Jai Ganesh , Maryam Akram , Wanjiao Liu , Venkat Raman

Audio-Driven Face Animation is an eagerly anticipated technique for applications such as VR/AR, games, and movie making. With the rapid development of 3D engines, there is an increasing demand for driving 3D faces with audio. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Haozhe Wu , Jia Jia , Junliang Xing , Hongwei Xu , Xiangyuan Wang , Jelo Wang

In order to build self-consistent personalized dialogue agents, previous research has mostly focused on textual persona that delivers personal facts or personalities. However, to fully describe the multi-faceted nature of persona, image…

Computation and Language · Computer Science 2023-05-30 Jaewoo Ahn , Yeda Song , Sangdoo Yun , Gunhee Kim

While many models are purposed for detecting the occurrence of significant events in financial systems, the task of providing qualitative detail on the developments is not usually as well automated. We present a deep learning approach for…

Computation and Language · Computer Science 2018-02-01 Samuel Rönnqvist , Peter Sarlin

In this article, we present a Web-based System called M2LADS, which supports the integration and visualization of multimodal data recorded in learning sessions in a MOOC in the form of Web-based Dashboards. Based on the edBB platform, the…

Human-Computer Interaction · Computer Science 2023-05-23 Álvaro Becerra , Roberto Daza , Ruth Cobos , Aythami Morales , Mutlu Cukurova , Julian Fierrez

Understanding charts requires models to jointly reason over geometric visual patterns, structured numerical data, and natural language -- a capability where current vision-language models (VLMs) remain limited. We introduce ChartNet, a…

In this paper, we present a novel dataset named MVB (Multi View Baggage) for baggage ReID task which has some essential differences from person ReID. The features of MVB are three-fold. First, MVB is the first publicly released large-scale…

Computer Vision and Pattern Recognition · Computer Science 2019-07-29 Zhulin Zhang , Dong Li , Jinhua Wu , Yunda Sun , Li Zhang

Micro-expressions (MEs) are subtle, fleeting nonverbal cues that reveal an individual's genuine emotional state. Their analysis has attracted considerable interest due to its promising applications in fields such as healthcare, criminal…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Chuang Ma , Yu Pei , Jianhang Zhang , Shaokai Zhao , Bowen Ji , Liang Xie , Ye Yan , Erwei Yin

With the rapid advancement of e-commerce, exploring general representations rather than task-specific ones has attracted increasing research attention. For product understanding, although existing discriminative dual-flow architectures…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Daoze Zhang , Chenghan Fu , Zhanheng Nie , Jianyu Liu , Wanxian Guan , Yuan Gao , Jun Song , Pengjie Wang , Jian Xu , Bo Zheng

Today, the acquisition of various behavioral log data has enabled deeper understanding of customer preferences and future behaviors in the marketing field. In particular, multimodal deep learning has achieved highly accurate predictions by…

Computational Engineering, Finance, and Science · Computer Science 2024-05-14 Junichiro Niimi

The advancement of artificial intelligence in oral healthcare relies on the availability of large-scale multimodal datasets that capture the complexity of clinical practice. In this paper, we present a comprehensive multimodal dataset,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Haoxin Lv , Ijazul Haq , Jin Du , Jiaxin Ma , Binnian Zhu , Xiaobing Dang , Chaoan Liang , Ruxu Du , Yingjie Zhang , Muhammad Saqib

Multimodal deep learning has been used to predict clinical endpoints and diagnoses from clinical routine data. However, these models suffer from scaling issues: they have to learn pairwise interactions between each piece of information in…

Multi-modality images have been widely used and provide comprehensive information for medical image analysis. However, acquiring all modalities among all institutes is costly and often impossible in clinical settings. To leverage more…

Image and Video Processing · Electrical Eng. & Systems 2022-09-13 Qi Chang , Hui Qu , Zhennan Yan , Yunhe Gao , Lohendran Baskaran , Dimitris Metaxas

The progress we are currently witnessing in many computer vision applications, including automatic face analysis, would not be made possible without tremendous efforts in collecting and annotating large scale visual databases. To this end,…

Computer Vision and Pattern Recognition · Computer Science 2018-06-15 Shiyang Cheng , Irene Kotsia , Maja Pantic , Stefanos Zafeiriou

In today's world, emotional support is increasingly essential, yet it remains challenging for both those seeking help and those offering it. Multimodal approaches to emotional support show great promise by integrating diverse data sources…

‹ Prev 1 4 5 6 7 8 10 Next ›