中文
相关论文

相关论文: HyperFusion: Hierarchical Multimodal Ensemble Lear…

200 篇论文

Predicting user influence in social networks is a critical problem, and hypergraphs, as a prevalent higher-order modeling approach, provide new perspectives for this task. However, the absence of explicit cascade or infection probability…

社会与信息网络 · 计算机科学 2025-08-22 Su-Su Zhang , JinFeng Xie , Yang Chen , Min Gao , Cong Li , Chuang Liu , Xiu-Xiu Zhan

Social media users articulate their opinions on a broad spectrum of subjects and share their experiences through posts comprising multiple modes of expression, leading to a notable surge in such multimodal content on social media platforms.…

信息检索 · 计算机科学 2024-12-17 Shubhi Bansal , Mohit Kumar , Chandravardhan Singh Raghaw , Nagendra Kumar

In this work, we present a multi-modal model for commercial product classification, that combines features extracted by multiple neural network models from textual (CamemBERT and FlauBERT) and visual data (SE-ResNeXt-50), using simple…

人工智能 · 计算机科学 2022-07-12 Tsegaye Misikir Tashu , Sara Fattouh , Peter Kiss , Tomas Horvath

Accurate estimation of user location is important for many online services. Previous neural network based methods largely ignore the hierarchical structure among locations. In this paper, we propose a hierarchical location prediction neural…

社会与信息网络 · 计算机科学 2019-10-30 Binxuan Huang , Kathleen M. Carley

Multi-agent trajectory prediction in autonomous driving requires a comprehensive understanding of complex social dynamics. Existing methods, however, often struggle to capture the full richness of these dynamics, particularly the…

人工智能 · 计算机科学 2025-12-11 Bingqing Wei , Lianmin Chen , Zhongyu Xia , Yongtao Wang

Predicting popularity, or the total volume of information outbreaks, is an important subproblem for understanding collective behavior in networks. Each of the two main types of recent approaches to the problem, feature-driven and generative…

社会与信息网络 · 计算机科学 2016-08-31 Swapnil Mishra , Marian-Andrei Rizoiu , Lexing Xie

Self-Supervised learning from multimodal image and text data allows deep neural networks to learn powerful features with no need of human annotated data. Web and Social Media platforms provide a virtually unlimited amount of this multimodal…

计算机视觉与模式识别 · 计算机科学 2019-01-09 Raul Gomez , Lluis Gomez , Jaume Gibert , Dimosthenis Karatzas

Multimodal emotion recognition from physiological signals is receiving an increasing amount of attention due to the impossibility to control them at will unlike behavioral reactions, thus providing more reliable information. Existing deep…

人机交互 · 计算机科学 2023-10-12 Eleonora Lopez , Eleonora Chiarantano , Eleonora Grassucci , Danilo Comminiello

Recommender systems are designed to predict user preferences over collections of items. These systems process users' previous interactions to decide which items should be ranked higher to satisfy their desires. An ensemble recommender…

信息检索 · 计算机科学 2023-06-23 Alireza Gharahighehi , Celine Vens , Konstantinos Pliakos

Political activity on social media presents a data-rich window into political behavior, but the vast amount of data means that almost all content analyses of social media require a data labeling step. However, most automated machine…

计算机视觉与模式识别 · 计算机科学 2021-09-24 Patrick Y. Wu , Walter R. Mebane

Multimodal information (e.g., visual, acoustic, and textual) has been widely used to enhance representation learning for micro-video recommendation. For integrating multimodal information into a joint representation of micro-video,…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Han Liu , Yinwei Wei , Fan Liu , Wenjie Wang , Liqiang Nie , Tat-Seng Chua

Many vision-related tasks benefit from reasoning over multiple modalities to leverage complementary views of data in an attempt to learn robust embedding spaces. Most deep learning-based methods rely on a late fusion technique whereby…

计算机视觉与模式识别 · 计算机科学 2020-03-04 Austin Reiter , Menglin Jia , Pu Yang , Ser-Nam Lim

Effective mining of social media, which consists of a large number of users is a challenging task. Traditional approaches rely on the analysis of text data related to users to accomplish this task. However, text data lacks significant…

社会与信息网络 · 计算机科学 2020-07-30 Syed Afaq Ali Shah , Weifeng Deng , Jianxin Li , Muhammad Aamir Cheema , Abdul Bais

Group cohesiveness is a compelling and often studied composition in group dynamics and group performance. The enormous number of web images of groups of people can be used to develop an effective method to detect group cohesiveness. This…

计算机视觉与模式识别 · 计算机科学 2019-10-04 Bin Zhu , Xin Guo , Kenneth Barner , Charles Boncelet

With the rapid information explosion on online social network sites (SNSs), it becomes difficult for users to seek new friends or broaden their social networks in an efficient way. Link prediction, which can effectively conquer this…

社会与信息网络 · 计算机科学 2022-01-26 Huizi Wu , Shiyi Wang , Hui Fang

This paper develops a spatiotemporal model for the visualization of dynamic topologies of hybrid spaces. The visualization of spatiotemporal data is a well-known problem, for example in digital twins in urban planning. There is also a lack…

计算机与社会 · 计算机科学 2024-03-11 Wolfgang Höhl

Multimodal sentiment analysis has emerged as a critical tool for understanding human emotions across diverse communication channels. While existing methods have made significant strides, they often struggle to effectively differentiate and…

机器学习 · 计算机科学 2025-04-01 Jiahao Qin , Feng Liu , Lu Zong

Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data. Current methods primarily utilize task-specific models, while recent foundation models for…

机器学习 · 计算机科学 2025-10-16 Dominik J. Mühlematter , Lin Che , Ye Hong , Martin Raubal , Nina Wiedemann

Speech emotion recognition (SER) remains a challenging yet crucial task due to the inherent complexity and diversity of human emotions. To address this problem, researchers attempt to fuse information from other modalities via multimodal…

声音 · 计算机科学 2024-12-10 Feng Li , Jiusong Luo , Wanjun Xia

Cross view feature fusion is the key to address the occlusion problem in human pose estimation. The current fusion methods need to train a separate model for every pair of cameras making them difficult to scale. In this work, we introduce…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Rongchang Xie , Chunyu Wang , Yizhou Wang