English
Related papers

Related papers: Multimodal Representation-disentangled Information…

200 papers

The interactive recommender systems involve users in the recommendation procedure by receiving timely user feedback to update the recommendation policy. Therefore, they are widely used in real application scenarios. Previous interactive…

Information Retrieval · Computer Science 2020-12-25 Qinxu Ding , Yong Liu , Chunyan Miao , Fei Cheng , Haihong Tang

Current multimodal recommendation models have extensively explored the effective utilization of multimodal information; however, their reliance on ID embeddings remains a performance bottleneck. Even with the assistance of multimodal…

Information Retrieval · Computer Science 2024-10-28 Kangning Zhang , Jiarui Jin , Yingjie Qin , Ruilong Su , Jianghao Lin , Yong Yu , Weinan Zhang

Multi-modal entity alignment (MMEA) aims to identify equivalent entities between multi-modal knowledge graphs (MMKGs), where the entities can be associated with related images. Most existing studies integrate multi-modal information heavily…

Computation and Language · Computer Science 2024-07-30 Taoyu Su , Jiawei Sheng , Shicheng Wang , Xinghua Zhang , Hongbo Xu , Tingwen Liu

Sequential recommendation predicts user preferences over time and has achieved remarkable success. However, the growing length of user interaction sequences and the complex entanglement of evolving user interests and intentions introduce…

Information Retrieval · Computer Science 2025-08-06 Haoran Zhang , Jingtong Liu , Jiangzhou Deng , Junpeng Guo

Deep learning models trained on audio-visual data have been successfully used to achieve state-of-the-art performance for emotion recognition. In particular, models trained with multitask learning have shown additional performance…

Image and Video Processing · Electrical Eng. & Systems 2021-02-15 Raghuveer Peri , Srinivas Parthasarathy , Charles Bradshaw , Shiva Sundaram

Learning from demonstrations (LfD) typically relies on large amounts of action-labeled expert trajectories, which fundamentally constrains the scale of available training data. A promising alternative is to learn directly from unlabeled…

Robotics · Computer Science 2025-08-13 Haoyu Zhang , Long Cheng

Multimodal information (e.g., visual, acoustic, and textual) has been widely used to enhance representation learning for micro-video recommendation. For integrating multimodal information into a joint representation of micro-video,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Han Liu , Yinwei Wei , Fan Liu , Wenjie Wang , Liqiang Nie , Tat-Seng Chua

Deep Neural Nets (DNNs) learn latent representations induced by their downstream task, objective function, and other parameters. The quality of the learned representations impacts the DNN's generalization ability and the coherence of the…

Machine Learning · Computer Science 2024-02-13 Nir Weingarten , Zohar Yakhini , Moshe Butman , Ran Gilad-Bachrach

Multimodal recommendation systems are increasingly popular for their potential to improve performance by integrating diverse data types. However, the actual benefits of this integration remain unclear, raising questions about when and how…

Information Retrieval · Computer Science 2025-08-08 Hongyu Zhou , Yinan Zhang , Aixin Sun , Zhiqi Shen

Food recommendation systems serve as pivotal components in the realm of digital lifestyle services, designed to assist users in discovering recipes and food items that resonate with their unique dietary predilections. Typically, multi-modal…

Information Retrieval · Computer Science 2025-02-28 Yixin Zhang , Xin Zhou , Qianwen Meng , Fanglin Zhu , Yonghui Xu , Zhiqi Shen , Lizhen Cui

Radiologists must utilize multiple modal images for tumor segmentation and diagnosis due to the limitations of medical imaging and the diversity of tumor signals. This leads to the development of multimodal learning in segmentation.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Chuyun Shen , Wenhao Li , Haoqing Chen , Xiaoling Wang , Fengping Zhu , Yuxin Li , Xiangfeng Wang , Bo Jin

Multi-behavior recommendation faces a critical challenge in practice: auxiliary behaviors (e.g., clicks, carts) are often noisy, weakly correlated, or semantically misaligned with the target behavior (e.g., purchase), which leads to biased…

Information Retrieval · Computer Science 2026-01-22 Miaomiao Cai , Zhijie Zhang , Junfeng Fang , Zhiyong Cheng , Xiang Wang , Meng Wang

Task-oriented semantic communication emerges as a crucial paradigm for next-generation wireless networks, aiming to efficiently transmit task-relevant information while reducing interference and redundancy across multiple users. Existing…

Signal Processing · Electrical Eng. & Systems 2026-04-14 Jiaxiang Wang , Zhaohui Yang , Yahao Ding , Ye Hu , Mohammad Shikh-Bahaei

Normalization is fundamental to deep learning, but existing approaches such as BatchNorm, LayerNorm, and RMSNorm are variance-centric by enforcing zero mean and unit variance, stabilizing training without controlling how representations…

Machine Learning · Computer Science 2026-01-30 Xiandong Zou , Jia Li , Xiaotong Yuan , Pan Zhou

We propose a new GAN-based unsupervised model for disentangled representation learning. The new model is discovered in an attempt to utilize the Information Bottleneck (IB) framework to the optimization of GAN, thereby named IB-GAN. The…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Insu Jeon , Wonkwang Lee , Myeongjang Pyeon , Gunhee Kim

Purpose: Different Magnetic resonance imaging (MRI) modalities of the same anatomical structure are required to present different pathological information from the physical level for diagnostic needs. However, it is often difficult to…

Image and Video Processing · Electrical Eng. & Systems 2021-09-15 Yuchen Fei , Bo Zhan , Mei Hong , Xi Wu , Jiliu Zhou , Yan Wang

Self-supervised learning aims to learn representation that can be effectively generalized to downstream tasks. Many self-supervised approaches regard two views of an image as both the input and the self-supervised signals, assuming that…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Liangjian Wen , Xiasi Wang , Jianzhuang Liu , Zenglin Xu

Multimodal models trained on complete modality data often exhibit a substantial decrease in performance when faced with imperfect data containing corruptions or missing modalities. To address this robustness challenge, prior methods have…

Multimedia · Computer Science 2023-10-24 Mengxi Chen , Jiangchao Yao , Linyu Xing , Yu Wang , Ya Zhang , Yanfeng Wang

A real-world application or setting involves interaction between different modalities (e.g., video, speech, text). In order to process the multimodal information automatically and use it for an end application, Multimodal Representation…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Abhinav Joshi , Naman Gupta , Jinang Shah , Binod Bhattarai , Ashutosh Modi , Danail Stoyanov

Multimodal recommendation has emerged as a promising solution to alleviate the cold-start and sparsity problems in collaborative filtering by incorporating rich content information, such as product images and textual descriptions. However,…

Information Retrieval · Computer Science 2025-06-03 Sibei Liu , Yuanzhe Zhang , Xiang Li , Yunbo Liu , Chengwei Feng , Hao Yang