English
Related papers

Related papers: Captions Speak Louder than Images: Generalizing Fo…

200 papers

With tremendous efforts on developing effective e-commerce models, conventional e-commerce models show limited success in generalist e-commerce modeling, and suffer from unsatisfactory performance on new users and new products - a typical…

Computation and Language · Computer Science 2024-08-06 Bo Peng , Xinyi Ling , Ziru Chen , Huan Sun , Xia Ning

LLMs and MLLMs have become indispensable tools across a wide range of applications. E-commerce, however, poses distinctive challenges -- including intricate domain knowledge, long-tail product evidence, heterogeneous visual data, and the…

Databases · Computer Science 2026-05-14 Yong Liu , Ximan Liu , Guoqing Yang , Bing Bai , Xiaoqiang Xu , Zhen Chen , Ke Zhang , Yan Li

E-commerce platforms are rich in multimodal data, featuring a variety of images that depict product details. However, this raises an important question: do these images always enhance product understanding, or can they sometimes introduce…

Computation and Language · Computer Science 2025-11-14 Xinyi Ling , Hanwen Du , Zhihui Zhu , Xia Ning

This paper aims to establish a generic multi-modal foundation model that has the scalable capability to massive downstream applications in E-commerce. Recently, large-scale vision-language pretraining approaches have achieved remarkable…

Computer Vision and Pattern Recognition · Computer Science 2023-04-07 Yang Jin , Yongzhi Li , Zehuan Yuan , Yadong Mu

With the rapid advancement of e-commerce, exploring general representations rather than task-specific ones has attracted increasing research attention. For product understanding, although existing discriminative dual-flow architectures…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Daoze Zhang , Chenghan Fu , Zhanheng Nie , Jianyu Liu , Wanxian Guan , Yuan Gao , Jun Song , Pengjie Wang , Jian Xu , Bo Zheng

Multimodal foundation models that can holistically process text alongside images, video, audio, and other sensory modalities are increasingly used in a variety of real-world applications. However, it is challenging to characterize and study…

Over the past decade, significant advances have been made in the field of image search for e-commerce applications. Traditional image-to-image retrieval models, which focus solely on image details such as texture, tend to overlook useful…

Computer Vision and Pattern Recognition · Computer Science 2024-02-09 Chang Liu , Peng Hou , Anxiang Zeng , Han Yu

Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challenges: (i) the modality imbalance induced by modality mixed training; (ii) underutilization of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Zhanheng Nie , Chenghan Fu , Daoze Zhang , Junxian Wu , Wanxian Guan , Pengjie Wang , Jian Xu , Bo Zheng

In recent years, multimodal benchmarks for general domains have guided the rapid development of multimodal models on general tasks. However, the financial field has its peculiarities. It features unique graphical images (e.g., candlestick…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Ziliang Gan , Yu Lu , Dong Zhang , Haohan Li , Che Liu , Jian Liu , Ji Liu , Haipang Wu , Chaoyou Fu , Zenglin Xu , Rongjunchen Zhang , Yong Dai

Missing-modality information on e-commerce platforms, such as absent product images or textual descriptions, often arises from annotation errors or incomplete metadata, impairing both product presentation and downstream applications such as…

Multimedia · Computer Science 2026-01-29 Junchen Fu , Wenhao Deng , Kaiwen Zheng , Ioannis Arapakis , Yu Ye , Yongxin Ni , Joemon M. Jose , Xuri Ge

The development of Multimodal Large Language Models (MLLMs) has seen significant advancements with increasing demands in various fields (e.g., multimodal agents, embodied intelligence). While model-driven approaches attempt to enhance MLLMs…

The Contrastive Language-Image Pre-training (CLIP) framework has become a widely used approach for multimodal representation learning, particularly in image-text retrieval and clustering. However, its efficacy is constrained by three key…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Tiancheng Gu , Kaicheng Yang , Ziyong Feng , Xingjun Wang , Yanzhao Zhang , Dingkun Long , Yingda Chen , Weidong Cai , Jiankang Deng

Foundation models have emerged as a powerful approach for processing electronic health records (EHRs), offering flexibility to handle diverse medical data modalities. In this study, we present a comprehensive benchmark that evaluates the…

Machine Learning · Computer Science 2025-07-22 Kunyu Yu , Rui Yang , Jingchi Liao , Siqi Li , Huitao Li , Irene Li , Yifan Peng , Rishikesan Kamaleswaran , Nan Liu

Multimodal large language models are playing an increasingly significant role in empowering the financial domain, however, the challenges they face, such as multimodal and high-density information and cross-modal multi-hop reasoning, go…

Expensive multi-objective optimization is a prevalent and crucial concern in many real-world scenarios, where sample-efficiency is vital due to the limited evaluations to recover the true Pareto front for decision making. Existing works…

Machine Learning · Computer Science 2026-02-03 Yiming Yao , Fei Liu , Liang Zhao , Xi Lin , Yilu Liu , Qingfu Zhang

Large Language Models (LLMs) excel on general-purpose NLP benchmarks, yet their capabilities in specialized domains remain underexplored. In e-commerce, existing evaluations-such as EcomInstruct, ChineseEcomQA, eCeLLM, and Shopping…

Artificial Intelligence · Computer Science 2025-10-24 Shuyi Xie , Ziqin Liew , Hailing Zhang , Haibo Zhang , Ling Hu , Zhiqiang Zhou , Shuman Liu , Anxiang Zeng

Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions. The capabilities of MLLMs depend on two crucial factors: the model architecture to facilitate the feature…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Tianyu Yu , Jinyi Hu , Yuan Yao , Haoye Zhang , Yue Zhao , Chongyi Wang , Shan Wang , Yinxv Pan , Jiao Xue , Dahai Li , Zhiyuan Liu , Hai-Tao Zheng , Maosong Sun

In this paper, we address multi-modal pretraining of product data in the field of E-commerce. Current multi-modal pretraining methods proposed for image and text modalities lack robustness in the face of modality-missing and modality-noise,…

Computer Vision and Pattern Recognition · Computer Science 2021-09-03 Yushan Zhu , Huaixiao Tou , Wen Zhang , Ganqiang Ye , Hui Chen , Ningyu Zhang , Huajun Chen

Classifying products into categories precisely and efficiently is a major challenge in modern e-commerce. The high traffic of new products uploaded daily and the dynamic nature of the categories raise the need for machine learning models…

Computer Vision and Pattern Recognition · Computer Science 2016-11-30 Tom Zahavy , Alessandro Magnani , Abhinandan Krishnan , Shie Mannor

With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention. Although recent multimodal large language models (MLLMs) have driven significant progress in product…

Machine Learning · Computer Science 2026-04-03 Junxian Wu , Chenghan Fu , Zhanheng Nie , Daoze Zhang , Bowen Wan , Wanxian Guan , Chuan Yu , Jian Xu , Bo Zheng
‹ Prev 1 2 3 10 Next ›