English
Related papers

Related papers: RingMoE: Mixture-of-Modality-Experts Multi-Modal F…

200 papers

The multimodal language models (MLMs) based on generative pre-trained Transformer are considered powerful candidates for unifying various domains and tasks. MLMs developed for remote sensing (RS) have demonstrated outstanding performance in…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Qingyun Li , Yushi Chen , Xinya Shu , Dong Chen , Xin He , Yi Yu , Xue Yang

Self-supervised frameworks for representation learning have recently stirred up interest among the remote sensing community, given their potential to mitigate the high labeling costs associated with curating large satellite image datasets.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hugo Chan-To-Hing , Bharadwaj Veeravalli

Advances in multimodal models have greatly improved how interactions relevant to various tasks are modeled. Today's multimodal models mainly focus on the correspondence between images and text, using this for tasks like image-text matching.…

Computation and Language · Computer Science 2024-09-27 Haofei Yu , Zhengyang Qi , Lawrence Jang , Ruslan Salakhutdinov , Louis-Philippe Morency , Paul Pu Liang

Healthcare relies on multiple types of data, such as medical images, genetic information, and clinical records, to improve diagnosis and treatment. However, missing data is a common challenge due to privacy restrictions, cost, and technical…

Machine Learning · Computer Science 2025-03-13 Nazanin Moradinasab , Saurav Sengupta , Jiebei Liu , Sana Syed , Donald E. Brown

While Large Language Models (LLMs) are commonly fine-tuned to handle domain-specific tasks before being applied to vertical applications, adapting them to complex scenarios with diverse specialized knowledge remains challenging. Meanwhile,…

Machine Learning · Computer Science 2026-05-26 Mengyang Sun , Maochuan Dou , Tao Feng , Dan Zhang , Yihao Wang , Junpeng Liu , Yifan Zhu , Jie Tang

Multimodal fusion of remote sensing images serves as a core technology for overcoming the limitations of single-source data and improving the accuracy of surface information extraction, which exhibits significant application value in fields…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Siyu Zhang , Lianlei Shan , Runhe Qiu

The Mixtures-of-Experts (MoE) model is a widespread distributed and integrated learning method for large language models (LLM), which is favored due to its ability to sparsify and expand models efficiently. However, the performance of MoE…

Machine Learning · Computer Science 2024-05-24 Jing Li , Zhijie Sun , Xuan He , Li Zeng , Yi Lin , Entong Li , Binfan Zheng , Rongqian Zhao , Xin Chen

Confusion and forgetting of object classes have been challenges of prime interest in Few-Shot Object Detection (FSOD). To overcome these pitfalls in metric learning based FSOD techniques, we introduce a novel Submodular Mutual Information…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Anay Majee , Ryan Sharp , Rishabh Iyer

In this study, we propose MoME, a Mixture of Visual Language Medical Experts, for Medical Image Segmentation. MoME adapts the successful Mixture of Experts (MoE) paradigm, widely used in Large Language Models (LLMs), for medical…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Arghavan Rezvani , Xiangyi Yan , Anthony T. Wu , Kun Han , Pooya Khosravi , Xiaohui Xie

Preference-based reinforcement learning offers a scalable alternative to manual reward engineering by learning reward structures from comparative feedback. However, large-scale preference datasets, whether collected from crowdsourced…

Robotics · Computer Science 2026-05-04 Ziqin Yuan , Ruiqi Wang , Dezhong Zhao , Baijian Yang , Byung-Cheol Min

The Mixture of Experts (MoE) architecture enables the scaling of Large Language Models (LLMs) to trillions of parameters by activating a sparse subset of weights for each input, maintaining constant computational cost during inference.…

Machine Learning · Computer Science 2026-01-08 Shihao Ji , Zihui Song

Effectively managing missing modalities is a fundamental challenge in real-world multimodal learning scenarios, where data incompleteness often results from systematic collection errors or sensor failures. Sparse Mixture-of-Experts (SMoE)…

Machine Learning · Computer Science 2026-05-12 Liangwei Nathan Zheng , Wei Emma Zhang , Mingyu Guo , Olaf Maennel , Weitong Chen

Functional connectivity (FC) derived from resting-state fMRI plays a critical role in personalized predictions such as age and cognitive performance. However, applying foundation models(FM) to fMRI data remains challenging due to its high…

Neurons and Cognition · Quantitative Biology 2025-08-26 Yanwen Wang , Xinglin Zhao , Yijin Song , Xiaobo Liu , Yanrong Hao , Rui Cao , Xin Wen

Large-scale foundation models in Earth Observation can learn versatile, label-efficient representations by leveraging massive amounts of unlabeled data. However, existing public datasets are often limited in scale, geographic coverage, or…

Vision foundation models like DINOv2 demonstrate remarkable potential in medical imaging despite their origin in natural image domains. However, their design inherently works best for uni-modal image analysis, limiting their effectiveness…

Image and Video Processing · Electrical Eng. & Systems 2025-09-09 Daniel Scholz , Ayhan Can Erdur , Viktoria Ehm , Anke Meyer-Baese , Jan C. Peeken , Daniel Rueckert , Benedikt Wiestler

Foundation models have advanced machine learning across various modalities, including images. Recently multiple teams trained foundation models specialized for remote sensing applications. This line of research is motivated by the distinct…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Ani Vanyan , Alvard Barseghyan , Hakob Tamazyan , Tigran Galstyan , Vahan Huroyan , Naira Hovakimyan , Hrant Khachatrian

Survival analysis, as a challenging task, requires integrating Whole Slide Images (WSIs) and genomic data for comprehensive decision-making. There are two main challenges in this task: significant heterogeneity and complex inter- and…

Image and Video Processing · Electrical Eng. & Systems 2024-06-17 Conghao Xiong , Hao Chen , Hao Zheng , Dong Wei , Yefeng Zheng , Joseph J. Y. Sung , Irwin King

Mixture of Experts (MoE) is able to scale up vision transformers effectively. However, it requires prohibiting computation resources to train a large MoE transformer. In this paper, we propose Residual Mixture of Experts (RMoE), an…

Computer Vision and Pattern Recognition · Computer Science 2022-10-05 Lemeng Wu , Mengchen Liu , Yinpeng Chen , Dongdong Chen , Xiyang Dai , Lu Yuan

Multi-task learning (MTL) leverages a shared model to accomplish multiple tasks and facilitate knowledge transfer. Recent research on task arithmetic-based MTL demonstrates that merging the parameters of independently fine-tuned models can…

Machine Learning · Computer Science 2024-10-30 Li Shen , Anke Tang , Enneng Yang , Guibing Guo , Yong Luo , Lefei Zhang , Xiaochun Cao , Bo Du , Dacheng Tao

In the field of information extraction (IE), tasks across a wide range of modalities and their combinations have been traditionally studied in isolation, leaving a gap in deeply recognizing and analyzing cross-modal information. To address…

Multimedia · Computer Science 2024-06-12 Meishan Zhang , Hao Fei , Bin Wang , Shengqiong Wu , Yixin Cao , Fei Li , Min Zhang
‹ Prev 1 8 9 10 Next ›