中文
相关论文

相关论文: OFASys: A Multi-Modal Multi-Task Learning System f…

200 篇论文

The success of vision-language models is primarily attributed to effective alignment across modalities such as vision and language. However, modality gaps persist in existing alignment algorithms and appear necessary for human perception as…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Hanqi Yan , Xiangxiang Cui , Lu Yin , Jindong Gu , Paul Pu Liang , Yulan He , Yifei Wang

One of the grand enduring goals of AI is to create generalist agents that can learn multiple different tasks from diverse data via multitask learning (MTL). However, in practice, applying gradient descent (GD) on the average loss across all…

机器学习 · 计算机科学 2023-10-31 Bo Liu , Yihao Feng , Peter Stone , Qiang Liu

Recently, we have witnessed the great success of the generalist model in natural language processing. The generalist model is a general framework trained with massive data and is able to process various downstream tasks simultaneously.…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Ziyi Wang , Yongming Rao , Shuofeng Sun , Xinrun Liu , Yi Wei , Xumin Yu , Zuyan Liu , Yanbo Wang , Hongmin Liu , Jie Zhou , Jiwen Lu

Generating natural language requires conveying content in an appropriate style. We explore two related tasks on generating text of varying formality: monolingual formality transfer and formality-sensitive machine translation. We propose to…

计算与语言 · 计算机科学 2018-06-13 Xing Niu , Sudha Rao , Marine Carpuat

Foundational models (FMs), pretrained on extensive datasets using self-supervised techniques, are capable of learning generalized patterns from large amounts of data. This reduces the need for extensive labeled datasets for each new task,…

机器学习 · 计算机科学 2024-06-19 Quan M. Tran , Suong N. Hoang , Lam M. Nguyen , Dzung Phan , Hoang Thanh Lam

Tracking and segmentation play essential roles in video understanding, providing basic positional information and temporal association of objects within video sequences. Despite their shared objective, existing approaches often tackle these…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Tianlu Zhang , Qiang Zhang , Guiguang Ding , Jungong Han

Recent studies leverage large language models with multi-tasking capabilities, using natural language prompts to guide the model's behavior and surpassing performance of task-specific models. Motivated by this, we ask: can we build a single…

How can we build generalist robot systems? Scale may not be enough due to the significant multimodality of robotics tasks, lack of easily accessible data and the challenges of deploying on physical hardware. Meanwhile, most deployed robotic…

机器人学 · 计算机科学 2025-03-11 Murtaza Dalal

Learning multi-modal representations is an essential step towards real-world robotic applications, and various multi-modal fusion models have been developed for this purpose. However, we observe that existing models, whose objectives are…

机器学习 · 计算机科学 2021-06-22 Chenzhuang Du , Tingle Li , Yichen Liu , Zixin Wen , Tianyu Hua , Yue Wang , Hang Zhao

Missing modality issues are common in real-world applications, arising from factors such as equipment failures and privacy concerns. When fine-tuning pre-trained models on downstream datasets with missing modalities, performance can degrade…

机器学习 · 计算机科学 2025-03-04 Zirun Guo , Shulei Wang , Wang Lin , Weicai Yan , Yangyang Wu , Tao Jin

This paper presents MOCAS, a multimodal dataset dedicated for human cognitive workload (CWL) assessment. In contrast to existing datasets based on virtual game stimuli, the data in MOCAS was collected from realistic closed-circuit…

数据库 · 计算机科学 2024-06-11 Wonse Jo , Ruiqi Wang , Su Sun , Revanth Krishna Senthilkumaran , Daniel Foti , Byung-Cheol Min

We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curriculum-inspired progressive training strategy that transitions…

多媒体 · 计算机科学 2025-12-01 Meituan LongCat Team , Bairui Wang , Bayan , Bin Xiao , Bo Zhang , Bolin Rong , Borun Chen , Chang Wan , Chao Zhang , Chen Huang , Chen Chen , Chen Chen , Chengxu Yang , Chengzuo Yang , Cong Han , Dandan Peng , Delian Ruan , Detai Xin , Disong Wang , Dongchao Yang , Fanfan Liu , Fengjiao Chen , Fengyu Yang , Gan Dong , Gang Huang , Gang Xu , Guanglu Wan , Guoqiang Tan , Guoqiao Yu , Haibo Qiu , Hao Lu , Hongbo Liu , Hongyu Xiang , Jiaheng Wu , Jian Yang , Jiaxing Liu , Jing Huang , Jingang Wang , Jinrui Ding , Juchao Jiang , Jun Kuang , Jun Wang , Junhui Mei , Ke Ding , Kefeng Zhang , Lei Chen , Liang Shi , Limeng Qiao , Liming Zheng , Lin Ma , Liuyang Guo , Liya Ma , Luying Sun , Man Gao , Mengshen Zhu , Miao Cao , Minliang Lin , Nuo Xu , Peng Shi , Qi Zhang , Qian Fang , Qian Wang , Qian Yang , Quanxiu Wang , Rongxiang Weng , Rongxin Guo , Ruoxuan Liang , Senbin Yang , Shanbo Xu , Shanglin Lei , Shengze Ye , Shimin Chen , Shuaiqi Chen , Shujie Hu , Shuo Li , Siqi Yang , Siyu Xu , Siyu Ren , Song Li , Songxiang Liu , Tianhao Bai , Tianye Dai , Wei Hong , Wei Wang , Weixiao Zhao , Wengang Cao , Wenlong Zhu , Wenlong He , Xi Su , Xi Nan , Xiaohan Zhao , Xiaohao Wang , Xiaoyu Zhao , Xiaoyu Wang , Xiaoyu Li , Xin Pan , Xin Chen , Xiusong Sun , Xu Xiang , Xudong Xing , Xuezhi Cao , Xunliang Cai , Yang Yang , Yanli Tan , Yao Yao , Yerui Sun , Yi Chen , Yifan Lu , Yin Gong , Yining Zhang , Yitian Chen , Yiyang Gan , Yuchen Tang , Yuchen Xie , Yueqian Wang , Yuewen Zheng , Yufei Zhang , Yufeng Zhong , Yulei Qian , Yuqi Peng , Yuqian Li , Yuwei Jiang , Zeyang Hu , Zheng Zhang , Zhengkun Tian , Zhiqing Hong , Zhixiong Zeng , Zhuqi Mi , Ziran Li , Ziwen Wang , Ziyi Zhao , Ziyuan Zhuang , Zizhe Zhao

The development of artificial intelligence systems is transitioning from creating static, task-specific models to dynamic, agent-based systems capable of performing well in a wide range of applications. We propose an Interactive Agent…

Multimodal learning has become a prominent research area, with the potential of substantial performance gains by combining information across modalities. At the same time, model development has trended toward increasingly complex deep…

机器学习 · 计算机科学 2026-05-08 Tillmann Rheude , Roland Eils , Benjamin Wild

In this paper, we propose SimMLM, a simple yet powerful framework for multimodal learning with missing modalities. Unlike existing approaches that rely on sophisticated network architectures or complex data imputation techniques, SimMLM…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Sijie Li , Chen Chen , Jungong Han

Continuously learning to solve unseen tasks with limited experience has been extensively pursued in meta-learning and continual learning, but with restricted assumptions such as accessible task distributions, independently and identically…

机器学习 · 计算机科学 2020-12-01 Mengdi Xu , Wenhao Ding , Jiacheng Zhu , Zuxin Liu , Baiming Chen , Ding Zhao

Recent approaches to multi-task learning (MTL) have focused on modelling connections between tasks at the decoder level. This leads to a tight coupling between tasks, which need retraining if a new task is inserted or removed. We argue that…

机器学习 · 计算机科学 2022-04-13 Jaime Spencer , Richard Bowden , Simon Hadfield

People perceive the world with multiple senses (e.g., through hearing sounds, reading words and seeing objects). However, most existing AI systems only process an individual modality. This paper presents an approach that excels at handling…

计算与语言 · 计算机科学 2022-05-13 Yong Dai , Duyu Tang , Liangxin Liu , Minghuan Tan , Cong Zhou , Jingquan Wang , Zhangyin Feng , Fan Zhang , Xueyu Hu , Shuming Shi

Training large language models (LLMs) is constrained by memory requirements, with activations accounting for a substantial fraction of the total footprint. Existing approaches reduce memory using low-rank weight parameterizations or…

机器学习 · 计算机科学 2026-04-13 Sakshi Choudhary , Utkarsh Saxena , Kaushik Roy

We address a challenging lifelong few-shot image generation task for the first time. In this situation, a generative model learns a sequence of tasks using only a few samples per task. Consequently, the learned model encounters both…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Juwon Seo , Ji-Su Kang , Gyeong-Moon Park