中文
相关论文

相关论文: Towards Flexible, Scalable, and Adaptive Multi-Mod…

200 篇论文

Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily focus on RGB video generation and lack the ability to support…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Guile Wu , David Huang , Dongfeng Bai , Bingbing Liu

Medical multi-modal pre-training has revealed promise in computer-aided diagnosis by leveraging large-scale unlabeled datasets. However, existing methods based on masked autoencoders mainly rely on data-level reconstruction tasks, but lack…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Yupei Zhang , Li Pan , Qiushi Yang , Tan Li , Zhen Chen

In the diverse field of medical imaging, automatic segmentation has numerous applications and must handle a wide variety of input domains, such as different types of Computed Tomography (CT) scans and Magnetic Resonance (MR) images. This…

图像与视频处理 · 电气工程与系统科学 2024-11-26 Chengyin Li , Hui Zhu , Rafi Ibn Sultan , Hassan Bagher Ebadian , Prashant Khanduri , Chetty Indrin , Kundan Thind , Dongxiao Zhu

In this paper, we study the challenging unconstrained set-based face recognition problem where each subject face is instantiated by a set of media (images and videos) instead of a single image. Naively aggregating information from all the…

计算机视觉与模式识别 · 计算机科学 2019-03-26 Jian Zhao , Jianshu Li , Xiaoguang Tu , Fang Zhao , Yuan Xin , Junliang Xing , Hengzhu Liu , Shuicheng Yan , Jiashi Feng

Multimodal Magnetic Resonance (MR) Imaging plays a crucial role in disease diagnosis due to its ability to provide complementary information by analyzing a relationship between multimodal images on the same subject. Acquiring all MR…

图像与视频处理 · 电气工程与系统科学 2024-02-02 Jihoon Cho , Xiaofeng Liu , Fangxu Xing , Jinsong Ouyang , Georges El Fakhri , Jinah Park , Jonghye Woo

Multi-modal representation learning has become a pivotal area in artificial intelligence, enabling the integration of diverse modalities such as vision, text, and audio to solve complex problems. However, existing approaches predominantly…

机器学习 · 计算机科学 2025-05-01 Sangyeon Cho , Jangyeong Jeon , Mingi Kim , Junyeong Kim

Understanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Dingkang Yang , Mingcheng Li , Linhao Qu , Kun Yang , Peng Zhai , Song Wang , Lihua Zhang

Current video models fail as world model as they lack fine-graiend control. General-purpose household robots require real-time fine motor control to handle delicate tasks and urgent situations. In this work, we introduce fine-grained…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Yichen Li , Antonio Torralba

We present multimodal conditioning modules (MCM) for enabling conditional image synthesis using pretrained diffusion models. Previous multimodal synthesis works rely on training networks from scratch or fine-tuning pretrained networks, both…

计算机视觉与模式识别 · 计算机科学 2023-05-22 Cusuh Ham , James Hays , Jingwan Lu , Krishna Kumar Singh , Zhifei Zhang , Tobias Hinz

In the field of deep learning applied to face recognition, securing large-scale, high-quality datasets is vital for attaining precise and reliable results. However, amassing significant volumes of high-quality real data faces hurdles such…

计算机视觉与模式识别 · 计算机科学 2023-05-18 Omer Granoviter , Alexey Gruzdev , Vladimir Loginov , Max Kogan , Orly Zvitia

Existing facial editing methods have achieved remarkable results, yet they often fall short in supporting multimodal conditional local facial editing. One of the significant evidences is that their output image quality degrades dramatically…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Wanglong Lu , Jikai Wang , Xiaogang Jin , Xianta Jiang , Hanli Zhao

Existing face-swapping methods often deliver competitive results in constrained settings but exhibit substantial quality degradation when handling extreme facial poses. To improve facial pose robustness, explicit geometric features are…

计算机视觉与模式识别 · 计算机科学 2026-01-26 Jongmin Yu , Hyeontaek Oh , Zhongtian Sun , Angelica I Aviles-Rivero , Moongu Jeon , Jinhong Yang

Synthetically generated face images have shown to be indistinguishable from real images by humans and as such can lead to a lack of trust in digital content as they can, for instance, be used to spread misinformation. Therefore, the need to…

计算机视觉与模式识别 · 计算机科学 2024-12-31 M. Ibsen , C. Rathgeb , S. Marcel , C. Busch

Employing deep learning-based approaches for fine-grained facial expression analysis, such as those involving the estimation of Action Unit (AU) intensities, is difficult due to the lack of a large-scale dataset of real faces with…

计算机视觉与模式识别 · 计算机科学 2019-05-31 Zhilei Liu , Guoxian Song , Jianfei Cai , Tat-Jen Cham , Juyong Zhang

Multimodal learning has seen great success mining data features from multiple modalities with remarkable model performance improvement. Meanwhile, federated learning (FL) addresses the data sharing problem, enabling privacy-preserved…

机器学习 · 计算机科学 2023-03-29 Rongyu Zhang , Xiaowei Chi , Guiliang Liu , Wenyi Zhang , Yuan Du , Fangxin Wang

Multimodal variational autoencoders have demonstrated their ability to learn the relationships between different modalities by mapping them into a latent representation. Their design and capacity to perform any-to-any conditional and…

机器学习 · 计算机科学 2025-02-04 Daniel Wesego , Pedram Rooshenas

Magnetic resonance imaging (MRI) is a widely used neuroimaging technique that can provide images of different contrasts (i.e., modalities). Fusing this multi-modal data has proven particularly effective for boosting model performance in…

计算机视觉与模式识别 · 计算机科学 2020-02-13 Tao Zhou , Huazhu Fu , Geng Chen , Jianbing Shen , Ling Shao

In the application of face recognition, eyeglasses could significantly degrade the recognition accuracy. A feasible method is to collect large-scale face images with eyeglasses for training deep learning methods. However, it is difficult to…

计算机视觉与模式识别 · 计算机科学 2021-02-11 Jianzhu Guo , Xiangyu Zhu , Zhen Lei , Stan Z. Li

Research on multi-modal learning dominantly aligns the modalities in a unified space at training, and only a single one is taken for prediction at inference. However, for a real machine, e.g., a robot, sensors could be added or removed at…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Yuanhuiyi Lyu , Xu Zheng , Dahun Kim , Lin Wang

Reconstructing 3D face from a single unconstrained image remains a challenging problem due to diverse conditions in unconstrained environments. Recently, learning-based methods have achieved notable results by effectively capturing complex…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Danling Cao