中文
相关论文

相关论文: OtterHD: A High-Resolution Multi-modality Model

200 篇论文

The advent of deep learning has led to significant progress in monocular human reconstruction. However, existing representations, such as parametric models, voxel grids, meshes and implicit neural representations, have difficulties…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Qiao Feng , Yebin Liu , Yu-Kun Lai , Jingyu Yang , Kun Li

Early and accurate detection of the bone fracture is paramount to initiating treatment as early as possible and avoiding any delay in patient treatment and outcomes. Interpretation of X-ray image is a time consuming and error prone task,…

图像与视频处理 · 电气工程与系统科学 2025-08-07 Md. Ehsanul Haque , Abrar Fahim , Shamik Dey , Syoda Anamika Jahan , S. M. Jahidul Islam , Sakib Rokoni , Md Sakib Morshed

This research introduces a transformative framework for integrating Vision-Enhanced Large Language Models (LLMs) with advanced transformer-based architectures to tackle challenges in high-resolution image synthesis and multimodal data…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Karthikeya KV

Efficient and accurate extraction of key information from 2D engineering drawings is essential for advancing digital manufacturing workflows. Such information includes geometric dimensioning and tolerancing (GD&T), measures, material…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Muhammad Tayyab Khan , Lequn Chen , Zane Yong , Jun Ming Tan , Wenhe Feng , Seung Ki Moon

Emerging holographic display technology offers unique capabilities for next-generation virtual reality systems. Current holographic near-eye displays, however, only support a small \'etendue, which results in a direct tradeoff between…

图形学 · 计算机科学 2024-11-26 Brian Chao , Manu Gopakumar , Suyeon Choi , Jonghyun Kim , Liang Shi , Gordon Wetzstein

Universal multimodal embedding models have achieved great success in capturing semantic relevance between queries and candidates. However, current methods either condense queries and candidates into a single vector, potentially limiting the…

信息检索 · 计算机科学 2026-04-08 Zilin Xiao , Qi Ma , Mengting Gu , Chun-cheng Jason Chen , Xintao Chen , Vicente Ordonez , Vijai Mohan

Recent evidence suggests that modeling higher-order interactions (HOIs) in functional magnetic resonance imaging (fMRI) data can enhance the diagnostic accuracy of machine learning systems. However, effectively extracting and utilizing HOIs…

机器学习 · 计算机科学 2025-11-04 Kunyu Zhang , Qiang Li , Shujian Yu

Recent advancements in large-scale models have showcased remarkable generalization capabilities in various tasks. However, integrating multimodal processing into these models presents a significant challenge, as it often comes with a high…

多媒体 · 计算机科学 2024-07-17 Hao Sun , Yu Song , Xinyao Yu , Jiaqing Liu , Yen-Wei Chen , Lanfen Lin

Understanding human-to-human interactions, especially in contexts like public security surveillance, is critical for monitoring and maintaining safety. Traditional activity recognition systems are limited by fixed vocabularies, predefined…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Lala Shakti Swarup Ray , Bo Zhou , Sungho Suh , Paul Lukowicz

Multimodal Large Language Models (MM-LLMs) have seen significant advancements in the last year, demonstrating impressive performance across tasks. However, to truly democratize AI, models must exhibit strong capabilities and be able to run…

机器学习 · 计算机科学 2024-09-04 Jainaveen Sundaram , Ravi Iyer

The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpart. In this paper, we introduce Baichuan-omni, the first…

This article introduces a benchmark designed to evaluate the capabilities of multimodal models in analyzing and interpreting images. The benchmark focuses on seven key visual aspects: main object, additional objects, background, detail,…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Evgenii Evstafev

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multimodal models fail to provide satisfactory results in describing occluded objects for…

计算机视觉与模式识别 · 计算机科学 2024-10-03 Wenmo Qiu , Xinhan Di

Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Run Ling , Ke Cao , Jian Lu , Ao Ma , Haowei Liu , Runze He , Changwei Wang , Rongtao Xu , Yihua Shao , Zhanjie Zhang , Peng Wu , Guibing Guo , Wei Feng , Zheng Zhang , Jingjing Lv , Junjie Shen , Ching Law , Xingwei Wang

We introduce Motif-2-12.7B, a new open-weight foundation model that pushes the efficiency frontier of large language models by combining architectural innovation with system-level optimization. Designed for scalable language understanding…

High-dimensional tensor models are notoriously computationally expensive to train. We present a meta-learning algorithm, MMT, that can significantly speed up the process for spatial tensor models. MMT leverages the property that spatial…

机器学习 · 计算机科学 2018-03-01 Stephan Zheng , Rose Yu , Yisong Yue

The advent of foundation models has heralded a new era in medical artificial intelligence (AI), enabling the extraction of generalizable representations from large-scale unlabeled datasets. However, current ophthalmic AI paradigms are…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Tienyu Chang , Zhen Chen , Renjie Liang , Jinyu Ding , Jie Xu , Sunu Mathew , Amir Reza Hajrasouliha , Andrew J. Saykin , Ruogu Fang , Yu Huang , Jiang Bian , Qingyu Chen

We present OCTCube-M, a 3D OCT-based multi-modal foundation model for jointly analyzing OCT and en face images. OCTCube-M first developed OCTCube, a 3D foundation model pre-trained on 26,685 3D OCT volumes encompassing 1.62 million 2D OCT…

图像与视频处理 · 电气工程与系统科学 2024-12-20 Zixuan Liu , Hanwen Xu , Addie Woicik , Linda G. Shapiro , Marian Blazes , Yue Wu , Verena Steffen , Catherine Cukras , Cecilia S. Lee , Miao Zhang , Aaron Y. Lee , Sheng Wang

Multi-modal large language models (MLLMs) have achieved remarkable success in fine-grained visual understanding across a range of tasks. However, they often encounter significant challenges due to inadequate alignment for fine-grained…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Wei Wang , Zhaowei Li , Qi Xu , Linfeng Li , YiQing Cai , Botian Jiang , Hang Song , Xingcan Hu , Pengyu Wang , Li Xiao

One critical challenge in 6D object pose estimation from a single RGBD image is efficient integration of two different modalities, i.e., color and depth. In this work, we tackle this problem by a novel Deep Fusion Transformer~(DFTr) block…

计算机视觉与模式识别 · 计算机科学 2023-08-11 Jun Zhou , Kai Chen , Linlin Xu , Qi Dou , Jing Qin