中文
相关论文

相关论文: OtterHD: A High-Resolution Multi-modality Model

200 篇论文

Large language models have been widely adopted but require significant GPU memory for inference. We develop a procedure for Int8 matrix multiplication for feed-forward and attention projection layers in transformers, which cut the memory…

机器学习 · 计算机科学 2022-11-11 Tim Dettmers , Mike Lewis , Younes Belkada , Luke Zettlemoyer

In human-centered environments such as restaurants, homes, and warehouses, robots often face challenges in accurately recognizing 3D objects. These challenges stem from the complexity and variability of these environments, including diverse…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Songsong Xiong , Hamidreza Kasaei

Within the field of robotics, computer vision remains a significant barrier to progress, with many tasks hindered by inefficient vision systems. This research proposes a generalized vision module leveraging YOLOv9, a state-of-the-art…

机器人学 · 计算机科学 2025-10-16 Nicolas Pottier , Meng Cheng Lau

The modeling and simulation of high-dimensional multiscale systems is a critical challenge across all areas of science and engineering. It is broadly believed that even with today's computer advances resolving all spatiotemporal scales…

Neural networks commonly execute on hardware accelerators such as NPUs and GPUs for their size and computation overhead. These accelerators are costly and it is hard to scale their resources to handle real-time workload fluctuations. We…

机器学习 · 计算机科学 2025-10-06 Jaemin Kim , Hongjun Um , Sungkyun Kim , Yongjun Park , Jiwon Seo

Multimodal deep learning has been used to predict clinical endpoints and diagnoses from clinical routine data. However, these models suffer from scaling issues: they have to learn pairwise interactions between each piece of information in…

A large-scale labeled dataset is a key factor for the success of supervised deep learning in computer vision. However, a limited number of annotated data is very common, especially in ophthalmic image analysis, since manual annotation is…

图像与视频处理 · 电气工程与系统科学 2022-03-15 Zhiyuan Cai , Li Lin , Huaqing He , Xiaoying Tang

This tutorial explores recent advancements in multimodal pretrained and large models, capable of integrating and processing diverse data forms such as text, images, audio, and video. Participants will gain an understanding of the…

计算与语言 · 计算机科学 2024-10-10 Soyeon Caren Han , Feiqi Cao , Josiah Poon , Roberto Navigli

We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve fluent and high-quality interaction across modalities without…

Object pose estimation is a long-standing problem in computer vision. Recently, attention-based vision transformer models have achieved state-of-the-art results in many computer vision applications. Exploiting the permutation-invariant…

计算机视觉与模式识别 · 计算机科学 2023-12-14 Arul Selvam Periyasamy , Vladimir Tsaturyan , Sven Behnke

In this paper, we introduce LLaVA-Octopus, a novel video multimodal large language model. LLaVA-Octopus adaptively weights features from different visual projectors based on user instructions, enabling us to leverage the complementary…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Boyuan Sun , Jiaxing Zhao , Xiang Chen , Xihan Wei , Qibin Hou

AI tasks in the car interior like identifying and localizing externally introduced objects is crucial for response quality of personal assistants. However, computational resources of on-board systems remain highly constrained, restricting…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Sebastian Schmidt , Bálint Mészáros , Ahmet Firintepe , Stephan Günnemann

Robots and other smart devices need efficient object-based scene representations from their on-board vision systems to reason about contact, physics and occlusion. Recognized precise object models will play an important role alongside…

计算机视觉与模式识别 · 计算机科学 2020-04-10 Kentaro Wada , Edgar Sucar , Stephen James , Daniel Lenton , Andrew J. Davison

In recent years, Transformer networks are beginning to replace pure convolutional neural networks (CNNs) in the field of computer vision due to their global receptive field and adaptability to input. However, the quadratic computational…

计算机视觉与模式识别 · 计算机科学 2023-08-31 Yuwei Qiu , Kaihao Zhang , Chenxi Wang , Wenhan Luo , Hongdong Li , Zhi Jin

Optical character recognition (OCR) and multilingual text understanding remain major failure modes of multimodal large language models (MLLMs), particularly in real-world images containing cluttered layouts, small fonts, blur, occlusion,…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Qinwu Xu , Yifan Jiang , Haoyu Ren

Generalist models have achieved remarkable success in both language and vision-language tasks, showcasing the potential of unified modeling. However, effectively integrating fine-grained perception tasks like detection and segmentation into…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Hao Tang , Chenwei Xie , Haiyang Wang , Xiaoyi Bao , Tingyu Weng , Pandeng Li , Yun Zheng , Liwei Wang

Medical imaging is critical for diagnostics, but clinical adoption of advanced AI-driven imaging faces challenges due to patient variability, image artifacts, and limited model generalization. While deep learning has transformed image…

图像与视频处理 · 电气工程与系统科学 2025-06-02 Abdul-mojeed Olabisi Ilyas , Adeleke Maradesa , Jamal Banzi , Jianpan Huang , Henry K. F. Mak , Kannie W. Y. Chan

Optical Coherence Tomography Angiography (OCTA) and its derived en-face projections provide high-resolution visualization of the retinal and choroidal vasculature, which is critical for the rapid and accurate diagnosis of retinal diseases.…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Pooya Khosravi , Kun Han , Anthony T. Wu , Arghavan Rezvani , Zexin Feng , Xiaohui Xie

Most of the face recognition works focus on specific modules or demonstrate a research idea. This paper presents a pose-invariant 3D-aided 2D face recognition system (UR2D) that is robust to pose variations as large as 90? by leveraging…

计算机视觉与模式识别 · 计算机科学 2017-09-20 Xiang Xu , Pengfei Dou , Ha A. Le , Ioannis A. Kakadiaris

Optic nerve head (ONH) detection has been a crucial area of study in ophthalmology for years. However, the significant discrepancy between fundus image datasets, each generated using a single type of fundus camera, poses challenges to the…

图像与视频处理 · 电气工程与系统科学 2024-06-04 Jiayi Wang , Yi-An Mao , Xiaoyu Ma , Sicen Guo , Yuting Shao , Xiao Lv , Wenting Han , Mark Christopher , Linda M. Zangwill , Yanlong Bi , Rui Fan
‹ 上一页 1 8 9 10 下一页 ›