中文
相关论文

相关论文: Dolphin v1.0 Technical Report

200 篇论文

Multimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Jiaqi Tang , Jianmin Chen , Wei Wei , Xiaogang Xu , Runtao Liu , Xiangyu Wu , Qipeng Xie , Jiafei Wu , Lei Zhang , Qifeng Chen

Ultrasound imaging is widely used in clinical diagnosis due to its non-invasive nature and real-time capabilities. However, traditional ultrasound diagnostics relies heavily on physician expertise and is often hampered by suboptimal image…

图像与视频处理 · 电气工程与系统科学 2025-12-18 Yuncheng Jiang , Chun-Mei Feng , Jinke Ren , Jun Wei , Zixun Zhang , Yiwen Hu , Yunbi Liu , Rui Sun , Xuemei Tang , Juan Du , Xiang Wan , Yong Xu , Bo Du , Xin Gao , Guangyu Wang , Shaohua Zhou , Shuguang Cui , Zhen Li

Measurement-critical ultrasound tasks often depend on a small anatomical region, making global reconstruction metrics an unreliable proxy for clinical fidelity. We propose an ROI-aware representation learning framework and instantiate it…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Ines Abbes , Mahmood Alzubaidi , Mowafa Househ , Khalid Alyafei , Marco Agus , Samir Brahim Belhaouari

DeepSeek-R1 is a cutting-edge open-source large language model (LLM) developed by DeepSeek, showcasing advanced reasoning capabilities through a hybrid architecture that integrates mixture of experts (MoE), chain of thought (CoT) reasoning,…

计算与语言 · 计算机科学 2025-06-03 Jiancheng Ye , Sophie Bronstein , Jiarui Hai , Malak Abu Hashish

Annotation and labeling of images are some of the biggest challenges in applying deep learning to medical data. Current processes are time and cost-intensive and, therefore, a limiting factor for the wide adoption of the technology.…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Manuel Zahn , Douglas P. Perrin

Deep-learning (DL) algorithms are becoming the standard for processing ultrasound (US) fetal images. Despite a large number of survey papers already present in this field, most of them are focusing on a broader area of medical-image…

图像与视频处理 · 电气工程与系统科学 2022-11-08 Maria Chiara Fiorentino , Francesca Pia Villani , Mariachiara Di Cosmo , Emanuele Frontoni , Sara Moccia

The development of audio foundation models has accelerated rapidly since the emergence of GPT-4o. However, the lack of comprehensive evaluation has become a critical bottleneck for further progress in the field, particularly in audio…

We introduce NVLM 1.0, a family of frontier-class multimodal large language models (LLMs) that achieve state-of-the-art results on vision-language tasks, rivaling the leading proprietary models (e.g., GPT-4o) and open-access models (e.g.,…

Despite their success, current training pipelines for reasoning VLMs focus on a limited range of tasks, such as mathematical and logical reasoning. As a result, these models face difficulties in generalizing their reasoning capabilities to…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Yuheng Zha , Kun Zhou , Yujia Wu , Yushu Wang , Jie Feng , Zhi Xu , Shibo Hao , Zhengzhong Liu , Eric P. Xing , Zhiting Hu

Analyzing underwater fish imagery is critical for ecological monitoring but remains difficult due to visual degradation and costly annotations. We introduce FishDetector-R1, a unified MLLM-based framework for fish detection, segmentation,…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Yi Liu , Jingyu Song , Vedanth Kallakuri , Katherine A. Skinner

Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm. Although several multimodal reasoning models have been explored in the medical domain, most of them…

人工智能 · 计算机科学 2025-09-11 Ruiqi Wu , Yuang Yao , Tengfei Ma , Chenran Zhang , Na Su , Tao Zhou , Geng Chen , Wen Fan , Yi Zhou

Effective reasoning remains a core challenge for large language models (LLMs) in the financial domain, where tasks often require domain-specific knowledge, precise numerical calculations, and strict adherence to compliance rules. We propose…

人工智能 · 计算机科学 2025-04-23 Jie Zhu , Qian Chen , Huaixia Dou , Junhui Li , Lifan Guo , Feng Chen , Chi Zhang

Ultrasound (US) report generation is a challenging task due to the variability of US images, operator dependence, and the need for standardized text. Unlike X-ray and CT, US imaging lacks consistent datasets, making automation difficult. In…

图像与视频处理 · 电气工程与系统科学 2025-05-20 Peixuan Ge , Tongkun Su , Faqin Lv , Baoliang Zhao , Peng Zhang , Chi Hong Wong , Liang Yao , Yu Sun , Zenan Wang , Pak Kin Wong , Ying Hu

The diagnosis of pathological images is often limited by expert availability and regional disparities, highlighting the importance of automated diagnosis using Vision-Language Models (VLMs). Traditional multimodal models typically emphasize…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Jianyu Wu , Hao Yang , Xinhua Zeng , Guibing He , Zhiyu Chen , Zihui Li , Xiaochuan Zhang , Yangyang Ma , Run Fang , Yang Liu

Audio-visual speech separation (AVSS) methods leverage visual cues to extract target speech and have demonstrated strong separation quality in noisy acoustic environments. However, these methods usually involve a large number of parameters…

声音 · 计算机科学 2026-03-12 Kai Li , Kejun Gao , Xiaolin Hu

This study proposes a novel perspective on multimodal deep learning for biomedical signal classification, systematically analyzing how complementary feature domains impact model performance. While fusing multiple domains often presumes…

机器学习 · 计算机科学 2025-08-05 Timothy Oladunni , Alex Wong

A ubiquitous task in processing electronic medical data is the assignment of standardized codes representing diagnoses and/or procedures to free-text documents such as medical reports. This is a difficult natural language processing task…

Reasoning is a critical frontier for advancing medical image analysis, where transparency and trustworthiness play a central role in both clinician trust and regulatory approval. Although Medical Visual Language Models (VLMs) show promise…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Jiazhen Pan , Che Liu , Junde Wu , Fenglin Liu , Jiayuan Zhu , Hongwei Bran Li , Chen Chen , Cheng Ouyang , Daniel Rueckert

Recent advances in clinical AI have enabled remarkable progress across many clinical domains. However, existing benchmarks and models are primarily limited to a small set of modalities and tasks, which hinders the development of large-scale…

机器学习 · 计算机科学 2025-03-21 Wei Dai , Peilin Chen , Malinda Lu , Daniel Li , Haowen Wei , Hejie Cui , Paul Pu Liang

Long-horizon video-audio reasoning and fine-grained pixel understanding impose conflicting requirements on omnimodal models: dense temporal coverage demands many low-resolution frames, whereas precise grounding calls for high-resolution…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Hao Zhong , Muzhi Zhu , Zongze Du , Zheng Huang , Canyu Zhao , Mingyu Liu , Wen Wang , Hao Chen , Chunhua Shen