中文
相关论文

相关论文: Vision-Language Model for Accurate Crater Detectio…

200 篇论文

Open-vocabulary object detection focusing on detecting novel categories guided by natural language. In this report, we propose Open-Vocabulary Light-Weighted Detection Transformer (OVLW-DETR), a deployment friendly open-vocabulary detector…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Yu Wang , Xiangbo Su , Qiang Chen , Xinyu Zhang , Teng Xi , Kun Yao , Errui Ding , Gang Zhang , Jingdong Wang

The deployment of vision-language models (VLMs) in dermatology is hindered by the trilemma of high computational costs, extreme data scarcity, and the black-box nature of deep learning. To address these challenges, we present SkinCLIP-VL, a…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Zhixiang Lu , Shijie Xu , Kaicheng Yan , Xuyue Cai , Chong Zhang , Yulong Li , Angelos Stefanidis , Anh Nguyen , Jionglong Su

We introduce a new computer aided detection and diagnosis system for lung cancer screening with low-dose CT scans that produces meaningful probability assessments. Our system is based entirely on 3D convolutional neural networks and…

计算机视觉与模式识别 · 计算机科学 2020-01-22 Onur Ozdemir , Rebecca L. Russell , Andrew A. Berlin

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-language backbones. However, most existing VLAs rely primarily on 2D visual representations,…

机器人学 · 计算机科学 2026-05-21 Shizhe Chen , Paul Pacaud , Cordelia Schmid

The automotive industry is currently expanding digital display options with every new model that comes onto the market. This entails not just an expansion in dimensions, resolution, and customization choices, but also the capability to…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Cornelius Bürkle , Fabian Oboril , Kay-Ulrich Scholl

Autonomous drones capable of interpreting and executing high-level language instructions in unstructured environments remain a long-standing goal. Yet existing approaches are constrained by their dependence on hand-crafted skills, extensive…

机器人学 · 计算机科学 2026-05-19 Qianzhong Chen , Naixiang Gao , Suning Huang , JunEn Low , Timothy Chen , Jiankai Sun , Mac Schwager

Vision-Language-Action (VLA) models have emerged as a dominant paradigm for generalist robotic manipulation, unifying perception and control within a single end-to-end architecture. However, despite their success in controlled environments,…

Accurate prediction of malignancy in renal tumors is crucial for informing clinical decisions and optimizing treatment strategies. However, existing imaging modalities lack the necessary accuracy to reliably predict malignancy before…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Zhengkang Fan , Chengkun Sun , Russell Terry , Jie Xu , Longin Jan Latecki

Landing robotic spacecrafts and humans on Mars has become one of the inevitable technological necessities for humans. Effectuating perfect Mars expedition requires landing of enormous cargoes, crewed modules, ascent vehicles, and scientific…

科普物理 · 物理学 2020-08-31 Malaya Kumar Biswal M , Ramesh Naidu Annavarapu

Incorporating deep learning (DL) classification models into unmanned aerial vehicles (UAVs) can significantly augment search-and-rescue operations and disaster management efforts. In such critical situations, the UAV's ability to promptly…

图像与视频处理 · 电气工程与系统科学 2023-06-21 Gao Yu Lee , Tanmoy Dam , Md Meftahul Ferdaus , Daniel Puiu Poenar , Vu N. Duong

The condition monitoring (CM) of synthetic fibre ropes (SFRs) used in offshore, maritime, and industrial settings demands more than a classifier: inspectors need continuous severity estimates, maintenance recommendations, anomaly flags,…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Anju Rani , Daniel Ortiz-Arroyo , Petar Durdevic

Vision-Language-Action (VLA) models offer a compelling framework for tackling complex robotic manipulation tasks, but they are often expensive to train. In this paper, we propose a novel VLA approach that leverages the competitive…

机器人学 · 计算机科学 2025-12-23 Max Argus , Jelena Bratulic , Houman Masnavi , Maxim Velikanov , Nick Heppert , Abhinav Valada , Thomas Brox

The selection of an optimal pacing site, which is ideally scar-free and late activated, is critical to the response of cardiac resynchronization therapy (CRT). Despite the success of current approaches formulating the detection of such late…

图像与视频处理 · 电气工程与系统科学 2022-11-14 Jiarui Xing , Shuo Wang , Kenneth C. Bilchick , Frederick H. Epstein , Amit R. Patel , Miaomiao Zhang

While originally designed for natural language processing tasks, the self-attention mechanism has recently taken various computer vision areas by storm. However, the 2D nature of images brings three challenges for applying self-attention in…

计算机视觉与模式识别 · 计算机科学 2022-07-12 Meng-Hao Guo , Cheng-Ze Lu , Zheng-Ning Liu , Ming-Ming Cheng , Shi-Min Hu

Deep neural network based object detection hasbecome the cornerstone of many real-world applications. Alongwith this success comes concerns about its vulnerability tomalicious attacks. To gain more insight into this issue, we proposea…

计算机视觉与模式识别 · 计算机科学 2020-08-20 Shengnan Hu , Yang Zhang , Sumit Laha , Ankit Sharma , Hassan Foroosh

Large language models (LLMs) demonstrate strong task-specific capabilities through fine-tuning, but merging multiple fine-tuned models often leads to degraded performance due to overlapping instruction-following components. Task Arithmetic…

计算与语言 · 计算机科学 2025-02-28 Yan-Lun Chen , Yi-Ru Wei , Chia-Yi Hsu , Chia-Mu Yu , Chun-Ying Huang , Ying-Dar Lin , Yu-Sung Wu , Wei-Bin Lee

Understanding space weather is vital for the protection of our terrestrial and space infrastructure. In order to predict space weather accurately, large amounts of data are required, particularly in the extreme ultraviolet (EUV) spectrum.…

太阳与恒星天体物理 · 物理学 2024-09-02 Manuel Indaco , Daniel Gass , William James Fawcett , Richard Galvez , Paul J. Wright , Andrés Muñoz-Jaramillo

Free space optical communication techniques have been the subject of numerous investigations in recent years, with multiple missions expected to fly in the near future. Existing methods require high pointing accuracies, drastically driving…

网络与互联网体系结构 · 计算机科学 2018-01-04 Sihao Huang , Haowen Lin

Current Vision-Language-Action (VLA) models rely on fixed computational depth, expending the same amount of compute on simple adjustments and complex multi-step manipulation. While Chain-of-Thought (CoT) prompting enables variable…

机器人学 · 计算机科学 2026-02-10 Yalcin Tur , Jalal Naghiyev , Haoquan Fang , Wei-Chuan Tsai , Jiafei Duan , Dieter Fox , Ranjay Krishna

Deep learning brought boosts to auto diabetic retinopathy (DR) diagnosis, thus, greatly helping ophthalmologists for early disease detection, which contributes to preventing disease deterioration that may eventually lead to blindness. It…

图像与视频处理 · 电气工程与系统科学 2024-08-15 Xue Xia , Kun Zhan , Yuming Fang , Wenhui Jiang , Fei Shen