English
Related papers

Related papers: Video CLIP Model for Multi-View Echocardiography I…

200 papers

Semi-supervised image classification has shown substantial progress in learning from limited labeled data, but recent advances remain largely untested for clinical applications. Motivated by the urgent need to improve timely diagnosis of…

Computer Vision and Pattern Recognition · Computer Science 2021-08-03 Zhe Huang , Gary Long , Benjamin Wessler , Michael C. Hughes

Echocardiogram video plays a crucial role in analysing cardiac function and diagnosing cardiac diseases. Current deep neural network methods primarily aim to enhance diagnosis accuracy by incorporating prior knowledge, such as segmenting…

Image and Video Processing · Electrical Eng. & Systems 2024-10-29 Jiewen Yang , Yiqun Lin , Bin Pu , Jiarong Guo , Xiaowei Xu , Xiaomeng Li

We present a knowledge augmentation strategy for assessing the diagnostic groups and gait impairment from monocular gait videos. Based on a large-scale pre-trained Vision Language Model (VLM), our model learns and improves visual, textual,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Diwei Wang , Kun Yuan , Candice Muller , Frédéric Blanc , Nicolas Padoy , Hyewon Seo

Listening to heart and lung sounds - auscultation - is one of the first and most fundamental steps in a clinical examination. Despite being fast and non-invasive, it demands years of experience to interpret subtle audio cues. Recent deep…

Machine Learning · Computer Science 2026-03-03 Yishan Wang , Tsai-Ning Wang , Mathias Funk , Aaqib Saeed

Vision-Language Models (VLMs) are able to process increasingly longer videos. Yet, important visual information is easily lost throughout the entire context and missed by VLMs. Also, it is important to design tools that enable…

Computation and Language · Computer Science 2026-01-09 Galann Pennec , Zhengyuan Liu , Nicholas Asher , Philippe Muller , Nancy F. Chen

This technical report summarizes our method for the Video-And-Language Understanding Evaluation (VALUE) challenge (https://value-benchmark.github.io/challenge\_2021.html). We propose a CLIP-Enhanced method to incorporate the image-text…

Computer Vision and Pattern Recognition · Computer Science 2021-10-15 Guohao Li , Feng He , Zhifan Feng

Advances in self-supervised learning (SSL) have shown that self-supervised pretraining on medical imaging data can provide a strong initialization for downstream supervised classification and segmentation. Given the difficulty of obtaining…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Gregory Holste , Evangelos K. Oikonomou , Bobak J. Mortazavi , Zhangyang Wang , Rohan Khera

Medical vision-and-language models (MVLMs) have attracted substantial interest due to their capability to offer a natural language interface for interpreting complex medical data. Their applications are versatile and have the potential to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Qi Chen , Ruoshan Zhao , Sinuo Wang , Vu Minh Hieu Phan , Anton van den Hengel , Johan Verjans , Zhibin Liao , Minh-Son To , Yong Xia , Jian Chen , Yutong Xie , Qi Wu

Transthoracic echo is one of the most common means of cardiac studies in the clinical routines. During the echo exam, the sonographer captures a set of standard cross sections (echo views) of the heart. Each 2D echo view cuts through the 3D…

Image and Video Processing · Electrical Eng. & Systems 2019-12-09 Amir H. Abdi , Mohammad H. Jafari , Sidney Fels , Theresa Tsang , Purang Abolmaesumi

Accurate segmentation is essential for echocardiography-based assessment of cardiovascular diseases (CVDs). However, the variability among sonographers and the inherent challenges of ultrasound images hinder precise segmentation. By…

Computer Vision and Pattern Recognition · Computer Science 2023-10-06 Rabin Adhikari , Manish Dhakal , Safal Thapaliya , Kanchan Poudel , Prasiddha Bhandari , Bishesh Khanal

In this paper, we initiate an attempt of developing an end-to-end chat-centric video understanding system, coined as VideoChat. It integrates video foundation models and large language models via a learnable neural interface, excelling in…

Computer Vision and Pattern Recognition · Computer Science 2024-01-05 KunChang Li , Yinan He , Yi Wang , Yizhuo Li , Wenhai Wang , Ping Luo , Yali Wang , Limin Wang , Yu Qiao

This paper presents Audio-Visual LLM, a Multimodal Large Language Model that takes both visual and auditory inputs for holistic video understanding. A key design is the modality-augmented training, which involves the integration of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Fangxun Shu , Lei Zhang , Hao Jiang , Cihang Xie

Synthetic data generation represents a significant advancement in boosting the performance of machine learning (ML) models, particularly in fields where data acquisition is challenging, such as echocardiography. The acquisition and labeling…

Machine Learning · Computer Science 2025-08-28 Nima Kondori , Hanwen Liang , Hooman Vaseli , Bingyu Xie , Christina Luong , Purang Abolmaesumi , Teresa Tsang , Renjie Liao

This paper presents VideoStreaming, an advanced vision-language large model (VLLM) for video understanding, that capably understands arbitrary-length video with a constant number of video tokens streamingly encoded and adaptively selected.…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Rui Qian , Xiaoyi Dong , Pan Zhang , Yuhang Zang , Shuangrui Ding , Dahua Lin , Jiaqi Wang

Automating crash video analysis is essential to leverage the growing availability of driving video data for traffic safety research and accountability attribution in autonomous driving. Crash video analysis is a challenging multitask…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Kaidi Liang , Ke Li , Xianbiao Hu , Ruwen Qin

Analysis of cardiac ultrasound images is commonly performed in routine clinical practice for quantification of cardiac function. Its increasing automation frequently employs deep learning networks that are trained to predict disease or…

Computer Vision and Pattern Recognition · Computer Science 2021-08-09 Agisilaos Chartsias , Shan Gao , Angela Mumith , Jorge Oliveira , Kanwal Bhatia , Bernhard Kainz , Arian Beqiri

Facial Expression Recognition (FER) is a crucial task in affective computing, but its conventional focus on the seven basic emotions limits its applicability to the complex and expanding emotional spectrum. To address the issue of new and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Niki Maria Foteinopoulou , Ioannis Patras

Vision-language foundation models have emerged as powerful general-purpose representation learners with strong potential for multimodal understanding, but their deterministic embeddings often fail to provide the reliability required for…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Ahmad Elallaf , Yu Zhang , Yuktha Priya Masupalli , Jeong Yang , Young Lee , Zechun Cao , Gongbo Liang

Multi-lead electrocardiogram (ECG) provides clinical information of heartbeats from several fixed viewpoints determined by the lead positioning. However, it is often not satisfactory to visualize ECG signals in these fixed and limited…

Signal Processing · Electrical Eng. & Systems 2022-04-05 Jintai Chen , Xiangshang Zheng , Hongyun Yu , Danny Z. Chen , Jian Wu

With the availability of large-scale, comprehensive, and general-purpose vision-language (VL) datasets such as MSCOCO, vision-language pre-training (VLP) has become an active area of research and proven to be effective for various VL tasks…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Li Xu , Bo Liu , Ameer Hamza Khan , Lu Fan , Xiao-Ming Wu
‹ Prev 1 3 4 5 6 7 10 Next ›