English
Related papers

Related papers: Echo-CoPilot: A Multiple-Perspective Agentic Frame…

200 papers

In clinical practice of echocardiography examinations, multiple planes containing the heart structures of different view are usually required in screening, diagnosis and treatment of cardiac disease. AI models for echocardiography have to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Jiongtong Hu , Wei Zhuo , Jun Cheng , Yingying Liu , Wufeng Xue , Dong Ni

Text-to-image synthesis has made remarkable progress, yet accurately interpreting complex and lengthy prompts remains challenging, often resulting in semantic inconsistencies and missing details. Existing solutions, such as fine-tuning, are…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Wen Ye , Zhaocheng Liu , Yuwei Gui , Tingyu Yuan , Yunyue Su , Bowen Fang , Chaoyang Zhao , Qiang Liu , Liang Wang

As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to…

Artificial Intelligence · Computer Science 2026-05-28 Yunqi Liu , Tong Niu , Zitong Wang , Zhenlong Dai , Yuqi Qing , Weiqiang Wang , Jian Liu

Echocardiography segmentation for cardiac analysis is time-consuming and resource-intensive due to the variability in image quality and the necessity to process scans from various standard views. While current automated segmentation methods…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Sekeun Kim , Hui Ren , Peng Guo , Abder-Rahman Ali , Patrick Zhang , Kyungsang Kim , Xiang Li , Quanzheng Li

Understanding the interaction of neural and cardiac systems during cognitive activity is critical to advancing physiological computing. Although EEG has been the gold standard for assessing mental workload, its limited portability restricts…

Machine Learning · Computer Science 2026-01-06 Akshay Sasi , Malavika Pradeep , Nusaibah Farrukh , Rahul Venugopal , Elizabeth Sherly

Diagnosing a whole-slide image is an interactive, multi-stage process of changing magnification and moving between fields. Although recent pathology foundation models demonstrated superior performances, practical agentic systems that decide…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Sheng Wang , Ruiming Wu , Charles Herndon , Yihang Liu , Shunsuke Koga , Jeanne Shen , Zhi Huang

Interpreting and communicating electrocardiogram (ECG) findings are crucial yet challenging tasks in cardiovascular diagnosis, traditionally requiring significant expertise and precise clinical communication. This paper introduces…

Signal Processing · Electrical Eng. & Systems 2025-10-30 Koustav Mallick , Neel Singh , Mohammedreza Hajiarbabi

Multimedia event extraction (M2E2) aims to predict triggers, ground arguments across text and images, and then assemble them into schema-consistent event records. Recent LLM-based approaches have shown strong potential for M2E2, but their…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Hailong Chu , Hongbing Li , Yunlong Chu , Shutai Huang , Xingyue Zhang , Tinghe Yan , Jinsong Zhang , Shuo Zhang , Lei Li

Although the heart has complex three-dimensional (3D) anatomy, conventional medical imaging with cardiac ultrasound relies on a series of 2D videos showing individual cardiac structures. 3D echocardiography is a developing modality that now…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Milos Vukadinovic , Hirotaka Ieki , Yuki Sahashi , David Ouyang , Bryan He

Deploying Audio-Language Models (Audio-LLMs) on edge infrastructure exposes a persistent tension between perception depth and computational efficiency. Lightweight local models tend to produce passive perception - generic summaries that…

Sound · Computer Science 2026-01-23 Hengfan Zhang , Yueqian Lin , Hai Helen Li , Yiran Chen

Objective: This work aims at providing a new method for the automatic detection of atrial fibrillation, other arrhythmia and noise on short single lead ECG signals, emphasizing the importance of the interpretability of the classification…

Artificial Intelligence · Computer Science 2021-12-09 Tomás Teijeiro , Constantino A. García , Daniel Castro , Paulo Félix

As large language models empower healthcare, intelligent clinical decision support has developed rapidly. Longitudinal electronic health records (EHR) provide essential temporal evidence for accurate clinical diagnosis and analysis.…

Computation and Language · Computer Science 2026-05-15 Zihan Deng , Xiaozhen Zhong , Chuanzhi Xu

The accurate interpretation of Electrocardiogram (ECG) signals is pivotal for diagnosing cardiovascular diseases. Integrating ECG signals with accompanying textual reports further holds immense potential to enhance clinical diagnostics by…

Machine Learning · Computer Science 2025-05-08 Hung Manh Pham , Aaqib Saeed , Dong Ma

Autonomous agent frameworks still struggle to reconcile long-term experiential learning with real-time, context-sensitive decision-making. In practice, this gap appears as static cognition, rigid workflow dependence, and inefficient context…

Artificial Intelligence · Computer Science 2026-03-11 Xiaoxing Wang , Ning Liao , Shikun Wei , Chen Tang , Feiyu Xiong

Electrocardiography (ECG) serves as an indispensable diagnostic tool in clinical practice, yet existing multimodal large language models (MLLMs) remain unreliable for ECG interpretation, often producing plausible but clinically incorrect…

Computation and Language · Computer Science 2026-05-26 Jiarui Jin , Haoyu Wang , Xingliang Wu , Xiaocheng Fang , Xiang Lan , Zihan Wang , Deyun Zhang , Bo Liu , Yingying Zhang , Xian Wu , Hongyan Li , Shenda Hong

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feeding frame-level…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Noriyuki Kugo , Xiang Li , Zixin Li , Ashish Gupta , Arpandeep Khatua , Nidhish Jain , Chaitanya Patel , Yuta Kyuragi , Yasunori Ishii , Masamoto Tanabiki , Kazuki Kozuka , Ehsan Adeli

Maintaining narrative coherence and visual consistency remains a central challenge in open-domain video generation. Existing text-to-video models often treat each shot independently, resulting in identity drift, scene inconsistency, and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Qinglin Zeng , Kaitong Cai , Ruiqi Chen , Qinhan Lv , Keze Wang

We introduce EchoXFlow, a clinical echocardiography dataset for learning from ultrasound in its native acquisition geometry rather than from scan-converted Cartesian videos. Existing public datasets offer limited opportunities to study…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Elias Stenhede , Joanna Sulkowska , Eivind Bjørkan Orstad , Henrik Schirmer , Arian Ranjbar

AI-powered coding assistants, like GitHub Copilot, are increasingly used to boost developers' productivity. However, their output quality hinges on the contextual richness of the prompts. Meanwhile, gaze behaviour carries rich cognitive…

Human-Computer Interaction · Computer Science 2025-11-12 Yasmine Elfares , Gül Çalikli , Mohamed Khamis

Controllable medical video generation has achieved remarkable progress, but it still lacks interpretability, which requires the alignment of generated contents with physical priors and faithful clinical manifestations. To push the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Junhu Fu , Ke Chen , Weidong Guo , Shuyu Liang , Jie Xu , Chen Ma , Kehao Wang , Shengli Lin , Zeju Li , Yuanyuan Wang , Yi Guo , Shuo Li
‹ Prev 1 3 4 5 6 7 10 Next ›