中文
相关论文

相关论文: ChartFormer: A Large Vision Language Model for Con…

200 篇论文

Understanding charts requires models to jointly reason over geometric visual patterns, structured numerical data, and natural language -- a capability where current vision-language models (VLMs) remain limited. We introduce ChartNet, a…

Pretraining Vision Transformers (ViTs) has achieved great success in visual recognition. A following scenario is to adapt a ViT to various image and video recognition tasks. The adaptation is challenging because of heavy computation and…

计算机视觉与模式识别 · 计算机科学 2022-10-18 Shoufa Chen , Chongjian Ge , Zhan Tong , Jiangliu Wang , Yibing Song , Jue Wang , Ping Luo

Despite the ubiquity of visualization examples published on the web, retargeting existing custom chart implementations to new datasets remains difficult, time-intensive, and tedious. The adaptation process assumes author familiarity with…

人机交互 · 计算机科学 2026-01-21 Luke S. Snyder , Chenglong Wang , Steven M. Drucker

Chart understanding requires models to effectively analyze and reason about numerical data, textual elements, and complex visual components. Our observations reveal that the perception capabilities of existing large vision-language models…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Junteng Liu , Weihao Zeng , Xiwen Zhang , Yijun Wang , Zifei Shan , Junxian He

Simulators perform an important role in prototyping, debugging, and benchmarking new advances in robotics and learning for control. Although many physics engines exist, some aspects of the real world are harder than others to simulate. One…

机器人学 · 计算机科学 2022-02-14 Shaoxiong Wang , Mike Lambeta , Po-Wei Chou , Roberto Calandra

The use of natural language interfaces (NLIs) to create charts is becoming increasingly popular due to the intuitiveness of natural language interactions. One key challenge in this approach is to accurately capture user intents and…

人机交互 · 计算机科学 2025-01-22 Yuan Tian , Weiwei Cui , Dazhen Deng , Xinjing Yi , Yurun Yang , Haidong Zhang , Yingcai Wu

Tactile maps are commonly used to give visually impaired users access to geographical representations. Although those relief maps are efficient tools for acquisition of spatial knowledge, they present several limitations and issues such as…

人机交互 · 计算机科学 2022-09-01 Julie Ducasse , Anke Brock , Christophe Jouffrais

Transformers have demonstrated remarkable performance in natural language processing and computer vision. However, existing vision Transformers struggle to learn from limited medical data and are unable to generalize on diverse medical…

图像与视频处理 · 电气工程与系统科学 2023-04-06 Yunhe Gao , Mu Zhou , Di Liu , Zhennan Yan , Shaoting Zhang , Dimitris N. Metaxas

Traditional accessibility methods like alternative text and data tables typically underrepresent data visualization's full potential. Keyboard-based chart navigation has emerged as a potential solution, yet efficient data exploration…

人机交互 · 计算机科学 2024-08-20 Joshua Gorniak , Yoon Kim , Donglai Wei , Nam Wook Kim

Tables form a central component in both exploratory data analysis and formal reporting procedures across many industries. These tables are often complex in their conceptual structure and in the computations that generate their individual…

统计计算 · 统计学 2023-06-30 Gabriel Becker , Adrian Waddell

Data extraction from line-chart images is an essential component of the automated document understanding process, as line charts are a ubiquitous data visualization format. However, the amount of visual and structural variations in…

计算机视觉与模式识别 · 计算机科学 2023-05-04 Jay Lal , Aditya Mitkari , Mahesh Bhosale , David Doermann

Text documents with numerical values involved are widely used in various applications such as scientific research, economy, public health and journalism. However, it is difficult for readers to quickly interpret such data-involved texts and…

人机交互 · 计算机科学 2024-11-08 Songheng Zhang , Lei Wang , Toby Jia-Jun Li , Qiaomu Shen , Yixin Cao , Yong Wang

This manuscript introduces SARFormer, a modified Vision Transformer (ViT) architecture designed for processing one or multiple synthetic aperture radar (SAR) images. Given the complex image geometry of SAR data, we propose an acquisition…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Jonathan Prexl , Michael Recla , Michael Schmitt

Charts provide visual representations of data and are widely used for analyzing information, addressing queries, and conveying insights to others. Various chart-related downstream tasks have emerged recently, such as question-answering and…

计算与语言 · 计算机科学 2024-03-15 Ahmed Masry , Mehrad Shahmohammadi , Md Rizwan Parvez , Enamul Hoque , Shafiq Joty

Diagrams often appear as node-link representations in many contexts, such as taxonomies, mind maps and networks in textbooks. Despite their pervasiveness, they present significant accessibility challenges for blind and low-vision people. To…

人机交互 · 计算机科学 2024-05-08 Yichun Zhao , Miguel A. Nacenta , Mahadeo A. Sukhai , Sowmya Somanath

Automatic chart to text summarization is an effective tool for the visually impaired people along with providing precise insights of tabular data in natural language to the user. A large and well-structured dataset is always a key part for…

Often, the needs and visual abilities differ between the annotator group and the end user group. Generating detailed diagram descriptions for blind and low-vision (BLV) users is one such challenging domain. Sighted annotators could describe…

人工智能 · 计算机科学 2025-03-18 Wan Ju Kang , Eunki Kim , Na Min An , Sangryul Kim , Haemin Choi , Ki Hoon Kwak , James Thorne

Our previous interview study explores the needs and uses of diagrammatic information by the Blind and Low Vision (BLV) community, resulting in a framework called the Ladder of Diagram Access. The framework outlines five levels of…

人机交互 · 计算机科学 2024-04-02 Yichun Zhao , Miguel A. Nacenta

Infographic charts are a powerful medium for communicating abstract data by combining visual elements (e.g., charts, images) with textual information. However, their visual and structural richness poses challenges for large vision-language…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Zhen Li , Duan Li , Yukai Guo , Xinyuan Guo , Bowen Li , Lanxi Xiao , Shenyu Qiao , Jiashu Chen , Zijian Wu , Hui Zhang , Xinhuan Shu , Shixia Liu

Scalable Vector Graphics (SVGs) are vital for modern image rendering due to their scalability and versatility. Previous SVG generation methods have focused on curve-based vectorization, lacking semantic understanding, often producing…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Juan A. Rodriguez , Abhay Puri , Shubham Agarwal , Issam H. Laradji , Pau Rodriguez , Sai Rajeswar , David Vazquez , Christopher Pal , Marco Pedersoli