English
Related papers

Related papers: StereoMath: An Accessible and Musical Equation Edi…

200 papers

Training vision-language models on cognitively-plausible amounts of data requires rethinking how models integrate multimodal information. Within the constraints of the Vision track for the BabyLM Challenge 2025, we propose a lightweight…

Artificial Intelligence · Computer Science 2025-10-10 Bianca-Mihaela Ganescu , Suchir Salhan , Andrew Caines , Paula Buttery

Advances in large language models (LLMs) offer new possibilities for enhancing math education by automating support for both teachers and students. While prior work has focused on generating math problems and high-quality distractors, the…

Artificial Intelligence · Computer Science 2025-03-11 Jaewook Lee , Jeongah Lee , Wanyong Feng , Andrew Lan

Current multimodal large language models (MLLMs) often underperform on mathematical problem-solving tasks that require fine-grained visual understanding. The limitation is largely attributable to inadequate perception of geometric…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Shan Zhang , Aotian Chen , Yanpeng Sun , Jindong Gu , Yi-Yu Zheng , Piotr Koniusz , Kai Zou , Anton van den Hengel , Yuan Xue

GPS and smartphones enable users to place location-based annotations, capturing rich environmental context. Previous research demonstrates that blind and low vision (BLV) people can use annotations to explore unfamiliar areas. However,…

Auxiliary lines are essential for solving complex geometric problems but remain challenging for large vision-language models (LVLMs). Recent attempts construct auxiliary lines via code-driven rendering, a strategy that relies on accurate…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Shasha Guo , Liang Pang , Xi Wang , Yanling Wang , Huawei Shen , Jing Zhang

Recent advances in Vision-Language Models (VLMs) have achieved impressive progress in multimodal mathematical reasoning. Yet, how much visual information truly contributes to reasoning remains unclear. Existing benchmarks report strong…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yuandong Wang , Yao Cui , Yuxin Zhao , Zhen Yang , Yangfu Zhu , Zhenzhou Shao

Communicating linear algebra in written form is challenging: mathematicians must choose between writing in languages that produce well-formatted but semantically-underdefined representations such as LaTeX; or languages with well-defined…

Programming Languages · Computer Science 2021-09-28 Yong Li , Shoaib Kamil , Alec Jacobson , Yotam Gingold

Typing mathematics is sometimes difficult with text editor functions for students with motor impairment and other associated impairments (visual, cognitive). Based on the HandiMathKey software keyboard, a user-centred design method…

Human-Computer Interaction · Computer Science 2023-11-27 Frédéric Vella , Nathalie Dubus , Eloise Grolleau , Marjorie Deleau , Cécile Malet , Christine Gallard , Véronique Ades , Nadine Vigouroux

Computer-based tests (CBTs) play an important role in the professional career of any person. Universities use CBTs for admissions. Further, many large courses use CBTs for evaluation and grading. Almost all software companies use CBTs to…

Human-Computer Interaction · Computer Science 2019-05-07 Pawan Kr Patel , Amey Karkare

Visuals are valuable tools for teaching math word problems (MWPs), helping young learners interpret textual descriptions into mathematical expressions before solving them. However, creating such visuals is labor-intensive and there is a…

Computation and Language · Computer Science 2025-06-05 Junling Wang , Anna Rutkiewicz , April Yi Wang , Mrinmaya Sachan

Despite the recent surge of research efforts to make data visualizations accessible to people who are blind or have low vision (BLV), how to support BLV people's data analysis remains an important and challenging question. As refreshable…

Human-Computer Interaction · Computer Science 2025-01-17 Samuel Reinders , Matthew Butler , Ingrid Zukerman , Bongshin Lee , Lizhen Qu , Kim Marriott

Multimodal Large Language Models (MLLMs) hold immense promise as assistive technologies for the blind and visually impaired (BVI) community. However, we identify a critical failure mode that undermines their trustworthiness in real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xiantao Zhang

Online shopping has become a valuable modern convenience, but blind or low vision (BLV) users still face significant challenges using it, because of: 1) inadequate image descriptions and 2) the inability to filter large amounts of…

Human-Computer Interaction · Computer Science 2021-02-02 Ruolin Wang , Zixuan Chen , Mingrui "Ray" Zhang , Zhaoheng Li , Zhixiu Liu , Zihan Dang , Chun Yu , Xiang "Anthony" Chen

The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning. A recent line of work explores learning spatial reasoning directly from multi-view images,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Kanghee Lee , Injae Lee , Minseok Kwak , Jungi Hong , Kwonyoung Ryu , Jaesik Park

Existing benchmarks for evaluating mathematical reasoning in large language models (LLMs) rely primarily on competition problems, formal proofs, or artificially challenging questions -- failing to capture the nature of mathematics…

Artificial Intelligence · Computer Science 2025-10-21 Jie Zhang , Cezara Petrui , Kristina Nikolić , Florian Tramèr

Real-world environments evolve continuously, yet blind and low-vision (BLV) individuals often have limited access to understanding how they change over time. Unexpected or relocated objects, layout modifications, and content updates (e.g.,…

Human-Computer Interaction · Computer Science 2026-04-28 Ruei-Che Chang , Xirui Jiang , Rosiana Natalie , Hao Chen , Vlad Roznyatovskiy , Jianzhong Zhang , Kang G. Shin , Ke Sun , Anhong Guo

Blind and Visually Impaired (BVI) Individuals face significant challenges in science due to the discipline's reliance on visual elements such as graphs, diagrams, and laboratory work. Traditional learning materials, such as Braille and…

Physics Education · Physics 2025-11-25 Ludovic Petitdemange , Salomé Nashed

Recent advancements in language and vision assistants have showcased impressive capabilities but suffer from a lack of transparency, limiting broader research and reproducibility. While open-source models handle general image tasks…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Geewook Kim , Minjoon Seo

Writing is a universal cultural technology that reuses vision for symbolic communication. Humans display striking resilience: we readily recognize words even when characters are fragmented, fused, or partially occluded. This paper…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Jie Zhang , Ting Xu , Gelei Deng , Runyi Hu , Han Qiu , Tianwei Zhang , Qing Guo , Ivor Tsang

The mathematical capabilities of Multi-modal Large Language Models (MLLMs) remain under-explored with three areas to be improved: visual encoding of math diagrams, diagram-language alignment, and chain-of-thought (CoT) reasoning. This draws…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Renrui Zhang , Xinyu Wei , Dongzhi Jiang , Ziyu Guo , Shicheng Li , Yichi Zhang , Chengzhuo Tong , Jiaming Liu , Aojun Zhou , Bin Wei , Shanghang Zhang , Peng Gao , Chunyuan Li , Hongsheng Li
‹ Prev 1 4 5 6 7 8 10 Next ›