中文
相关论文

相关论文: Comics for Everyone: Generating Accessible Text De…

200 篇论文

As virtual 3D environments become more prevalent, equitable access is essential for blind and low-vision (BLV) users, who face challenges with spatial awareness, navigation, and interaction. Prior work has explored supplementing visual…

人机交互 · 计算机科学 2026-02-10 Xinyun Cao , Kexin Phyllis Ju , Chenglin Li , Venkatesh Potluri , Dhruv Jain

Writing a coherent and engaging story is not easy. Creative writers use their knowledge and worldview to put disjointed elements together to form a coherent storyline, and work and rework iteratively toward perfection. Automated visual…

计算与语言 · 计算机科学 2021-07-08 Chi-Yang Hsu , Yun-Wei Chu , Ting-Hao 'Kenneth' Huang , Lun-Wei Ku

The frequent need for analysts to create visualizations to derive insights from data has driven extensive research into the generation of natural Language to Visualization (NL2VIS). While recent progress in large language models (LLMs)…

人机交互 · 计算机科学 2025-12-12 Xinyu Wang , Chenwei Liang , Shunyuan Zheng , Jinyuan Liang , Guozheng Li , Yu Zhang , Chi Harold Liu

When the visual style of text is considered, a wide variety can be observed in font, color, and size. However, when a word is read, its meaning is independent of the style in which it has been written or rendered. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Xiaomeng Wang , Martha Larson , Zhengyu Zhao

Manga, or Japanese comics, is a richly multimodal narrative form that blends images and text in complex ways. Teaching large multimodal models (LMMs) to understand such narratives at a human-like level could help manga creators reflect on…

计算与语言 · 计算机科学 2026-01-27 Jeonghun Baek , Kazuki Egashira , Shota Onohara , Atsuyuki Miyai , Yuki Imajuku , Hikaru Ikuta , Kiyoharu Aizawa

This work explores a closure task in comics, a medium where visual and textual elements are intricately intertwined. Specifically, Text-cloze refers to the task of selecting the correct text to use in a comic panel, given its neighboring…

计算机视觉与模式识别 · 计算机科学 2024-03-07 Emanuele Vivoli , Joan Lafuente Baeza , Ernest Valveny Llobet , Dimosthenis Karatzas

Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages recent advancements in…

计算与语言 · 计算机科学 2025-06-11 Mohamed Gado , Towhid Taliee , Muhammad Memon , Dmitry Ignatov , Radu Timofte

Access to textual and visual information for visually impaired persons becomes very difficult with screen readers which are not adapted to different websites.This paper analyses the use of different technologies for access digital content…

人机交互 · 计算机科学 2019-11-18 Katerine Romeo , Edwige Pissaloux , Frédéric Serin

Creating meaningful visual narratives through human-AI collaboration requires understanding how text-image intertextuality emerges when textual intentions meet AI-generated visuals. We conducted a three-phase qualitative study with 15…

人机交互 · 计算机科学 2025-11-06 Mengyao Guo , Kexin Nie , Ze Gao , Black Sun , Xueyang Wang , Jinda Han , Xingting Wu

We introduce MotionScript, a novel framework for generating highly detailed, natural language descriptions of 3D human motions. Unlike existing motion datasets that rely on broad action labels or generic captions, MotionScript provides…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Payam Jome Yazdian , Rachel Lagasse , Hamid Mohammadi , Eric Liu , Li Cheng , Angelica Lim

In recent years, efforts have been made to use text information for better user profiling and item characterization in recommendations. However, text information can sometimes be of low quality, hindering its effectiveness for real-world…

人工智能 · 计算机科学 2024-02-15 Yingpeng Du , Ziyan Wang , Zhu Sun , Haoyan Chua , Hongzhi Liu , Zhonghai Wu , Yining Ma , Jie Zhang , Youchen Sun

Few images on the Web receive alt-text descriptions that would make them accessible to blind and low vision (BLV) users. Image-based NLG systems have progressed to the point where they can begin to address this persistent societal problem,…

Natural language and visualization are two complementary modalities of human communication that play a crucial role in conveying information effectively. While visualizations help people discover trends, patterns, and anomalies in data,…

计算与语言 · 计算机科学 2024-10-01 Enamul Hoque , Mohammed Saidul Islam

AI models capable of comprehending humor hold real-world promise -- for example, enhancing engagement in human-machine interactions. To gauge and diagnose the capacity of multimodal large language models (MLLMs) for humor understanding, we…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Zhengpeng Shi , Yanpeng Zhao , Jianqun Zhou , Yuxuan Wang , Qinrong Cui , Wei Bi , Songchun Zhu , Bo Zhao , Zilong Zheng

We explore how the lens of fictional superpowers can help characterize how visualizations empower people and provide inspiration for new visualization systems. Researchers and practitioners often tout visualizations' ability to "make the…

Understanding how humans communicate and perceive narratives is important for media technology research and development. This is particularly important in current times when there are tools and algorithms that are easily available for…

人工智能 · 计算机科学 2023-12-15 Yi-Chun Chen , Arnav Jhala

Artificial Intelligence algorithms have now become pervasive in multiple high-stakes domains. However, their internal logic can be obscure to humans. Explainable Artificial Intelligence aims to design tools and techniques to illustrate the…

人机交互 · 计算机科学 2024-04-29 Eleonora Cappuccio , Daniele Fadda , Rosa Lanzilotti , Salvatore Rinzivillo

Text-rich visual understanding-the ability to process environments where dense textual content is integrated with visuals-is crucial for multimodal large language models (MLLMs) to interact effectively with structured environments. To…

计算机视觉与模式识别 · 计算机科学 2024-11-07 Junpeng Liu , Tianyue Ou , Yifan Song , Yuxiao Qu , Wai Lam , Chenyan Xiong , Wenhu Chen , Graham Neubig , Xiang Yue

Generative models have received a lot of attention in many areas of academia and the industry. Their capabilities span many areas, from the invention of images given a prompt to the generation of concrete code to solve a certain programming…

人机交互 · 计算机科学 2024-03-12 Pere-Pau Vázquez

While people with visual impairments are interested in artwork as much as their sighted peers, their experience is limited to few selective artworks that are exhibited at certain museums. To enable people with visual impairments to access…

人机交互 · 计算机科学 2021-03-03 Nahyun Kwon , Yunjung Lee , Uran Oh