English
Related papers

Related papers: DescribePro: Collaborative Audio Description with …

200 papers

Large Language Models (LLMs) have assisted humans in several writing tasks, including text revision and story generation. However, their effectiveness in supporting domain-specific writing, particularly in business contexts, is relatively…

Human-Computer Interaction · Computer Science 2024-07-17 Minhwa Lee , Zae Myung Kim , Vivek Khetan , Dongyeop Kang

This study investigates innovative interaction designs for communication and collaborative learning between learners of mixed hearing and signing abilities, leveraging advancements in mixed reality technologies like Apple Vision Pro and…

Human-Computer Interaction · Computer Science 2025-03-11 Si Chen , Haocong Cheng , Suzy Su , Stephanie Patterson , Raja Kushalnagar , Qi Wang , Yun Huang

We propose CoLyricist, an AI-assisted lyric writing tool designed to support the typical workflows of experienced lyricists and enhance their creative efficiency. While lyricists have unique processes, many follow common stages. Tools that…

Human-Computer Interaction · Computer Science 2026-02-27 Masahiro Yoshida , Bingxuan Li , Songyan Zhao , Qinyi Zhou , Shiwei Hu , Xiang Anthony Chen , Nanyun Peng

The advancement of Machine learning (ML), Large Audio Language Models (LALMs), and autonomous AI agents in Music Information Retrieval (MIR) necessitates a shift from static tagging to rich, human-aligned representation learning. However,…

Audio description (AD) is a crucial accessibility service provided to blind persons and persons with visual impairment, designed to convey visual information in acoustic form. Despite recent advancements in multilingual machine translation…

Computation and Language · Computer Science 2024-11-25 Lukas Fischer , Yingqiang Gao , Alexa Lintner , Sarah Ebling

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality of human…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Chaoyu Li , Sid Padmanabhuni , Maryam Cheema , Hasti Seifi , Pooyan Fazli

The growing popularity of AI writing assistants presents exciting opportunities to craft tools that cater to diverse user needs. This study explores how personality shapes preferences for AI writing companions and how personalized designs…

Human-Computer Interaction · Computer Science 2025-09-16 Mengke Wu , Kexin Quan , Weizi Liu , Mike Yao , Jessie Chin

Automated live visual descriptions can aid blind people in understanding their surroundings with autonomy and independence. However, providing descriptions that are rich, contextual, and just-in-time has been a long-standing challenge in…

Human-Computer Interaction · Computer Science 2024-08-14 Ruei-Che Chang , Yuxuan Liu , Anhong Guo

While supervised learning has achieved significant success in computer vision tasks, acquiring high-quality annotated data remains a bottleneck. This paper explores both scholarly and non-scholarly works in AI-assistive deep learning image…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Moseli Mots'oehli

A big part of achieving Artificial General Intelligence(AGI) is to build a machine that can see and listen like humans. Much work has focused on designing models for image classification, video classification, object detection, pose…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Ruotian Luo

Chatbots have long been explored as tools to support learning, and recent advances in large language models have significantly expanded the availability of platforms for educators to author AI tutoring chatbots. Yet effective authorship…

Human-Computer Interaction · Computer Science 2026-05-19 Miina Koyama , Ruiwei Xiao , John Stamper

This study introduces a novel training paradigm, audio difference learning, for improving audio captioning. The fundamental concept of the proposed learning method is to create a feature representation space that preserves the relationship…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Tatsuya Komatsu , Yusuke Fujita , Kazuya Takeda , Tomoki Toda

Large language models can be used for collaborative storytelling. In this work we report on using GPT-3 \cite{brown2020language} to co-narrate stories. The AI system must track plot progression and character arcs while the human actors…

Human-Computer Interaction · Computer Science 2021-10-01 Boyd Branch , Piotr Mirowski , Kory W. Mathewson

To make an engaging video, people sequence interesting moments and add visuals such as B-rolls or text. While video editing requires time and effort, AI has recently shown strong potential to make editing easier through suggestions and…

Human-Computer Interaction · Computer Science 2025-02-17 Mina Huh , Dingzeyu Li , Kim Pimmel , Hijung Valentina Shin , Amy Pavel , Mira Dontcheva

Audio Descriptions (AD) are essential for making visual content accessible to individuals with visual impairments. Recent works have shown a promising step towards automating AD, but they have been limited to describing high-quality movie…

Multimedia · Computer Science 2025-11-13 Lipisha Chaudhary , Trisha Mittal , Subhadra Gopalakrishnan , Ifeoma Nwogu , Jaclyn Pytlarz

Storytelling plays a central role in human socializing and entertainment. However, much of the research on automatic storytelling generation assumes that stories will be generated by an agent without any human interaction. In this paper, we…

Computation and Language · Computer Science 2020-11-23 Eric Nichols , Leo Gao , Randy Gomez

While audio description (AD) is the standard approach for making videos accessible to blind and low vision (BLV) people, existing AD guidelines do not consider BLV users' varied preferences across viewing scenarios. These scenarios range…

Human-Computer Interaction · Computer Science 2024-03-19 Lucy Jiang , Crescentia Jung , Mahika Phutane , Abigale Stangl , Shiri Azenkot

Large language model-based agents are rapidly evolving from simple conversational assistants into autonomous systems capable of performing complex, professional-level tasks in various domains. While these advancements promise significant…

Labeling neural network submodules with human-legible descriptions is useful for many downstream tasks: such descriptions can surface failures, guide interventions, and perhaps even explain important model behaviors. To date, most…

Computation and Language · Computer Science 2023-12-11 Sarah Schwettmann , Tamar Rott Shaham , Joanna Materzynska , Neil Chowdhury , Shuang Li , Jacob Andreas , David Bau , Antonio Torralba

Creators struggle to edit long-form, narrative-rich videos not because of UI complexity, but due to the cognitive demands of searching, storyboarding, and sequencing hours of footage. Existing transcript- or embedding-based methods fall…

Artificial Intelligence · Computer Science 2025-09-30 Zihan Ding , Xinyi Wang , Junlong Chen , Per Ola Kristensson , Junxiao Shen