English
Related papers

Related papers: TutoAI: A Cross-domain Framework for AI-assisted M…

200 papers

While users tend to perceive instructional videos as an experience rather than a lesson with a set of instructions, instructional videos are more effective and appealing than textual user manuals and eliminate the ambiguity in text-based…

Human-Computer Interaction · Computer Science 2023-11-22 Songsong Liu , Shu Wang , Kun Sun

Related tasks often have inter-dependence on each other and perform better when solved in a joint framework. In this paper, we present a deep multi-task learning framework that jointly performs sentiment and emotion analysis both. The…

Computation and Language · Computer Science 2019-05-16 Md Shad Akhtar , Dushyant Singh Chauhan , Deepanway Ghosal , Soujanya Poria , Asif Ekbal , Pushpak Bhattacharyya

The automatic summarization of surgical videos is essential for enhancing procedural documentation, supporting surgical training, and facilitating post-operative analysis. This paper presents a novel method at the intersection of artificial…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Hugo Georgenthum , Cristian Cosentino , Fabrizio Marozzo , Pietro Liò

Image-to-image translation is a general name for a task where an image from one domain is converted to a corresponding image in another domain, given sufficient training data. Traditionally different approaches have been proposed depending…

Computer Vision and Pattern Recognition · Computer Science 2018-05-09 Soumya Tripathy , Juho Kannala , Esa Rahtu

Text-editable and pose-controllable character video generation is a challenging but prevailing topic with practical applications. However, existing approaches mainly focus on single-object video generation with pose guidance, ignoring the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Beiyuan Zhang , Yue Ma , Chunlei Fu , Xinyang Song , Zhenan Sun , Ziqiang Li

Education materials for K-12 students often consist of multiple modalities, such as text and images, posing challenges for models to fully understand nuanced information in these materials. In this paper, we propose a unified language and…

Computation and Language · Computer Science 2025-10-10 Zhendong Chu , Jian Xie , Shen Wang , Zichao Wang , Qingsong Wen

Artificial Intelligence (AI) has made incredible progress recently. On the one hand, advanced foundation models like ChatGPT can offer powerful conversation, in-context learning and code generation abilities on a broad range of open-domain…

Artificial Intelligence · Computer Science 2023-03-30 Yaobo Liang , Chenfei Wu , Ting Song , Wenshan Wu , Yan Xia , Yu Liu , Yang Ou , Shuai Lu , Lei Ji , Shaoguang Mao , Yun Wang , Linjun Shou , Ming Gong , Nan Duan

Text-to-image (T2I) diffusion models have revolutionized visual content creation, but extending these capabilities to text-to-video (T2V) generation remains a challenge, particularly in preserving temporal consistency. Existing methods that…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Dohun Lee , Bryan S Kim , Geon Yeong Park , Jong Chul Ye

In mixed-initiative systems, the mode of AI assistance delivery can be as consequential as the assistance itself. We investigated two assistance delivery modes: on-demand help (users request via Button) and pre-scheduled help (assistance…

Human-Computer Interaction · Computer Science 2026-02-03 Yunhao Luo , Arthur Caetano , Avinash Ajit Nargund , Tobias Höllerer , Misha Sra

Understanding the structure of complex activities in untrimmed videos is a challenging task in the area of action recognition. One problem here is that this task usually requires a large amount of hand-annotated minute- or even hour-long…

Computer Vision and Pattern Recognition · Computer Science 2020-10-01 Rosaura G. VidalMata , Walter J. Scheirer , Anna Kukleva , David Cox , Hilde Kuehne

Scalable AI tutoring for procedural skill learning requires structured knowledge representations, yet constructing these representations remains a labor-intensive bottleneck. This paper introduces a new LLM-assisted text-to-model (TTM)…

Human-Computer Interaction · Computer Science 2026-05-05 Rahul K. Dass , Shubham Puri , Arpit Khandelwal , Xiao Jin , Ashok K. Goel

Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience. Video-to-Audio (V2A), as a particular type of automatic foley task, presents…

Sound · Computer Science 2024-09-12 Qi Yang , Binjie Mao , Zili Wang , Xing Nie , Pengfei Gao , Ying Guo , Cheng Zhen , Pengfei Yan , Shiming Xiang

Text-rich visual understanding-the ability to process environments where dense textual content is integrated with visuals-is crucial for multimodal large language models (MLLMs) to interact effectively with structured environments. To…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Junpeng Liu , Tianyue Ou , Yifan Song , Yuxiao Qu , Wai Lam , Chenyan Xiong , Wenhu Chen , Graham Neubig , Xiang Yue

Although large-scale video-language pre-training models, which usually build a global alignment between the video and the text, have achieved remarkable progress on various downstream tasks, the idea of adopting fine-grained information…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Weihong Zhong , Mao Zheng , Duyu Tang , Xuan Luo , Heng Gong , Xiaocheng Feng , Bing Qin

Many everyday tasks, ranging from appliance repair and cooking to car maintenance, require expert knowledge, particularly for complex, multi-step procedures. Despite growing interest in AI agents for augmented reality (AR) assistance,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Lavisha Aggarwal , Vikas Bahirwani , Andrea Colaco

The improved competence of generative models can help building multi-modal virtual assistants that leverage modalities beyond language. By observing humans performing multi-step tasks, one can build assistants that have situational…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Pha Nguyen , Sailik Sengupta , Girik Malik , Arshit Gupta , Bonan Min

Creation of images using generative adversarial networks has been widely adapted into multi-modal regime with the advent of multi-modal representation models pre-trained on large corpus. Various modalities sharing a common representation…

Sound · Computer Science 2022-06-10 Yoonjeon Kim , Joel Jang , Sumin Shin

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of existing interleaved…

Developers prefer to utilize third-party libraries when they implement some functionalities and Application Programming Interfaces (APIs) are frequently used by them. Facing an unfamiliar API, developers tend to consult tutorials as…

Software Engineering · Computer Science 2017-03-07 He Jiang , Jingxuan Zhang , Xiaochen Li , Zhilei Ren , David Lo

Cartoon videos have proven to be effective in learning vocabulary to preschool children.However, we have little knowledge about integrating AI into cartoon videos to provide systematic, multimodal vocabulary learning support. This…

Human-Computer Interaction · Computer Science 2025-02-19 Shiya Tsang , Ruiyao Miao , Junren Xiao , Hui Xiong