English
Related papers

Related papers: DocuBits: VR Document Decomposition for Procedural…

200 papers

We propose DocFormerv2, a multi-modal transformer for Visual Document Understanding (VDU). The VDU domain entails understanding documents (beyond mere OCR predictions) e.g., extracting information from a form, VQA for documents and other…

Computer Vision and Pattern Recognition · Computer Science 2023-06-05 Srikar Appalaraju , Peng Tang , Qi Dong , Nishant Sankaran , Yichu Zhou , R. Manmatha

In this contribution, we design, implement and evaluate the pedagogical benefits of a novel interactive note taking interface (iVRNote) in VR for the purpose of learning and reflection lectures. In future VR learning environments, students…

Human-Computer Interaction · Computer Science 2019-10-04 Yi-Ting Chen , Chi-Hsuan Hsu , Chih-Han Chung , Yu-Shuen Wang , Sabarish V. Babu

This report introduces PP-DocBee2, an advanced version of the PP-DocBee, designed to enhance multimodal document understanding. Built on a large multimodal model architecture, PP-DocBee2 addresses the limitations of its predecessor through…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Kui Huang , Xinrong Chen , Wenyu Lv , Jincheng Liao , Guanzhong Wang , Yi Liu

The use of visually-rich documents (VRDs) in various fields has created a demand for Document AI models that can read and comprehend documents like humans, which requires the overcoming of technical, linguistic, and cognitive barriers.…

Human-Computer Interaction · Computer Science 2023-10-24 Hao Wang , Qingxuan Wang , Yue Li , Changqing Wang , Chenhui Chu , Rui Wang

The visual dialog task attempts to train an agent to answer multi-turn questions given an image, which requires the deep understanding of interactions between the image and dialog history. Existing researches tend to employ the…

Computation and Language · Computer Science 2022-02-23 Tong Ye , Shijing Si , Jianzong Wang , Rui Wang , Ning Cheng , Jing Xiao

Document understanding tasks, in particular, Visually-rich Document Entity Retrieval (VDER), have gained significant attention in recent years thanks to their broad applications in enterprise AI. However, publicly available data have been…

Computation and Language · Computer Science 2023-10-27 Lijun Yu , Jin Miao , Xiaoyu Sun , Jiayi Chen , Alexander G. Hauptmann , Hanjun Dai , Wei Wei

The application and implementation of collaborative embodiment in virtual reality (VR) are a critical aspect of the computer science landscape, aiming to enhance multi-user interaction and teamwork in immersive environments. A notable and…

Human-Computer Interaction · Computer Science 2025-07-28 Hongyu Zhou , Yihao Dong , Masahiko Inami , Zhanna Sarsenbayeva , Anusha Withana

Recent advances in Visually-rich Document Understanding rely on large Vision-Language Models like Donut, which perform document-level Visual Question Answering without Optical Character Recognition. Despite their effectiveness, these models…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Adnan Ben Mansour , Ayoub Karine , David Naccache

Effective data visualization is a key part of the discovery process in the era of big data. It is the bridge between the quantitative content of the data and human intuition, and thus an essential component of the scientific path from data…

Deep learning is ubiquitous, but its lack of transparency limits its impact on several potential application areas. We demonstrate a virtual reality tool for automating the process of assigning data inputs to different categories. A dataset…

Human-Computer Interaction · Computer Science 2023-05-25 Hannes Kath , Bengt Lüers , Thiago S. Gouvêa , Daniel Sonntag

Vision Transformers (ViT) have emerged as the de-facto choice for numerous industry grade vision solutions. But their inference cost can be prohibitive for many settings, as they compute self-attention in each layer which suffers from…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Rajat Koner , Gagan Jain , Prateek Jain , Volker Tresp , Sujoy Paul

We present an approach to evaluate the efficacy of annotations in augmenting learning environments in the context of Virtual Reality. Our study extends previous work highlighting the benefits of learning based in virtual reality and…

Human-Computer Interaction · Computer Science 2025-02-24 Maximilian Enderling , Jan Hombeck , Kai Lawonn

Users often take notes for instructional videos to access key knowledge later without revisiting long videos. Automated note generation tools enable users to obtain informative notes efficiently. However, notes generated by existing…

Human-Computer Interaction · Computer Science 2025-08-21 Running Zhao , Zhihan Jiang , Xinchen Zhang , Chirui Chang , Handi Chen , Weipeng Deng , Luyao Jin , Xiaojuan Qi , Xun Qian , Edith C. H. Ngai

Document Visual Question Answering (DocVQA) is a practical yet challenging task, which is to ask questions based on documents while referring to multiple pages and different modalities of information, e.g, images and tables. To handle…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Chelsi Jain , Yiran Wu , Yifan Zeng , Jiale Liu , S hengyu Dai , Zhenwen Shao , Qingyun Wu , Huazheng Wang

Virtual Reality (VR) offers a unique collaborative experience, with parallel views playing a pivotal role in Collaborative Virtual Environments by supporting the transfer and delivery of items. Sharing and manipulating partners' views…

Human-Computer Interaction · Computer Science 2025-03-10 Xian Wang , Luyao Shen , Lei Chen , Mingming Fan , Lik-Hang Lee

Recent progress has shown great potential of visual prompt tuning (VPT) when adapting pre-trained vision transformers to various downstream tasks. However, most existing solutions independently optimize prompts at each layer, thereby…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Nan Zhou , Jiaxin Chen , Di Huang

We present a systematic review on tasks, interactions, and visualization widgets (refer to tangible entities that are used to accomplish data exploration tasks through specific interactions) in the context of tangible data exploration.…

Human-Computer Interaction · Computer Science 2025-07-02 Haonan Yao , Lingyun Yu , Lijie Yao

The advent of multimodal learning has brought a significant improvement in document AI. Documents are now treated as multimodal entities, incorporating both textual and visual information for downstream analysis. However, works in this…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Nikitha SR , Tarun Ram Menta , Mausoom Sarkar

Despite the value of VR (Virtual Reality) for educational purposes, the instructional power of VR in Biology Laboratory education remains under-explored. Laboratory lectures can be challenging due to students' low motivation to learn…

Human-Computer Interaction · Computer Science 2023-04-20 Fei Xue , Rongchen Guo , Siyuan Yao , Luxin Wang , Kwan-Liu Ma

MimicKit is an open-source framework for training motion controllers using motion imitation and reinforcement learning. The codebase provides implementations of commonly-used motion-imitation techniques and RL algorithms. This framework is…

Graphics · Computer Science 2026-01-21 Xue Bin Peng