English
Related papers

Related papers: AirScript - Creating Documents in Air

200 papers

We propose an unsupervised method for parsing large 3D scans of real-world scenes with easily-interpretable shapes. This work aims to provide a practical tool for analyzing 3D scenes in the context of aerial surveying and mapping, without…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Romain Loiseau , Elliot Vincent , Mathieu Aubry , Loic Landrieu

This work aims to generate realistic anatomical deformations from static patient scans. Specifically, we present a method to generate these deformations/augmentations via deep learning driven respiratory motion simulation that provides the…

Computer Vision and Pattern Recognition · Computer Science 2023-01-30 Donghoon Lee , Ellen Yorke , Masoud Zarepisheh , Saad Nadeem , Yu-Chi Hu

We present a framework for learning to describe fine-grained visual differences between instances using attribute phrases. Attribute phrases capture distinguishing aspects of an object (e.g., "propeller on the nose" or "door near the wing"…

Computer Vision and Pattern Recognition · Computer Science 2017-08-30 Jong-Chyi Su , Chenyun Wu , Huaizu Jiang , Subhransu Maji

Video action detection (VAD) is a formidable vision task that involves the localization and classification of actions within the spatial and temporal dimensions of a video clip. Among the myriad VAD architectures, two-stage VAD methods…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Seok Hwan Lee , Taein Son , Soo Won Seo , Jisong Kim , Jun Won Choi

OCR (Optical Character Recognition) is a technology that offers comprehensive alphanumeric recognition of handwritten and printed characters at electronic speed by merely scanning the document. Recently, the understanding of visual data has…

Computer Vision and Pattern Recognition · Computer Science 2023-07-12 Atman Mishra , A. Sharath Ram , Kavyashree C

We present a target-aware video diffusion model that generates videos from an input image, in which an actor interacts with a specified target while performing a desired action. The target is defined by a segmentation mask, and the action…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Taeksoo Kim , Hanbyul Joo

We introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vision jointly. We…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Aggelina Chatziagapi , Louis-Philippe Morency , Hongyu Gong , Michael Zollhoefer , Dimitris Samaras , Alexander Richard

Human action-reaction synthesis, a fundamental challenge in modeling causal human interactions, plays a critical role in applications ranging from virtual reality to social robotics. While diffusion-based models have demonstrated promising…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Wentao Jiang , Jingya Wang , Kaiyang Ji , Baoxiong Jia , Siyuan Huang , Ye Shi

In the past, computer vision systems for digitized documents could rely on systematically captured, high-quality scans. Today, transactions involving digital documents are more likely to start as mobile phone photo uploads taken by…

Computer Vision and Pattern Recognition · Computer Science 2021-12-20 Nikhil Maddikunta , Huijun Zhao , Sumit Keswani , Alfy Samuel , Fu-Ming Guo , Nishan Srishankar , Vishwa Pardeshi , Austin Huang

General-purpose foundation models have led to recent breakthroughs in artificial intelligence. In remote sensing, self-supervised learning (SSL) and Masked Image Modeling (MIM) have been adopted to build foundation models. However, these…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Fan Liu , Delong Chen , Zhangqingyun Guan , Xiaocong Zhou , Jiale Zhu , Qiaolin Ye , Liyong Fu , Jun Zhou

As generative AI becomes part of everyday writing, questions of transparency and productive human effort are increasingly important. Educators, reviewers, and readers want to understand how AI shaped the process. Where was human effort…

Human-Computer Interaction · Computer Science 2025-09-30 Momin N. Siddiqui , Nikki Nasseri , Adam Coscia , Roy Pea , Hari Subramonyam

In the rapidly evolving landscape of digital content creation, the demand for fast, convenient, and autonomous methods of crafting detailed 3D reconstructions of humans has grown significantly. Addressing this pressing need, our AirNeRF…

Robotics · Computer Science 2024-07-16 Alexey Kotcov , Maria Dronova , Vladislav Cheremnykh , Sausar Karaf , Dzmitry Tsetserukou

Facial animation is a core component for creating digital characters in Computer Graphics (CG) industry. A typical production workflow relies on sparse, semantically meaningful keyframes to precisely control facial expressions. Enabling…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Jingchao Wu , Zejian Kang , Haibo Liu , Yuanchen Fei , Xiangru Huang

With the development of large language models, their ability to follow simple instructions has significantly improved. However, adhering to complex instructions remains a major challenge. Current approaches to generating complex…

Computation and Language · Computer Science 2025-02-28 Wei Liu , Yancheng He , Hui Huang , Chengwei Hu , Jiaheng Liu , Shilong Li , Wenbo Su , Bo Zheng

Most people think that their handwriting is unique and cannot be imitated by machines, especially not using completely new content. Current cursive handwriting synthesis is visually limited or needs user interaction. We show that…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Martin Mayr , Martin Stumpf , Anguelos Nicolaou , Mathias Seuret , Andreas Maier , Vincent Christlein

Geovisualizations are powerful tools for exploratory spatial analysis, enabling sighted users to discern patterns, trends, and relationships within geographic data. However, these visual tools have remained largely inaccessible to…

Human-Computer Interaction · Computer Science 2024-12-11 Chu Li , Rock Yuren Pang , Ather Sharif , Arnavi Chheda-Kothary , Jeffrey Heer , Jon E. Froehlich

Semantic segmentation for uncrewed aerial vehicles (UAVs) is fundamental for aerial scene understanding, yet existing RGB and RGB-T datasets remain limited in scale, diversity, and annotation efficiency due to the high cost of manual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Markus Gross , Sai Bharadhwaj Matha , Rui Song , Viswanathan Muthuveerappan , Conrad Christoph , Julius Huber , Daniel Cremers

Association Rule Mining (ARM) is the task of mining patterns among data features in the form of logical rules, with applications across a myriad of domains. However, high-dimensional datasets often result in an excessive number of rules,…

Artificial Intelligence · Computer Science 2026-01-01 Erkan Karabulut , Paul Groth , Victoria Degeler

Generating human grasping poses that accurately reflect both object geometry and user-specified interaction semantics is essential for natural hand-object interactions in AR/VR and embodied AI. However, existing semantic grasping approaches…

Robotics · Computer Science 2026-03-31 Xiaofei Wu , Yi Zhang , Yumeng Liu , Yuexin Ma , Yujiao Shi , Xuming He

Accurate and automated captioning of aerial imagery is crucial for applications like environmental monitoring, urban planning, and disaster management. However, this task remains challenging due to complex spatial semantics and domain…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Xing Zi , Tengjun Ni , Xianjing Fan , Xian Tao , Jun Li , Ali Braytee , Mukesh Prasad