中文
相关论文

相关论文: AutoPoster: A Highly Automatic and Content-aware D…

200 篇论文

Analyzing the layout of a document to identify headers, sections, tables, figures etc. is critical to understanding its content. Deep learning based approaches for detecting the layout structure of document images have been promising.…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Natraj Raman , Sameena Shah , Manuela Veloso

Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these generated outputs presents ongoing challenges. Although numerous…

计算与语言 · 计算机科学 2025-06-13 Tian Lan , Yang-Hao Zhou , Zi-Ao Ma , Fanshu Sun , Rui-Qing Sun , Junyu Luo , Rong-Cheng Tu , Heyan Huang , Chen Xu , Zhijing Wu , Xian-Ling Mao

We present ShapeCrafter, a neural network for recursive text-conditioned 3D shape generation. Existing methods to generate text-conditioned 3D shapes consume an entire text prompt to generate a 3D shape in a single step. However, humans…

计算机视觉与模式识别 · 计算机科学 2023-04-11 Rao Fu , Xiao Zhan , Yiwen Chen , Daniel Ritchie , Srinath Sridhar

We explore computational approaches for visual guidance to aid in creating aesthetically pleasing art and graphic design. Our work complements and builds on previous work that developed models for how humans look at images. Our approach…

计算机视觉与模式识别 · 计算机科学 2021-07-14 Qingyuan Zheng , Zhuoru Li , Adam Bargteil

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion.The applications of listener agent generation in virtual interaction…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Xi Liu , Ying Guo , Cheng Zhen , Tong Li , Yingying Ao , Pengfei Yan

In online advertising scenario, sellers often create multiple creatives to provide comprehensive demonstrations, making it essential to present the most appealing design to maximize the Click-Through Rate (CTR). However, sellers generally…

信息检索 · 计算机科学 2024-01-23 Hao Yang , Jianxin Yuan , Shuai Yang , Linhe Xu , Shuo Yuan , Yifan Zeng

Web scraping is a powerful technique that extracts data from websites, enabling automated data collection, enhancing data analysis capabilities, and minimizing manual data entry efforts. Existing methods, wrappers-based methods suffer from…

计算与语言 · 计算机科学 2024-09-27 Wenhao Huang , Zhouhong Gu , Chenghao Peng , Zhixu Li , Jiaqing Liang , Yanghua Xiao , Liqian Wen , Zulong Chen

Small business owners (SBOs) often lack the resources and design experience needed to produce high-quality advertisements. To address this, we developed ACAI (AI Co-Creation for Advertising and Inspiration), an GenAI-powered multimodal…

人机交互 · 计算机科学 2025-03-11 Nimisha Karnatak , Adrien Baranes , Rob Marchant , Triona Butler , Kristen Olson

Self-driving vehicles are the future of transportation. With current advancements in this field, the world is getting closer to safe roads with almost zero probability of having accidents and eliminating human errors. However, there is…

计算机视觉与模式识别 · 计算机科学 2021-11-22 Kareem Metwaly , Aerin Kim , Elliot Branson , Vishal Monga

Graphic design is an effective language for visual communication. Using complex composition of visual elements (e.g., shape, color, font) guided by design principles and aesthetics, design helps produce more visually-appealing content. The…

人机交互 · 计算机科学 2023-09-06 Danqing Huang , Jiaqi Guo , Shizhao Sun , Hanling Tian , Jieru Lin , Zheng Hu , Chin-Yew Lin , Jian-Guang Lou , Dongmei Zhang

The increasing adoption of neural networks in learning-augmented systems highlights the importance of model safety and robustness, particularly in safety-critical domains. Despite progress in the formal verification of neural networks,…

机器学习 · 计算机科学 2024-10-25 Shuowei Jin , Francis Y. Yan , Cheng Tan , Anuj Kalia , Xenofon Foukas , Z. Morley Mao

Recent generative models such as GPT-4o have shown strong capabilities in producing high-quality images with accurate text rendering. However, commercial design tasks like advertising banners demand more than visual fidelity -- they require…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Zhao Wang , Bowen Chen , Yotaro Shimose , Sota Moriyama , Heng Wang , Shingo Takamatsu

Subject-driven text-to-image generation still struggles to preserve high-frequency identity details such as logos, patterns, and text. Existing methods typically operate directly in RGB space, which often leads to detail degradation under…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Hanzhong Guo , Yizhou Yu

Graphic design forms the cornerstone of modern visual communication, serving as a vital medium for promoting cultural and commercial events. Recent advances have explored automating this process using Large Multimodal Models (LMMs), yet…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Jiazhe Wei , Ken Li , Tianyu Lao , Haofan Wang , Liang Wang , Caifeng Shan , Chenyang Si

Artistic text generation aims to amplify the aesthetic qualities of text while maintaining readability. It can make the text more attractive and better convey its expression, thus enjoying a wide range of application scenarios such as…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Yuhang Bai , Zichuan Huang , Wenshuo Gao , Shuai Yang , Jiaying Liu

Layout generation is a novel task in computer vision, which combines the challenges in both object localization and aesthetic appraisal, widely used in advertisements, posters, and slides design. An accurate and pleasant layout should…

计算机视觉与模式识别 · 计算机科学 2022-09-05 Yunning Cao , Ye Ma , Min Zhou , Chuanbin Liu , Hongtao Xie , Tiezheng Ge , Yuning Jiang

Generative models, such as large language models and text-to-image diffusion models, are increasingly used to create visual designs like user interfaces (UIs) and presentation slides. Finetuning and benchmarking these generative models have…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Yi-Hao Peng , Jeffrey P. Bigham , Jason Wu

As the volume of peer-reviewed research surges, scholars increasingly rely on social platforms for discovery, while authors invest considerable effort in promoting their work to ensure visibility and citations. To streamline this process…

State-of-the-art approaches in computer vision heavily rely on sufficiently large training datasets. For real-world applications, obtaining such a dataset is usually a tedious task. In this paper, we present a fully automated pipeline to…

计算机视觉与模式识别 · 计算机科学 2023-04-19 Alexander Naumann , Felix Hertlein , Benchun Zhou , Laura Dörr , Kai Furmans

Automatic 3D content creation seeks to replace labor-intensive modeling and scanning pipelines with systems that can synthesize or recover 3D assets directly from text or images. Its applications span video games, virtual reality, robotics,…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Jiahao Li