中文
相关论文

相关论文: LICA: Layered Image Composition Annotations for Gr…

200 篇论文

Instance shape reconstruction from a 3D scene involves recovering the full geometries of multiple objects at the semantic instance level. Many methods leverage data-driven learning due to the intricacies of scene complexity and significant…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Haolin Liu , Chongjie Ye , Yinyu Nie , Yingfan He , Xiaoguang Han

In this paper, we study the graphic layout generation problem of producing high-quality visual-textual presentation designs for given images. We note that image compositions, which contain not only global semantics but also spatial…

计算机视觉与模式识别 · 计算机科学 2022-07-14 Min Zhou , Chenchen Xu , Ye Ma , Tiezheng Ge , Yuning Jiang , Weiwei Xu

The digitization of documents allows for wider accessibility and reproducibility. While automatic digitization of document layout and text content has been a long-standing focus of research, this problem in regard to graphical elements,…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Omar Moured , Jiaming Zhang , Alina Roitberg , Thorsten Schwarz , Rainer Stiefelhagen

Infographic charts are a powerful medium for communicating abstract data by combining visual elements (e.g., charts, images) with textual information. However, their visual and structural richness poses challenges for large vision-language…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Zhen Li , Duan Li , Yukai Guo , Xinyuan Guo , Bowen Li , Lanxi Xiao , Shenyu Qiao , Jiashu Chen , Zijian Wu , Hui Zhang , Xinhuan Shu , Shixia Liu

We introduce FigureQA, a visual reasoning corpus of over one million question-answer pairs grounded in over 100,000 images. The images are synthetic, scientific-style figures from five classes: line plots, dot-line plots, vertical and…

计算机视觉与模式识别 · 计算机科学 2018-02-26 Samira Ebrahimi Kahou , Vincent Michalski , Adam Atkinson , Akos Kadar , Adam Trischler , Yoshua Bengio

Improving the accessibility and automation capabilities of mobile devices can have a significant positive impact on the daily lives of countless users. To stimulate research in this direction, we release a human-annotated dataset with…

人机交互 · 计算机科学 2022-10-07 Srinivas Sunkara , Maria Wang , Lijuan Liu , Gilles Baechler , Yu-Chung Hsiao , Jindong , Chen , Abhanshu Sharma , James Stout

Content-aware visual-textual presentation layout aims at arranging spatial space on the given canvas for pre-defined elements, including text, logo, and underlay, which is a key to automatic template-free creative graphic design. In…

计算机视觉与模式识别 · 计算机科学 2023-03-29 HsiaoYuan Hsu , Xiangteng He , Yuxin Peng , Hao Kong , Qing Zhang

We present Open Images V4, a dataset of 9.2M images with unified annotations for image classification, object detection and visual relationship detection. The images have a Creative Commons Attribution license that allows to share and adapt…

We present PosterIQ, a design-driven benchmark for poster understanding and generation, annotated across composition structure, typographic hierarchy, and semantic intent. It includes 7,765 image-annotation instances and 822 generation…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Yuheng Feng , Wen Zhang , Haodong Duan , Xingxing Zou

In controllable image synthesis, generating coherent and consistent images from multiple references with spatial layout awareness remains an open challenge. We present LAMIC, a Layout-Aware Multi-Image Composition framework that, for the…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Yuzhuo Chen , Zehua Ma , Jianhua Wang , Kai Kang , Shunyu Yao , Weiming Zhang

Questing for learned lossy image coding (LIC) with superior compression performance and computation throughput is challenging. The vital factor behind it is how to intelligently explore Adaptive Neighborhood Information Aggregation (ANIA)…

图像与视频处理 · 电气工程与系统科学 2022-10-13 Ming Lu , Fangdong Chen , Shiliang Pu , Zhan Ma

Graphical Abstracts (GAs) play a crucial role in visually conveying the key findings of scientific papers. Although recent research increasingly incorporates visual materials such as Figure 1 as de facto GAs, their potential to enhance…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Takuro Kawada , Shunsuke Kitada , Sota Nemoto , Hitoshi Iyatomi

Conventional document layout analysis (DLA) traditionally depends on empirical priors or a fixed set of learnable queries executed in a single forward pass. While sufficient for early-generation documents with a small, predetermined number…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Yufan Chen , Omar Moured , Ruiping Liu , Junwei Zheng , Kunyu Peng , Jiaming Zhang , Rainer Stiefelhagen

The availability of labeled image datasets has been shown critical for high-level image understanding, which continuously drives the progress of feature designing and models developing. However, constructing labeled image datasets is…

计算机视觉与模式识别 · 计算机科学 2019-03-04 Yazhou Yao , Jian Zhang , Fumin Shen , Li Liu , Fan Zhu , Dongxiang Zhang , Heng-Tao Shen

Lens visualization has been a prominent research area in the visualization community, fueled by the continuous need to mitigate visual clutter and occlusion resulting from the increase in data volume. Interactive lenses for spatial data,…

图形学 · 计算机科学 2025-06-09 Roberta Mota , Ehud Sharlin , Usman Alim

Composite visualization is a popular design strategy that represents complex datasets by integrating multiple visualizations in a meaningful and aesthetic layout, such as juxtaposition, overlay, and nesting. With this strategy, numerous…

人机交互 · 计算机科学 2022-11-04 Dazhen Deng , Weiwei Cui , Xiyu Meng , Mengye Xu , Yu Liao , Haidong Zhang , Yingcai Wu

Typical layout-to-image synthesis (LIS) models generate images for a closed set of semantic classes, e.g., 182 common objects in COCO-Stuff. In this work, we explore the freestyle capability of the model, i.e., how far can it generate…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Han Xue , Zhiwu Huang , Qianru Sun , Li Song , Wenjun Zhang

A picture is worth a thousand words. Albeit a clich\'e, for the fashion industry, an image of a clothing piece allows one to perceive its category (e.g., dress), sub-category (e.g., day dress) and properties (e.g., white colour with floral…

计算机视觉与模式识别 · 计算机科学 2018-06-26 Beatriz Quintino Ferreira , Luís Baía , João Faria , Ricardo Gamelas Sousa

Object compositing, the task of placing and harmonizing objects in images of diverse visual scenes, has become an important task in computer vision with the rise of generative models. However, existing datasets lack the diversity and scale…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Jinwoo Kim , Sangmin Han , Jinho Jeong , Jiwoo Choi , Dongyoung Kim , Seon Joo Kim

Composed Image Retrieval (CIR) allows users to search for images by combining a reference image with a text prompt that describes desired modifications. While vision-language models like CLIP have popularized this task by embedding multiple…

‹ 上一页 1 2 3 10 下一页 ›