中文
相关论文

相关论文: CalliMaster: Mastering Page-level Chinese Calligra…

200 篇论文

Previous works indicate that the glyph of Chinese characters contains rich semantic information and has the potential to enhance the representation of Chinese characters. The typical method to utilize the glyph features is by incorporating…

人工智能 · 计算机科学 2021-07-02 Yunxin Li , Yu Zhao , Baotian Hu , Qingcai Chen , Yang Xiang , Xiaolong Wang , Yuxin Ding , Lin Ma

We propose a novel hierarchical approach for text-to-image synthesis by inferring semantic layout. Instead of learning a direct mapping from text to image, our algorithm decomposes the generation process into multiple steps, in which it…

计算机视觉与模式识别 · 计算机科学 2018-07-27 Seunghoon Hong , Dingdong Yang , Jongwook Choi , Honglak Lee

This paper introduces the WordArt Designer API, a novel framework for user-driven artistic typography synthesis utilizing Large Language Models (LLMs) on ModelScope. We address the challenge of simplifying artistic typography for…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Jun-Yan He , Zhi-Qi Cheng , Chenyang Li , Jingdong Sun , Wangmeng Xiang , Yusen Hu , Xianhui Lin , Xiaoyang Kang , Zengke Jin , Bin Luo , Yifeng Geng , Xuansong Xie , Jingren Zhou

Painterly image harmonization aims to harmonize a photographic foreground object on the painterly background. Different from previous auto-encoder based harmonization networks, we develop a progressive multi-stage harmonization network,…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Li Niu , Yan Hong , Junyan Cao , Liqing Zhang

We present Zero-Painter, a novel training-free framework for layout-conditional text-to-image synthesis that facilitates the creation of detailed and controlled imagery from textual prompts. Our method utilizes object masks and individual…

计算机视觉与模式识别 · 计算机科学 2024-06-07 Marianna Ohanyan , Hayk Manukyan , Zhangyang Wang , Shant Navasardyan , Humphrey Shi

Architectural floor plan design demands joint reasoning over geometry, semantics, and spatial hierarchy, which remains a major challenge for current AI systems. Although recent diffusion and language models improve visual fidelity, they…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Sizhong Qin , Ramon Elias Weber , Xinzheng Lu

Balancing scientific exposition and narrative engagement is a central challenge in science communication. To examine how to achieve balance, we conducted a formative study with four science communicators and a literature review of science…

人机交互 · 计算机科学 2026-05-19 Kexue Fu , Jiaye Leng , Yawen Zhang , Jingfei Huang , Yihang Zuo , Runze Cai , Zijian Ding , Ray LC , Shengdong Zhao , Qinyuan Lei

Chinese characters have a complex and hierarchical graphical structure carrying both semantic and phonetic information. We use this structure to enhance the text model and obtain better results in standard NLP operations. First of all, to…

计算与语言 · 计算机科学 2014-05-22 Yannis Haralambous

Visual tokenizers play a central role in latent image generation by bridging high-dimensional images and tractable generative modeling. However, most existing tokenizers are still trained with reconstruction-dominated objectives, which…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Qingfeng Li , Haoxian Zhang , Xu He , Songlin Tang , Zhixue Fang , Xiaoqiang Liu , Pengfei Wan Guoqi Li

In manipulation tasks, a robot interacts with movable object(s). The configuration space in manipulation planning is thus the Cartesian product of the configuration space of the robot with those of the movable objects. It is the complex…

机器人学 · 计算机科学 2015-12-18 Puttichai Lertkultanon , Quang-Cuong Pham

Generating realistic building layouts for automatic building design has been studied in both the computer vision and architecture domains. Traditional approaches from the architecture domain, which are based on optimization techniques or…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Jiachen Liu , Yuan Xue , Haomiao Ni , Rui Yu , Zihan Zhou , Sharon X. Huang

It is common in graphic design humans visually arrange various elements according to their design intent and semantics. For example, a title text almost always appears on top of other elements in a document. In this work, we generate…

计算机视觉与模式识别 · 计算机科学 2021-08-03 Kotaro Kikuchi , Edgar Simo-Serra , Mayu Otani , Kota Yamaguchi

Text-guided diffusion models have greatly advanced image editing and generation. However, achieving physically consistent image retouching with precise parameter control (e.g., exposure, white balance, zoom) remains challenging. Existing…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Qirui Yang , Yang Yang , Ying Zeng , Xiaobin Hu , Bo Li , Huanjing Yue , Jingyu Yang , Peng-Tao Jiang

In this paper, we propose an end-to-end trainable framework for restoring historical documents content that follows the correct reading order. In this framework, two branches named character branch and layout branch are added behind the…

计算机视觉与模式识别 · 计算机科学 2020-07-15 Weihong Ma , Hesuo Zhang , Lianwen Jin , Sihang Wu , Jiapeng Wang , Yongpan Wang

Recently, leveraging large language models (LLMs) or multimodal large language models (MLLMs) for document understanding has been proven very promising. However, previous works that employ LLMs/MLLMs for document understanding have not…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Chuwei Luo , Yufan Shen , Zhaoqing Zhu , Qi Zheng , Zhi Yu , Cong Yao

We devise a 3D scene graph representation, contact graph+ (cg+), for efficient sequential task planning. Augmented with predicate-like attributes, this contact graph-based representation abstracts scene layouts with succinct geometric…

机器人学 · 计算机科学 2022-07-19 Ziyuan Jiao , Yida Niu , Zeyu Zhang , Song-Chun Zhu , Yixin Zhu , Hangxin Liu

Previous text-to-image synthesis algorithms typically use explicit textual instructions to generate/manipulate images accurately, but they have difficulty adapting to guidance in the form of coarsely matched texts. In this work, we attempt…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Mengyao Cui , Zhe Zhu , Shao-Ping Lu , Yulu Yang

Layout is important for graphic design and scene generation. We propose a novel Generative Adversarial Network, called LayoutGAN, that synthesizes layouts by modeling geometric relations of different types of 2D elements. The generator of…

计算机视觉与模式识别 · 计算机科学 2019-01-23 Jianan Li , Jimei Yang , Aaron Hertzmann , Jianming Zhang , Tingfa Xu

3D indoor layout synthesis is crucial for creating virtual environments. Traditional methods struggle with generalization due to fixed datasets. While recent LLM and VLM-based approaches offer improved semantic richness, they often lack…

机器人学 · 计算机科学 2025-10-03 Jialin Gao , Donghao Zhou , Mingjian Liang , Lihao Liu , Chi-Wing Fu , Xiaowei Hu , Pheng-Ann Heng

Text alignment finds application in tasks such as citation recommendation and plagiarism detection. Existing alignment methods operate at a single, predefined level and cannot learn to align texts at, for example, sentence and document…

计算与语言 · 计算机科学 2020-10-06 Xuhui Zhou , Nikolaos Pappas , Noah A. Smith