中文
相关论文

相关论文: Character-Adapter: Prompt-Guided Region Control fo…

200 篇论文

Within Convolutional Neural Network (CNN), the convolution operations are good at extracting local features but experience difficulty to capture global representations. Within visual transformer, the cascaded self-attention modules can…

计算机视觉与模式识别 · 计算机科学 2021-05-11 Zhiliang Peng , Wei Huang , Shanzhi Gu , Lingxi Xie , Yaowei Wang , Jianbin Jiao , Qixiang Ye

Attribute guided face image synthesis aims to manipulate attributes on a face image. Most existing methods for image-to-image translation can either perform a fixed translation between any two image domains using a single attribute or…

计算机视觉与模式识别 · 计算机科学 2019-05-02 Behzad Bozorgtabar , Mohammad Saeed Rad , Hazım Kemal Ekenel , Jean-Philippe Thiran

Retargeting motion across characters with varying body shapes while preserving interaction semantics, such as self-contact and near-body proximity, remains a challenging problem. While recent geometry-aware approaches address this by…

图形学 · 计算机科学 2026-05-20 Soojin Choi , Seokhyeon Hong , Chaelin Kim , Junghyun Nam , Junhyuk Jeon , Junyong Noh

Personalized image generation requires text-to-image generative models that capture the core features of a reference subject to allow for controlled generation across different contexts. Existing methods face challenges due to complex…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Emanuele Aiello , Umberto Michieli , Diego Valsesia , Mete Ozay , Enrico Magli

Memory mechanism is a core component of LLM-based agents, enabling reasoning and knowledge discovery over long-horizon contexts. Existing agent memory systems are typically designed within isolated paradigms (e.g., explicit, parametric, or…

人工智能 · 计算机科学 2026-02-10 Xin Zhang , Kailai Yang , Chenyue Li , Hao Li , Qiyu Wei , Jun'ichi Tsujii , Sophia Ananiadou

In mission-critical domains such as law enforcement and medical diagnosis, the ability to explain and interpret the outputs of deep learning models is crucial for ensuring user trust and supporting informed decision-making. Despite…

计算机视觉与模式识别 · 计算机科学 2024-11-07 Bharat Chandra Yalavarthi , Nalini Ratha

Drag-based image editing using generative models provides intuitive control over image structures. However, existing methods rely heavily on manually provided masks and textual prompts to preserve semantic fidelity and motion precision.…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Sheng-Hao Liao , Shang-Fu Chen , Tai-Ming Huang , Wen-Huang Cheng , Kai-Lung Hua

Long-form video generation presents a dual challenge: models must capture long-range dependencies while preventing the error accumulation inherent in autoregressive decoding. To address these challenges, we make two contributions. First,…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Xiaofei Wu , Guozhen Zhang , Zhiyong Xu , Yuan Zhou , Qinglin Lu , Xuming He

Nerf-based Generative models have shown impressive capacity in generating high-quality images with consistent 3D geometry. Despite successful synthesis of fake identity images randomly sampled from latent space, adopting these models for…

计算机视觉与模式识别 · 计算机科学 2022-12-01 Yu Yin , Kamran Ghasedi , HsiangTao Wu , Jiaolong Yang , Xin Tong , Yun Fu

Despite significant advancements in image generation using advanced generative frameworks, cross-image integration of content and style remains a key challenge. Current generative models, while powerful, frequently depend on vague textual…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Shaoxu Li , Ye Pan

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing…

计算机视觉与模式识别 · 计算机科学 2023-06-01 Chao Xu , Junwei Zhu , Jiangning Zhang , Yue Han , Wenqing Chu , Ying Tai , Chengjie Wang , Zhifeng Xie , Yong Liu

Customized image generation is essential for creating personalized content based on user prompts, allowing large-scale text-to-image diffusion models to more effectively meet individual needs. However, existing models often neglect the…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Qingyu Shi , Lu Qi , Jianzong Wu , Jinbin Bai , Jingbo Wang , Yunhai Tong , Xiangtai Li

Robotic grasping is one of the most fundamental robotic manipulation tasks and has been the subject of extensive research. However, swiftly teaching a robot to grasp a novel target object in clutter remains challenging. This paper attempts…

机器人学 · 计算机科学 2025-01-07 Yang Yang , Houjian Yu , Xibai Lou , Yuanhao Liu , Changhyun Choi

Text-driven video generation witnesses rapid progress. However, merely using text prompts is not enough to depict the desired subject appearance that accurately aligns with users' intents, especially for customized content creation. In this…

计算机视觉与模式识别 · 计算机科学 2023-12-04 Yuming Jiang , Tianxing Wu , Shuai Yang , Chenyang Si , Dahua Lin , Yu Qiao , Chen Change Loy , Ziwei Liu

Recent text-to-image generation models have demonstrated impressive capability of generating text-aligned images with high fidelity. However, generating images of novel concept provided by the user input image is still a challenging task.…

计算机视觉与模式识别 · 计算机科学 2023-05-24 Yufan Zhou , Ruiyi Zhang , Tong Sun , Jinhui Xu

Conventionally, autoencoders are unsupervised representation learning tools. In this work, we propose a novel discriminative autoencoder. Use of supervised discriminative learning ensures that the learned representation is robust to…

计算机视觉与模式识别 · 计算机科学 2019-12-30 Anupriya Gogna , Angshul Majumdar

Instance based photo cartoonization is one of the challenging image stylization tasks which aim at transforming realistic photos into cartoon style images while preserving the semantic contents of the photos. State-of-the-art Deep Neural…

计算机视觉与模式识别 · 计算机科学 2019-11-15 Yugang Chen , Muchun Chen , Chaoyue Song , Bingbing Ni

Traditional image codecs emphasize signal fidelity and human perception, often at the expense of machine vision tasks. Deep learning methods have demonstrated promising coding performance by utilizing rich semantic embeddings optimized for…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Sha Guo , Zhuo Chen , Yang Zhao , Ning Zhang , Xiaotong Li , Lingyu Duan

This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dialogue and emotional…

声音 · 计算机科学 2025-09-22 Zhiwen Qian , Jinhua Liang , Huan Zhang

The ability to synthesize personalized group photos and specify the positions of each identity offers immense creative potential. While such imagery can be visually appealing, it presents significant challenges for existing technologies. A…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Yimeng Zhang , Tiancheng Zhi , Jing Liu , Shen Sang , Liming Jiang , Qing Yan , Sijia Liu , Linjie Luo
‹ 上一页 1 8 9 10 下一页 ›