中文
相关论文

相关论文: A$^3$: Towards Advertising Aesthetic Assessment

200 篇论文

Textual-based prompt learning methods primarily employ multiple learnable soft prompts and hard class tokens in a cascading manner as text inputs, aiming to align image and text (category) spaces for downstream tasks. However, current…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Zheng Li , Yibing Song , Ming-Ming Cheng , Xiang Li , Jian Yang

The global fashion e-commerce market relies significantly on intelligent and aesthetic-aware outfit-completion tools to promote sales. While previous studies have approached the problem of fashion outfit-completion and compatible-item…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yuntian Wu , Xiaonan Hu , Ziqi Zhou , Hao Lu

Image cropping aims at improving the aesthetic quality of images by adjusting their composition. Most weakly supervised cropping methods (without bounding box supervision) rely on the sliding window mechanism. The sliding window mechanism…

计算机视觉与模式识别 · 计算机科学 2018-03-13 Debang Li , Huikai Wu , Junge Zhang , Kaiqi Huang

State-of-the-art T2I models are capable of generating high-resolution images given textual prompts. However, they still struggle with accurately depicting compositional scenes that specify multiple objects, attributes, and spatial…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Yixin Wan , Kai-Wei Chang

In this work, we share three insights for achieving state-of-the-art aesthetic quality in text-to-image generative models. We focus on three critical aspects for model improvement: enhancing color and contrast, improving generation across…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Daiqing Li , Aleks Kamko , Ehsan Akhgari , Ali Sabet , Linmiao Xu , Suhail Doshi

Pictures in physics education go beyond instructional functions and serve affective roles, such as attracting attention, creating fascination, and fostering engagement with the depicted content. Recognizing the importance of these affective…

物理教育 · 物理学 2024-11-08 Tatjana Zähringer , Raimund Girwidz , Andreas Müller

Multimodal large language models (MLLMs) are well suited to image aesthetic assessment, as they can capture high-level aesthetic features leveraging their cross-modal understanding capacity. However, the scarcity of multimodal aesthetic…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Boyang Liu , Yifan Hu , Senjie Jin , Shihan Dou , Gonglei Shi , Jie Shao , Tao Gui , Xuanjing Huang

Assessing the aesthetics of an image is challenging, as it is influenced by multiple factors including composition, color, style, and high-level semantics. Existing image aesthetic assessment (IAA) methods primarily rely on human-labeled…

计算机视觉与模式识别 · 计算机科学 2023-06-06 Junjie Ke , Keren Ye , Jiahui Yu , Yonghui Wu , Peyman Milanfar , Feng Yang

The widespread use of smartphones has made photography ubiquitous, yet a clear gap remains between ordinary users and professional photographers, who can identify aesthetic issues and provide actionable shooting guidance during capture. We…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Tianxiang Du , Hulingxiao He , Yuxin Peng

Emotion evoked by an advertisement plays a key role in influencing brand recall and eventual consumer choices. Automatic ad affect recognition has several useful applications. However, the use of content-based feature representations does…

计算机视觉与模式识别 · 计算机科学 2018-08-15 Abhinav Shukla , Harish Katti , Mohan Kankanhalli , Ramanathan Subramanian

Learning and improving large language models through human preference feedback has become a mainstream approach, but it has rarely been applied to the field of low-light image enhancement. Existing low-light enhancement evaluations…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Jun Yin , Yangfan He , Miao Zhang , Pengyu Zeng , Tianyi Wang , Shuai Lu , Xueqian Wang

Multimodal large language models (MLLMs), equipped with increasingly advanced planning and tool-use capabilities, are evolving into autonomous agents capable of performing multimodal web browsing and deep search in open-world environments.…

Faces and humans are crucial elements in social interaction and are widely included in everyday photos and videos. Therefore, a deep understanding of faces and humans will enable multi-modal assistants to achieve improved response quality…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Lixiong Qin , Shilong Ou , Miaoxuan Zhang , Jiangning Wei , Yuhang Zhang , Xiaoshuai Song , Yuchen Liu , Mei Wang , Weiran Xu

Automatic photo aesthetic assessment is a challenging artificial intelligence task. Existing computational approaches have focused on modeling a single aesthetic score or a class (good or bad), however these do not provide any details on…

计算机视觉与模式识别 · 计算机科学 2017-07-14 Gautam Malu , Raju S. Bapi , Bipin Indurkhya

AI alignment refers to models acting towards human-intended goals, preferences, or ethical principles. Given that most large-scale deep learning models act as black boxes and cannot be manually controlled, analyzing the similarity between…

计算机视觉与模式识别 · 计算机科学 2023-10-23 Jiyoung Lee , Seungho Kim , Seunghyun Won , Joonseok Lee , Marzyeh Ghassemi , James Thorne , Jaeseok Choi , O-Kil Kwon , Edward Choi

In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily target portraits and artistic images, lacking a systematic evaluation of interior scenes. We introduce…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Yuan Gao , Jin Song

Real-world applications could benefit from the ability to automatically generate a fine-grained ranking of photo aesthetics. However, previous methods for image aesthetics analysis have primarily focused on the coarse, binary categorization…

计算机视觉与模式识别 · 计算机科学 2016-07-28 Shu Kong , Xiaohui Shen , Zhe Lin , Radomir Mech , Charless Fowlkes

We tackle the problem of understanding visual ads where given an ad image, our goal is to rank appropriate human generated statements describing the purpose of the ad. This problem is generally addressed by jointly embedding images and…

计算机视觉与模式识别 · 计算机科学 2018-07-05 Karuna Ahuja , Karan Sikka , Anirban Roy , Ajay Divakaran

GAN-based image editing task aims at manipulating image attributes in the latent space of generative models. Most of the previous 2D and 3D-aware approaches mainly focus on editing attributes in images with ambiguous semantics or regions…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Zhijun Zhai , Zengmao Wang , Xiaoxiao Long , Kaixuan Zhou , Bo Du

The frontier of visual reasoning is shifting toward models like OpenAI o3, which can intelligently create and operate tools to transform images for problem-solving, also known as thinking-\textit{with}-images in chain-of-thought. Yet…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Ming Li , Jike Zhong , Shitian Zhao , Haoquan Zhang , Shaoheng Lin , Yuxiang Lai , Chen Wei , Konstantinos Psounis , Kaipeng Zhang