中文
相关论文

相关论文: Venus: Benchmarking and Empowering Multimodal Larg…

200 篇论文

With collective endeavors, multimodal large language models (MLLMs) are undergoing a flourishing development. However, their performances on image aesthetics perception remain indeterminate, which is highly desired in real-world…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Yipo Huang , Quan Yuan , Xiangfei Sheng , Zhichao Yang , Haoning Wu , Pengfei Chen , Yuzhe Yang , Leida Li , Weisi Lin

Assessing the aesthetic quality of graphic design is central to visual communication, yet remains underexplored in vision language models (VLMs). We investigate whether VLMs can evaluate design aesthetics in ways comparable to humans. Prior…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Arctanx An , Shizhao Sun , Danqing Huang , Mingxi Cheng , Yan Gao , Ji Li , Yu Qiao , Jiang Bian

Perceiving and producing aesthetic judgments is a fundamental yet underexplored capability for multimodal large language models (MLLMs). However, existing benchmarks for image aesthetic assessment (IAA) are narrow in perception scope or…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Guolong Wang , Heng Huang , Zhiqiang Zhang , Wentian Li , Feilong Ma , Xin Jin

With the vigorous development of mobile photography technology, major mobile phone manufacturers are scrambling to improve the shooting ability of equipments and the photo beautification algorithm of software. However, the improvement of…

计算机视觉与模式识别 · 计算机科学 2022-08-10 Xin Jin , Qiang Deng , Jianwen Lv , Heng Huang , Hao Lou , Chaoen Xiao

The highly abstract nature of image aesthetics perception (IAP) poses significant challenge for current multimodal large language models (MLLMs). The lack of human-annotated multi-modality aesthetic data further exacerbates this dilemma,…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Yipo Huang , Xiangfei Sheng , Zhichao Yang , Quan Yuan , Zhichao Duan , Pengfei Chen , Leida Li , Weisi Lin , Guangming Shi

Vision-language models (VLMs) have demonstrated impressive multimodal comprehension capabilities and are being deployed in an increasing number of online video understanding applications. While recent efforts extensively explore advancing…

分布式、并行与集群计算 · 计算机科学 2026-01-08 Shengyuan Ye , Bei Ouyang , Tianyi Qian , Liekang Zeng , Mu Yuan , Xiaowen Chu , Weijie Hong , Xu Chen

Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of these applications require an explicit aesthetic judgment. Most existing solutions reduce this…

Image Aesthetic Assessment (IAA) is a vital and intricate task that entails analyzing and assessing an image's aesthetic values, and identifying its highlights and areas for improvement. Traditional methods of IAA often concentrate on a…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Yuti Liu , Shice Liu , Junyuan Gao , Pengtao Jiang , Hao Zhang , Jinwei Chen , Bo Li

Visualization, a domain-specific yet widely used form of imagery, is an effective way to turn complex datasets into intuitive insights, and its value depends on whether data are faithfully represented, clearly communicated, and…

计算与语言 · 计算机科学 2026-03-03 Yupeng Xie , Zhiyang Zhang , Yifan Wu , Sirong Lu , Jiayi Zhang , Zhaoyang Yu , Jinlin Wang , Sirui Hong , Bang Liu , Chenglin Wu , Yuyu Luo

An aesthetics evaluation model is at the heart of predicting users' aesthetic experience and developing user interfaces with higher quality. However, previous methods on aesthetic evaluation largely ignore the interpretability of the model…

人机交互 · 计算机科学 2022-02-10 Xiaoran Wu

Visual layouts are essential in graphic design fields such as advertising, posters, and web interfaces. The application of generative models for content-aware layout generation has recently gained traction. However, these models fail to…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Sohan Patnaik , Rishabh Jain , Balaji Krishnamurthy , Mausoom Sarkar

The rapid expansion of mobile internet has resulted in a substantial increase in user-generated content (UGC) images, thereby making the thorough assessment of UGC images both urgent and essential. Recently, multimodal large language models…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Mingxing Li , Rui Wang , Lei Sun , Yancheng Bai , Xiangxiang Chu

Large multimodal models (LMMs) have demonstrated outstanding capabilities in various visual perception tasks, which has in turn made the evaluation of LMMs significant. However, the capability of video aesthetic quality assessment, which is…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Yunhao Li , Sijing Wu , Zhilin Gao , Zicheng Zhang , Qi Jia , Huiyu Duan , Xiongkuo Min , Guangtao Zhai

Evaluating image editing models remains challenging due to the coarse granularity and limited interpretability of traditional metrics, which often fail to capture aspects important to human perception and intent. Such metrics frequently…

Aspect ratio and spatial layout are two of the principal factors determining the aesthetic value of a photograph. But, incorporating these into the traditional convolution-based frameworks for the task of image aesthetics assessment is…

计算机视觉与模式识别 · 计算机科学 2022-06-29 Koustav Ghosal , Aljosa Smolic

Multimodal Large Language Models (MLLMs) are increasingly applied in Personalized Image Aesthetic Assessment (PIAA) as a scalable alternative to expert evaluations. However, their predictions may reflect subtle biases influenced by…

计算与语言 · 计算机科学 2025-09-16 Kun Li , Lai-Man Po , Hongzheng Yang , Xuyuan Xu , Kangcheng Liu , Yuzhi Zhao

Aesthetic-driven image cropping is crucial for applications like view recommendation and thumbnail generation, where visual appeal significantly impacts user engagement. A key factor in visual appeal is composition--the deliberate…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Yen-Hong Wong , Lai-Kuan Wong

We explore computational approaches for visual guidance to aid in creating aesthetically pleasing art and graphic design. Our work complements and builds on previous work that developed models for how humans look at images. Our approach…

计算机视觉与模式识别 · 计算机科学 2021-07-14 Qingyuan Zheng , Zhuoru Li , Adam Bargteil

Aesthetic image cropping aims to enhance the aesthetic quality of an image by improving its composition through spatial cropping. Previous methods often rely on saliency prediction or retrieval augmentation, ignoring the task's core…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Zhitong Dong , Chao Li , Jie Yu , Hao Chen

Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such scores do not necessarily imply faithful use of visual evidence. Prior studies have…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Jingru Chen , Yiming Liu , Mingtao Chen , Sijie Chen , Richeng Xuan , Liang Yang , Zhichao Hu , Fanyang Lu
‹ 上一页 1 2 3 10 下一页 ›