中文
相关论文

相关论文: Users prefer Guetzli JPEG over same-sized libjpeg

200 篇论文

Digital pathology offers a groundbreaking opportunity to transform clinical practice in histopathological image analysis, yet faces a significant hurdle: the substantial file sizes of pathological Whole Slide Images (WSI). While current…

Many User interactive systems are proposed all methods are trying to implement as a user friendly and various approaches proposed but most of the systems not reached to the use specifications like user friendly systems with user interest,…

计算机视觉与模式识别 · 计算机科学 2012-04-12 R. Venkata Ramana Chary , D. Rajya Lakshmi , K. V. N. Sunitha

This paper introduces Zimtohrli, a novel, full-reference audio similarity metric designed for efficient and perceptually accurate quality assessment. In an era dominated by computationally intensive deep learning models and proprietary…

音频与语音处理 · 电气工程与系统科学 2025-10-01 Jyrki Alakuijala , Martin Bruse , Sami Boukortt , Jozef Marus Coldenhoff , Milos Cernak

Large Language Models (LLMs) such as ChatGPT have shown remarkable abilities in producing human-like text. However, it is unclear how accurately these models internalize concepts that shape human thought and behavior. Here, we developed a…

机器学习 · 计算机科学 2025-07-01 Hiro Taiyo Hamada , Ippei Fujisawa , Genji Kawakita , Yuki Yamada

Current perceptual similarity metrics operate at the level of pixels and patches. These metrics compare images in terms of their low-level colors and textures, but fail to capture mid-level similarities and differences in image layout,…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Stephanie Fu , Netanel Tamir , Shobhita Sundaram , Lucy Chai , Richard Zhang , Tali Dekel , Phillip Isola

Watching movies is one of the social activities typically done in groups. Emotion is the most vital factor that affects movie viewers' preferences. So, the emotional aspect of the movie needs to be determined and analyzed for further…

人工智能 · 计算机科学 2024-04-23 Adilet Yerkin , Elnara Kadyrgali , Yerdauit Torekhan , Pakizar Shamoi

Understanding how human brains interpret and process information is important. Here, we investigated the selectivity and inter-individual differences in human brain responses to images via functional MRI. In our first experiment, we found…

定量方法 · 定量生物学 2023-04-20 Zijin Gu , Keith Jamison , Mert R. Sabuncu , Amy Kuceyeski

We introduce CatSIM, a new similarity metric for binary and multinary two- and three-dimensional images and volumes. CatSIM uses a structural similarity image quality paradigm and is robust to small perturbations in location so that…

计算机视觉与模式识别 · 计算机科学 2020-04-21 Geoffrey Z. Thompson , Ranjan Maitra

Lossy Image compression is necessary for efficient storage and transfer of data. Typically the trade-off between bit-rate and quality determines the optimal compression level. This makes the image quality metric an integral part of any…

计算机视觉与模式识别 · 计算机科学 2021-07-16 Juan Carlos Mier , Eddie Huang , Hossein Talebi , Feng Yang , Peyman Milanfar

This work presents a comparative evaluation of machine translation systems applied to images containing textual information, a task that lies at the intersection of computer vision and natural language processing. The study compares three…

计算与语言 · 计算机科学 2026-05-29 Blai Puchol , Sergio Gómez González , Miguel Domingo , Francisco Casacuberta

As text-to-image systems continue to grow in popularity with the general public, questions have arisen about bias and diversity in the generated images. Here, we investigate properties of images generated in response to prompts which are…

计算机与社会 · 计算机科学 2023-02-15 Kathleen C. Fraser , Svetlana Kiritchenko , Isar Nejadgholi

There is a wide variety of music similarity detection algorithms, while discussions about music plagiarism in the real world are often based on audience perceptions. Therefore, we aim to conduct a study to examine the key criteria of human…

声音 · 计算机科学 2026-01-07 Daeun Hwang , Hyeonbin Hwang

For people first impressions of someone are of determining importance. They are hard to alter through further information. This begs the question if a computer can reach the same judgement. Earlier research has already pointed out that age,…

计算机视觉与模式识别 · 计算机科学 2016-03-11 Rasmus Rothe , Radu Timofte , Luc Van Gool

Object tags denote concrete entities and are central to many computer vision tasks, whereas abstract tags capture higher-level information, which is relevant for tasks that require a contextual, potentially subjective scene understanding.…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Darya Baranouskaya , Andrea Cavallaro

Digital cameras digitize scene light into linear raw representations, which the image signal processor (ISP) converts into display-ready outputs. While raw data preserves full sensor information--valuable for editing and vision…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Mahmoud Afifi , Ran Zhang , Michael S. Brown

JPEG is a popular image compression method widely used by individuals, data center, cloud storage and network filesystems. However, most recent progress on image compression mainly focuses on uncompressed images while ignoring trillions of…

图像与视频处理 · 电气工程与系统科学 2022-03-31 Lina Guo , Xinjie Shi , Dailan He , Yuanyuan Wang , Rui Ma , Hongwei Qin , Yan Wang

Feature similarity matching, which transfers the information of the reference frame to the query frame, is a key component in semi-supervised video object segmentation. If surjective matching is adopted, background distractors can easily…

计算机视觉与模式识别 · 计算机科学 2023-02-02 Suhwan Cho , Woo Jin Kim , MyeongAh Cho , Seunghoon Lee , Minhyeok Lee , Chaewon Park , Sangyoun Lee

In an ideal design pipeline, user interface (UI) design is intertwined with user research to validate decisions, yet studies are often resource-constrained during early exploration. Recent advances in multimodal large language models…

While text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt. While previous work has evaluated T2I alignment by proposing metrics, benchmarks, and templates for…

Large language models (LLMs) often exhibit subtle yet distinctive characteristics in their outputs that users intuitively recognize, but struggle to quantify. These "vibes" -- such as tone, formatting, or writing style -- influence user…

计算与语言 · 计算机科学 2025-04-22 Lisa Dunlap , Krishna Mandal , Trevor Darrell , Jacob Steinhardt , Joseph E Gonzalez