English
Related papers

Related papers: SketchMind: A Multi-Agent Cognitive Framework for …

200 papers

In this paper, we delve into the intricate dynamics of Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) by addressing a critical yet overlooked aspect -- the choice of viewpoint during sketch creation. Unlike photo systems that…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Aneeshan Sain , Pinaki Nath Chowdhury , Subhadeep Koley , Ayan Kumar Bhunia , Yi-Zhe Song

Targeting the issues of "shortcuts" and insufficient contextual understanding in complex cross-modal reasoning of multimodal large models, this paper proposes a zero-shot multimodal reasoning component guided by human-like cognitive…

Artificial Intelligence · Computer Science 2025-09-16 Zhou-Peng Shou , Zhi-Qiang You , Fang Wang , Hai-Bo Liu

Large-scale distributed training of neural networks is often limited by network bandwidth, wherein the communication time overwhelms the local computation time. Motivated by the success of sketching methods in sub-linear/streaming…

Machine Learning · Computer Science 2020-01-24 Nikita Ivkin , Daniel Rothchild , Enayat Ullah , Vladimir Braverman , Ion Stoica , Raman Arora

While Multimodal Large Language Models (MLLMs) excel at visual understanding tasks through text reasoning, they often fall short in scenarios requiring visual imagination. Unlike current works that take predefined external toolkits or…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jintao Tong , Jiaqi Gu , Yujing Lou , Lubin Fan , Yixiong Zou , Yue Wu , Jieping Ye , Ruixuan Li

Text-to-image models can generate visually appealing images from text descriptions. Efforts have been devoted to improving model controls with prompt tuning and spatial conditioning. However, our formative study highlights the challenges…

Human-Computer Interaction · Computer Science 2025-02-12 Haichuan Lin , Yilin Ye , Jiazhi Xia , Wei Zeng

Imagining a colored realistic image from an arbitrarily drawn sketch is one of the human capabilities that we eager machines to mimic. Unlike previous methods that either requires the sketch-image pairs or utilize low-quantity detected…

Computer Vision and Pattern Recognition · Computer Science 2020-12-24 Bingchen Liu , Yizhe Zhu , Kunpeng Song , Ahmed Elgammal

Scientific illustrations demand both high information density and post-editability. However, current generative models have two major limitations: Frist, image generation models output rasterized images lacking semantic structure, making it…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Jianwen Sun , Fanrui Zhang , Yukang Feng , Chuanhao Li , Zizhen Li , Jiaxin Ai , Yifan Chang , Yu Dai , Kaipeng Zhang

Free-hand sketches are highly illustrative, and have been widely used by humans to depict objects or stories from ancient times to the present. The recent prevalence of touchscreen devices has made sketch creation a much easier task than…

Computer Vision and Pattern Recognition · Computer Science 2022-02-02 Peng Xu , Timothy M. Hospedales , Qiyue Yin , Yi-Zhe Song , Tao Xiang , Liang Wang

Advances in large language models (LLMs) are rapidly transforming scientific work, yet empirical evidence on how these systems reshape research activities remains limited. We report a mixed-methods pilot evaluation of an AI-orchestrated…

Computers and Society · Computer Science 2026-02-24 Yuan An

In this work, we propose an interactive general instruction framework SketchMeHow to guidance the common users to complete the daily tasks in real-time. In contrast to the conventional augmented reality-based instruction systems, the…

Human-Computer Interaction · Computer Science 2021-09-08 Haoran Xie , Yichen Peng , Hange Wang , Kazunori Miyata

Can a user create a deep generative model by sketching a single example? Traditionally, creating a GAN model has required the collection of a large-scale dataset of exemplars and specialized knowledge in deep learning. In contrast,…

Computer Vision and Pattern Recognition · Computer Science 2021-09-21 Sheng-Yu Wang , David Bau , Jun-Yan Zhu

Blended mathematical sensemaking in science (blended MSS) involves deep conceptual understanding of quantitative relationships describing phenomena in science and has been studies in various disciplines. However, no unified characterization…

Physics Education · Physics 2023-05-25 Leonora Kaldaras , Carl Wieman

Developing simple and expressive access controls -- interfaces to specify policies that define who should have access to resources and under what circumstances -- is a longstanding challenge in usable security. We present Sketch-based…

Human-Computer Interaction · Computer Science 2026-05-12 Kyzyl Monteiro , Sauvik Das

Assessing artistic creativity is foundational to creativity research and arts education, yet manual scoring (e.g., Torrance Tests of Creative Thinking) is labor-intensive at scale. Prior machine-learning approaches show promise for visual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Zhehan Zhang , Meihua Qian , Li Luo , Siyu Huang , Chaoyi Zhou , Ripon Saha , Xinxin Song

Multimodal large language models (LLMs) have demonstrated impressive capabilities in generating high-quality images from textual instructions. However, their performance in generating scientific images--a critical application for…

Generating medical images from human-drawn free-hand sketches holds promise for various important medical imaging applications. Due to the extreme difficulty in collecting free-hand sketch data in the medical domain, most deep…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Quan Huu Cap , Atsushi Fukuda

Chart question answering (CQA) has become a critical multimodal task for evaluating the reasoning capabilities of vision-language models. While early approaches have shown promising performance by focusing on visual features or leveraging…

Computation and Language · Computer Science 2025-05-30 Jingxuan Wei , Nan Xu , Junnan Zhu , Yanni Hao , Gaowei Wu , Bihui Yu , Lei Wang

Parsing sketches via semantic segmentation is attractive but challenging, because (i) free-hand drawings are abstract with large variances in depicting objects due to different drawing styles and skills; (ii) distorting lines drawn on the…

Computer Vision and Pattern Recognition · Computer Science 2019-10-15 Junkun Jiang , Ruomei Wang , Shujin Lin , Fei Wang

Machine learning practitioners often end up tunneling on low-level technical details like model architectures and performance metrics. Could early model development instead focus on high-level questions of which factors a model ought to pay…

Human-Computer Interaction · Computer Science 2023-03-07 Michelle S. Lam , Zixian Ma , Anne Li , Izequiel Freitas , Dakuo Wang , James A. Landay , Michael S. Bernstein

Scientific data visualization plays a crucial role in research by enabling the direct display of complex information and assisting researchers in identifying implicit patterns. Despite its importance, the use of Large Language Models (LLMs)…

Computation and Language · Computer Science 2024-03-20 Zhiyu Yang , Zihan Zhou , Shuo Wang , Xin Cong , Xu Han , Yukun Yan , Zhenghao Liu , Zhixing Tan , Pengyuan Liu , Dong Yu , Zhiyuan Liu , Xiaodong Shi , Maosong Sun
‹ Prev 1 3 4 5 6 7 10 Next ›