English
Related papers

Related papers: NeuroCLIP: Neuromorphic Data Understanding by CLIP…

200 papers

Neuromorphic engineering aims to advance computing by mimicking the brain's efficient processing, where data is encoded as asynchronous temporal events. This eliminates the need for a synchronisation clock and minimises power consumption…

Neural and Evolutionary Computing · Computer Science 2026-02-03 Ben Walters , Yeshwanth Bethi , Taylor Kergan , Binh Nguyen , Amirali Amirsoleimani , Jason K. Eshraghian , Saeed Afshar , Mostafa Rahimi Azghadi

Neuromorphic processors are well-suited for efficiently handling sparse events from event-based cameras. However, they face significant challenges in the growth of computing demand and hardware costs as the input resolution increases. This…

Computer Vision and Pattern Recognition · Computer Science 2024-06-26 Cina Arjmand , Yingfu Xu , Kevin Shidqi , Alexandra F. Dobrita , Kanishkan Vadivel , Paul Detterer , Manolis Sifalakis , Amirreza Yousefzadeh , Guangzhi Tang

CLIP has achieved impressive zero-shot performance after pre-training on a large-scale dataset consisting of paired image-text data. Previous works have utilized CLIP by incorporating manually designed visual prompts like colored circles…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Jiedong Zhuang , Jiaqi Hu , Lianrui Mu , Rui Hu , Xiaoyu Liang , Jiangnan Ye , Haoji Hu

Pretraining robust vision or multimodal foundation models (e.g., CLIP) relies on large-scale datasets that may be noisy, potentially misaligned, and have long-tail distributions. Previous works have shown promising results in augmenting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Qingqing Cao , Mahyar Najibi , Sachin Mehta

Large-scale vision-language models such as CLIP have achieved remarkable success in zero-shot image recognition, yet their predictions remain largely opaque to human understanding. In contrast, Concept Bottleneck Models provide…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Onat Ozdemir , Anders Christensen , Stephan Alaniz , Zeynep Akata , Emre Akbas

Robotic grasping plays an important role in the field of robotics. The current state-of-the-art robotic grasping detection systems are usually built on the conventional vision, such as RGB-D camera. Compared to traditional frame-based…

Computer Vision and Pattern Recognition · Computer Science 2020-05-04 Bin Li , Hu Cao , Zhongnan Qu , Yingbai Hu , Zhenke Wang , Zichen Liang

Neuromorphic vision sensors, or event cameras, differ from conventional cameras in that they do not capture images at a specified rate. Instead, they asynchronously log local brightness changes at each pixel. As a result, event cameras only…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Paul Kielty , Cian Ryan , Mehdi Sefidgar Dilmaghani , Waseem Shariff , Joe Lemley , Peter Corcoran

Contrastive Language-Image Pre-training (CLIP) is an approach that has advanced research and applications in computer vision, fueling modern recognition systems and generative models. We believe that the main ingredient to the success of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Hu Xu , Saining Xie , Xiaoqing Ellen Tan , Po-Yao Huang , Russell Howes , Vasu Sharma , Shang-Wen Li , Gargi Ghosh , Luke Zettlemoyer , Christoph Feichtenhofer

Neuromorphic chip refers to an unconventional computing architecture that is modelled on biological brains. It is ideally suited for processing sensory data for intelligence computing, decision-making or context cognition. Despite rapid…

Emerging Technologies · Computer Science 2016-09-09 Shuchao Qin , Fengqiu Wang , Yujie Liu , Qing Wan , Xinran Wang , Yongbing Xu , Yi Shi , Xiaomu Wang , Rong Zhang

Neuromorphic computing has emerged as a promising avenue towards building the next generation of intelligent computing systems. It has been proposed that memristive devices, which exhibit history-dependent conductivity modulation, could…

Spikes are the currency in central nervous systems for information transmission and processing. They are also believed to play an essential role in low-power consumption of the biological systems, whose efficiency attracts increasing…

Neural and Evolutionary Computing · Computer Science 2020-05-05 Qiang Yu , Shenglan Li , Huajin Tang , Longbiao Wang , Jianwu Dang , Kay Chen Tan

Contrastive Language-Image Pretraining (CLIP) performs zero-shot image classification by mapping images and textual class representation into a shared embedding space, then retrieving the class closest to the image. This work provides a new…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Fawaz Sammani , Nikos Deligiannis

Two main routes of learning methods exist at present including error-driven global learning and neuroscience-oriented local learning. Integrating them into one network may provide complementary learning capabilities for versatile learning…

Neural and Evolutionary Computing · Computer Science 2021-06-23 Yujie Wu , Rong Zhao , Jun Zhu , Feng Chen , Mingkun Xu , Guoqi Li , Sen Song , Lei Deng , Guanrui Wang , Hao Zheng , Jing Pei , Youhui Zhang , Mingguo Zhao , Luping Shi

Motion deblurring addresses the challenge of image blur caused by camera or scene movement. Event cameras provide motion information that is encoded in the asynchronous event streams. To efficiently leverage the temporal information of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Xiaopeng Lin , Yulong Huang , Hongwei Ren , Zunchang Liu , Yue Zhou , Haotian Fu , Bojun Cheng

CLIP models learn transferable multi-modal features via image-text contrastive learning on internet-scale data. They are widely used in zero-shot classification, multi-modal retrieval, text-to-image diffusion, and as image encoders in large…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Marc-Antoine Lavoie , Anas Mahmoud , Aldo Zaimi , Arsene Fansi Tchango , Steven L. Waslander

Contrastive vision-language models like CLIP have shown great progress in transfer learning. In the inference stage, the proper text description, also known as prompt, needs to be carefully designed to correctly classify the given images.…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Tony Huang , Jack Chu , Fangyun Wei

Most existing Vision-and-Language (V&L) models rely on pre-trained visual encoders, using a relatively small set of manually-annotated data (as compared to web-crawled data), to perceive the visual world. However, it has been observed that…

Computer Vision and Pattern Recognition · Computer Science 2021-07-15 Sheng Shen , Liunian Harold Li , Hao Tan , Mohit Bansal , Anna Rohrbach , Kai-Wei Chang , Zhewei Yao , Kurt Keutzer

We present a neuromorphic split-computing framework for energy-efficient low-latency inference over optical inter-satellite links. The system partitions a spiking neural network (SNN) between edge and core nodes. To transmit sparse spiking…

Image and Video Processing · Electrical Eng. & Systems 2025-11-21 Zihang Song , Petar Popovski

Sign-language recognition has achieved substantial gains in classification accuracy in recent years; however, the latency and power requirements of most existing methods limit their suitability for real-time deployment. Neuromorphic sensing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Sarka Liskova , Olha Vedmedenko , Mazdak Fatahi , Matej Hoffmann , P. Michael Furlong , Giulia D Angelo

Open-vocabulary segmentation, powered by large visual-language models like CLIP, has expanded 2D segmentation capabilities beyond fixed classes predefined by the dataset, enabling zero-shot understanding across diverse scenes. Extending…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Weiwen Hu , Niccolò Parodi , Marcus Zepp , Ingo Feldmann , Oliver Schreer , Peter Eisert