English
Related papers

Related papers: CLIPSwarm: Converting text into formations of robo…

200 papers

Robotic swarms and mobile sensor networks are used for environmental monitoring in various domains and areas of operation. Especially in otherwise inaccessible environments decentralized robotic swarms can be advantageous due to their high…

Multiagent Systems · Computer Science 2019-03-14 Hannes Hornischer , Joshua Cherian Varughese , Ronald Thenius , Franz Wotawa , Manfred Füllsack , Thomas Schmickl

Easily accessible sensors, like drones with diverse onboard sensors, have greatly expanded studying animal behavior in natural environments. Yet, analyzing vast, unlabeled video data, often spanning hours, remains a challenge for machine…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Duc Pham , Matthew Hansen , Félicie Dhellemmes , Jens Krause , Pia Bideau

One of the major motifs in collective or swarm intelligence is that, even though individuals follow simple rules, the resulting global behavior can be complex and intelligent. In artificial swarm systems, such as swarm robots, the goal is…

Neural and Evolutionary Computing · Computer Science 2014-07-02 Donghwa Jeong , Kiju Lee

Advancements in generative models have sparked significant interest in generating images while adhering to specific structural guidelines. Scene graph to image generation is one such task of generating images which are consistent with the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Rameshwar Mishra , A V Subramanyam

We present SignCLIP, which re-purposes CLIP (Contrastive Language-Image Pretraining) to project spoken language text and sign language videos, two classes of natural languages of distinct modalities, into the same space. SignCLIP is an…

Computation and Language · Computer Science 2024-10-08 Zifan Jiang , Gerard Sant , Amit Moryossef , Mathias Müller , Rico Sennrich , Sarah Ebling

Pre-trained vision-language models~(VLMs) are the de-facto foundation models for various downstream tasks. However, scene text recognition methods still prefer backbones pre-trained on a single modality, namely, the visual modality, despite…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Shuai Zhao , Ruijie Quan , Linchao Zhu , Yi Yang

In this paper, we demonstrate that CLIP can also be adapted to downstream tasks where its vision-language alignment is suboptimally learned during pre-training on web-crawled data, all without requiring fine-tuning. We explore the case of…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Sohee Kim , Jisu Kang , Dunam Kim , Seokju Lee

In this paper, we propose Lan-grasp, a novel approach towards more appropriate semantic grasping and placing. We leverage foundation models to equip the robot with a semantic understanding of object geometry, enabling it to identify the…

Swarm robots, which are inspired from the way insects behave collectively in order to achieve a common goal, have become a major part of research with applications involving search and rescue, area exploration, surveillance etc. In this…

Robotics · Computer Science 2024-08-31 Prateek , Pawan Wadhwani , Reshesh Kumar Pathak , Mayur Bhosale , A Helen Victoria

We present a method for generating colored 3D shapes from natural language. To this end, we first learn joint embeddings of freeform text descriptions and colored 3D shapes. Our model combines and extends learning by association and metric…

Computer Vision and Pattern Recognition · Computer Science 2018-03-23 Kevin Chen , Christopher B. Choy , Manolis Savva , Angel X. Chang , Thomas Funkhouser , Silvio Savarese

This paper presents a novel planning method that achieves navigation of multi-robot formations in cluttered environments, while maintaining the formation throughout the robots motion. The method utilises a decentralised approach to find…

Robotics · Computer Science 2023-07-17 Jeppe Heini Mikkelsen , Matteo Fumagalli

We propose a decentralized algorithm to collaboratively transport arbitrarily shaped objects using a swarm of robots. Our approach starts with a task allocation phase that sequentially distributes locations around the object to be…

Robotics · Computer Science 2021-06-08 Vivek Shankar Vardharajan , Karthik Soma , Giovanni Beltrame

Multi-hand semantic grasp generation aims to generate feasible and semantically appropriate grasp poses for different robotic hands based on natural language instructions. Although the task is highly valuable, due to the lack of multihand…

We present PAPERCLIP (Proposal Abstracts Provide an Effective Representation for Contrastive Language-Image Pre-training), a method which associates astronomical observations imaged by telescopes with natural language using a neural network…

Instrumentation and Methods for Astrophysics · Physics 2024-03-15 Siddharth Mishra-Sharma , Yiding Song , Jesse Thaler

Developing autonomous physical human-robot interaction (pHRI) systems is limited by the scarcity of large-scale training data to learn robust robot behaviors for real-world applications. In this paper, we introduce a zero-shot…

Robotics · Computer Science 2026-04-13 Junxiang Wang , Xinwen Xu , Tiancheng Wu , Julian Millan , Nir Pechuk , Zackory Erickson

Swarm behaviour engineering is an area of research that seeks to investigate methods and techniques for coordinating computation and action within groups of simple agents to achieve complex global goals like pattern formation, collective…

Artificial Intelligence · Computer Science 2025-08-13 Gianluca Aguzzi , Roberto Casadei , Mirko Viroli

Training a text-to-image generator in the general domain (e.g., Dall.e, CogView) requires huge amounts of paired text-image data, which is too expensive to collect. In this paper, we propose a self-supervised scheme named as CLIP-GEN for…

Computer Vision and Pattern Recognition · Computer Science 2022-03-02 Zihao Wang , Wei Liu , Qian He , Xinglong Wu , Zili Yi

Conventional robot social behavior generation has been limited in flexibility and autonomy, relying on predefined motions or human feedback. This study proposes CRISP (Critique-and-Replan for Interactive Social Presence), an autonomous…

Robotics · Computer Science 2026-03-23 Jiyu Lim , Youngwoo Yoon , Kwanghyun Park

Multimodal video summarization requires visual features that align semantically with language generation. Traditional approaches rely on CNN features trained for object classification, which represent visual concepts as discrete categories…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Maham Nazir , Muhammad Aqeel , Richong Zhang , Francesco Setti

Bearing measurements,as the most common modality in nature, have recently gained traction in multi-robot systems to enhance mutual localization and swarm collaboration. Despite their advantages, challenges such as sensory noise, obstacle…

Robotics · Computer Science 2024-01-17 Yingjian Wang , Xiangyong Wen , Fei Gao