中文
相关论文

相关论文: FlexiFly: Interfacing the Physical World with Foun…

200 篇论文

Research attention on natural user interfaces (NUIs) for drone flights are rising. Nevertheless, NUIs are highly diversified, and primarily evaluated by different physical environments leading to hard-to-compare performance between such…

人机交互 · 计算机科学 2022-08-01 Zheng Li , Yiming Huang , Yui-Pan Yau , Pan Hui , Lik-Hang Lee

The physical interaction of aerial robots with their environment has countless potential applications and is an emerging area with many open challenges. Fully-actuated multirotors have been introduced to tackle some of these challenges.…

机器人学 · 计算机科学 2022-07-08 Azarakhsh Keipour

This paper introduces WavesFM, a novel Wireless Foundation Model (WFM) framework, capable of supporting a wide array of communication, sensing, and localization tasks. Our proposed architecture combines a shared Vision Transformer (ViT)…

信号处理 · 电气工程与系统科学 2025-04-22 Ahmed Aboulfotouh , Elsayed Mohammed , Hatem Abou-Zeid

Accurate localisation in planetary robotics enables the advanced autonomy required to support the increased scale and scope of future missions. The successes of the Ingenuity helicopter and multiple planetary orbiters lay the groundwork for…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Lachlan Holden , Feras Dayoub , Alberto Candela , David Harvey , Tat-Jun Chin

Communication is fundamental for multi-robot collaboration, with accurate radio mapping playing a crucial role in predicting signal strength between robots. However, modeling radio signal propagation in large and occluded environments is…

机器人学 · 计算机科学 2025-04-22 Yiming Luo , Yunfei Wang , Hongming Chen , Chengkai Wu , Ximin Lyu , Jinni Zhou , Jun Ma , Fu Zhang , Boyu Zhou

Geospatial foundation models (GeoFMs) promise broad generalisation capacity for Earth observation (EO) tasks, particularly under data-limited conditions. However, their large size poses a barrier to deployment on resource-constrained space…

Foundation models, large-scale, pre-trained deep-learning models adapted to a wide range of downstream tasks have gained significant interest lately in various deep-learning problems undergoing a paradigm shift with the rise of these…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Bobby Azad , Reza Azad , Sania Eskandari , Afshin Bozorgpour , Amirhossein Kazerouni , Islem Rekik , Dorit Merhof

In the realm of geospatial analysis, the diversity of remote sensors, encompassing both optical and microwave technologies, offers a wealth of distinct observational capabilities. Recognizing this, we present msGFM, a multisensor geospatial…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Boran Han , Shuai Zhang , Xingjian Shi , Markus Reichstein

Federated learning (FL) offers privacy-preserving decentralized machine learning, optimizing models at edge clients without sharing private data. Simultaneously, foundation models (FMs) have gained traction in the artificial intelligence…

机器学习 · 计算机科学 2023-10-06 Sixing Yu , J. Pablo Muñoz , Ali Jannesari

Humanoid robots are drawing significant attention as versatile platforms for complex motor control, human-robot interaction, and general-purpose physical intelligence. However, achieving efficient whole-body control (WBC) in humanoids…

机器人学 · 计算机科学 2026-02-10 Mingqi Yuan , Tao Yu , Wenqi Ge , Xiuyong Yao , Huijiang Wang , Jiayu Chen , Bo Li , Wei Zhang , Wenjun Zeng , Hua Chen , Xin Jin

For effective interactions with the open world, robots should understand how interactions with known and novel objects help them towards their goal. A key aspect of this understanding lies in detecting an object's affordances, which…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Anne Kemmeren , Gertjan Burghouts , Michael van Bekkum , Wouter Meijer , Jelle van Mil

Foundation models like ChatGPT and Sora that are trained on a huge scale of data have made a revolutionary social impact. However, it is extremely challenging for sensors in many different fields to collect similar scales of natural images…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Chenyang Lei , Liyi Chen , Jun Cen , Xiao Chen , Zhen Lei , Felix Heide , Qifeng Chen , Zhaoxiang Zhang

The emergence of multi-modal foundation models has markedly transformed the technology for autonomous driving, shifting away from conventional and mostly hand-crafted design choices towards unified, foundation-model-based approaches,…

机器人学 · 计算机科学 2026-03-24 Kemal Oksuz , Alexandru Buburuzan , Anthony Knittel , Yuhan Yao , Puneet K. Dokania

Recent advances in foundation models, particularly Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs), have facilitated the development of intelligent agents capable of performing complex tasks. By leveraging the…

Humans can effortlessly locate desired objects in cluttered environments, relying on a cognitive mechanism known as visual search to efficiently filter out irrelevant information and focus on task-related regions. Inspired by this process,…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Geng Li , Jinglin Xu , Yunzhen Zhao , Yuxin Peng

End-to-end learning directly maps sensory inputs to actions, creating highly integrated and efficient policies for complex robotics tasks. However, such models often struggle to generalize beyond their training scenarios, limiting…

机器人学 · 计算机科学 2025-05-19 Makram Chahine , Alex Quach , Alaa Maalouf , Tsun-Hsuan Wang , Daniela Rus

Cloud segmentation is a critical challenge in remote sensing image interpretation, as its accuracy directly impacts the effectiveness of subsequent data processing and analysis. Recently, vision foundation models (VFM) have demonstrated…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Xuechao Zou , Shun Zhang , Kai Li , Shiying Wang , Junliang Xing , Lei Jin , Congyan Lang , Pin Tao

ffective Human-Robot Interaction (HRI) is crucial for enhancing accessibility and usability in real-world robotics applications. However, existing solutions often rely on gesture- only or language-only commands, making interaction…

人机交互 · 计算机科学 2026-05-19 Yuzhi Lai , Shenghai Yuan , Peizheng Li , Boya Zhang , Benjamin Kiefer , Tianchen Deng , Andreas Zell

The real world is messy and unstructured. Uncovering critical information often requires active, goal-driven exploration. It remains to be seen whether Vision-Language Models (VLMs), which recently emerged as a popular zero-shot tool in…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Adam Pardyl , Dominik Matuszek , Mateusz Przebieracz , Marek Cygan , Bartosz Zieliński , Maciej Wołczyk