English
Related papers

Related papers: Multi-Scale Control Signal-Aware Transformer for M…

200 papers

This paper presents a framework for automatic synthesis of a control sequence for multi-agent systems governed by continuous linear dynamics under timed constraints. First, the motion of the agents in the workspace is abstracted into…

Systems and Control · Computer Science 2017-03-09 Sofie Andersson , Alexandros Nikou , Dimos V. Dimarogonas

Inspired by the facts that retinal cells actually segregate the visual scene into different attributes (e.g., spatial details, temporal motion) for respective neuronal processing, we propose to first decompose the input video into…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Ming Lu , Tong Chen , Dandan Ding , Fengqing Zhu , Zhan Ma

Recent advances in dance generation have enabled the automatic synthesis of 3D dance motions. However, existing methods still face significant challenges in simultaneously achieving high realism, precise dance-music synchronization, diverse…

Grasping is a core task in robotics with various applications. However, most current implementations are primarily designed for rigid items, and their performance drops considerably when handling fragile or deformable materials that require…

Robotics · Computer Science 2025-09-29 Leonel Giacobbe , Jingdao Chen , Chuangchuang Sun

Deep neural network is an effective choice to automatically recognize human actions utilizing data from various wearable sensors. These networks automate the process of feature extraction relying completely on data. However, various noises…

Signal Processing · Electrical Eng. & Systems 2021-01-05 Tanvir Mahmud , A. Q. M. Sazzad Sayyed , Shaikh Anowarul Fattah , Sun-Yuan Kung

With the advent of large language models and large-scale robotic datasets, there has been tremendous progress in high-level decision-making for object manipulation. These generic models are able to interpret complex tasks using language…

Robotics · Computer Science 2023-11-03 Wentao Yuan , Adithyavairavan Murali , Arsalan Mousavian , Dieter Fox

Deep learning-based methods for video pedestrian detection and tracking require large volumes of training data to achieve good performance. However, data acquisition in crowded public environments raises data privacy concerns -- we are not…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Matteo Fabbri , Guillem Braso , Gianluca Maugeri , Orcun Cetintas , Riccardo Gasparini , Aljosa Osep , Simone Calderara , Laura Leal-Taixe , Rita Cucchiara

This paper proposes a new transformer-based framework to learn class-specific object localization maps as pseudo labels for weakly supervised semantic segmentation (WSSS). Inspired by the fact that the attended regions of the one-class…

Computer Vision and Pattern Recognition · Computer Science 2022-03-08 Lian Xu , Wanli Ouyang , Mohammed Bennamoun , Farid Boussaid , Dan Xu

In this project, we aim to build a Text-to-Speech system able to produce speech with a controllable emotional expressiveness. We propose a methodology for solving this problem in three main steps. The first is the collection of emotional…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-08 Noé Tits

Controlling the movements of dynamic objects and the camera within generated videos is a meaningful yet challenging task. Due to the lack of datasets with comprehensive 6D pose annotations, existing text-to-video methods can not…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Xincheng Shuai , Henghui Ding , Zhenyuan Qin , Hao Luo , Xingjun Ma , Dacheng Tao

Micro-Actions (MAs) are an important form of non-verbal communication in social interactions, with potential applications in human emotional analysis. However, existing methods in Micro-Action Recognition often overlook the inherent subtle…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Jihao Gu , Kun Li , Fei Wang , Yanyan Wei , Zhiliang Wu , Hehe Fan , Meng Wang

Multivariate Time Series (MTS) data capture temporal behaviors to provide invaluable insights into various physical dynamic phenomena. In smart mobility, MTS plays a crucial role in providing temporal dynamics of behaviors such as maneuver…

Machine Learning · Computer Science 2024-09-12 Thabang Lebese

We propose a sequence-to-sequence singing synthesizer, which avoids the need for training data with pre-aligned phonetic and acoustic features. Rather than the more common approach of a content-based attention mechanism combined with an…

Sound · Computer Science 2020-02-21 Merlijn Blaauw , Jordi Bonada

Pretraining robust vision or multimodal foundation models (e.g., CLIP) relies on large-scale datasets that may be noisy, potentially misaligned, and have long-tail distributions. Previous works have shown promising results in augmenting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Qingqing Cao , Mahyar Najibi , Sachin Mehta

Multimodal transformer exhibits high capacity and flexibility to align image and text for visual grounding. However, the existing encoder-only grounding framework (e.g., TransVG) suffers from heavy computation due to the self-attention…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Fengyuan Shi , Ruopeng Gao , Weilin Huang , Limin Wang

Incorporating the dynamics knowledge into the model is critical for achieving accurate trajectory prediction while considering the spatial and temporal characteristics of the vessel. However, existing methods rarely consider the underlying…

Machine Learning · Computer Science 2023-03-22 Huimin Qiang , Zhiyuan Guo , Shiyuan Xie , Xiaodong Peng

This paper presents a novel and efficient method for characteristic mode decomposition in multi-structure systems. By leveraging the translation and rotation matrices of vector spherical wavefunctions, our approach enables the synthesis of…

Computational Engineering, Finance, and Science · Computer Science 2025-07-18 Chenbo Shi , Xin Gu , Shichen Liang , Jin Pan , Le Zuo

Creating realistic characters that can react to the users' or another character's movement can benefit computer graphics, games and virtual reality hugely. However, synthesizing such reactive motions in human-human interactions is a…

Graphics · Computer Science 2021-10-04 Qianhui Men , Hubert P. H. Shum , Edmond S. L. Ho , Howard Leung

In this paper, we present a novel architecture to realize fine-grained style control on the transformer-based text-to-speech synthesis (TransformerTTS). Specifically, we model the speaking style by extracting a time sequence of local style…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-18 Li-Wei Chen , Alexander Rudnicky

Humans interact with an object in many different ways by making contact at different locations, creating a highly complex motion space that can be difficult to learn, particularly when synthesizing such human interactions in a controllable…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Xiaohan Zhang , Bharat Lal Bhatnagar , Vladimir Guzov , Sebastian Starke , Gerard Pons-Moll
‹ Prev 1 3 4 5 6 7 10 Next ›