English
Related papers

Related papers: New starting point registration method for tagged …

200 papers

We present a pipeline for unbiased and robust multimodal registration of neuroimaging modalities with minimal pre-processing. While typical multimodal studies need to use multiple independent processing pipelines, with diverse options and…

Computer Vision and Pattern Recognition · Computer Science 2024-01-26 Adria Casamitjana , Juan Eugenio Iglesias , Raul Tudela , Aida Ninerola-Baizan , Roser Sala-Llonch

Sensor-based human activity recognition (HAR) has predominantly focused on Inertial Measurement Units and vision data, often overlooking the capabilities unique to pressure sensors, which capture subtle body dynamics and shifts in the…

Artificial Intelligence · Computer Science 2025-05-06 Lala Shakti Swarup Ray , Lars Krupp , Vitor Fortes Rey , Bo Zhou , Sungho Suh , Paul Lukowicz

We present HyperMorph, a learning-based strategy for deformable image registration that removes the need to tune important registration hyperparameters during training. Classical registration methods solve an optimization problem to find a…

Computer Vision and Pattern Recognition · Computer Science 2021-05-06 Andrew Hoopes , Malte Hoffmann , Bruce Fischl , John Guttag , Adrian V. Dalca

We introduce a data-driven approach for enabling word-level timestamp prediction in the Canary model. Accurate timestamp information is crucial for a variety of downstream tasks such as speech content retrieval and timed subtitles. While…

Computation and Language · Computer Science 2025-05-22 Ke Hu , Krishna Puvvada , Elena Rastorgueva , Zhehuai Chen , He Huang , Shuoyang Ding , Kunal Dhawan , Hainan Xu , Jagadeesh Balam , Boris Ginsburg

In text-audio retrieval (TAR) tasks, due to the heterogeneity of contents between text and audio, the semantic information contained in the text is only similar to certain frames within the audio. Yet, existing works aggregate the entire…

Sound · Computer Science 2023-03-31 Yifei Xin , Dongchao Yang , Yuexian Zou

While Wav2Vec 2.0 has been proposed for speech recognition (ASR), it can also be used for speech emotion recognition (SER); its performance can be significantly improved using different fine-tuning strategies. Two baseline methods, vanilla…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-22 Li-Wei Chen , Alexander Rudnicky

Automatic Mean Opinion Score (MOS) prediction is crucial to evaluate the perceptual quality of the synthetic speech. While recent approaches using pre-trained self-supervised learning (SSL) models have shown promising results, they only…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-01 Hui Wang , Shiwan Zhao , Xiguang Zheng , Yong Qin

The affine rank minimization (ARM) problem arises in many real-world applications. The goal is to recover a low-rank matrix from a small amount of noisy affine measurements. The original problem is NP-hard, and so directly solving the…

Information Theory · Computer Science 2020-01-08 Zhipeng Xue , Xiaojun Yuan , Junjie Ma , Yi Ma

Text entry is a critical capability for any modern computing experience, with lightweight augmented reality (AR) glasses being no exception. Designed for all-day wearability, a limitation of lightweight AR glass is the restriction to the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-25 Junxiao Shen , Roger Boldu , Arpit Kalla , Michael Glueck , Hemant Bhaskar Surale Amy Karlson

Tracking any point (TAP) is a fundamental yet challenging task in computer vision, requiring high precision and long-term motion reasoning. Recent attempts to combine RGB frames and event streams have shown promise, yet they typically rely…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Jiaxiong Liu , Zhen Tan , Jinpu Zhang , Yi Zhou , Hui Shen , Xieyuanli Chen , Dewen Hu

Physical motions are inherently continuous, and higher camera frame rates typically contribute to improved smoothness and temporal coherence. For the first time, we explore continuous representations of human motion sequences, featuring the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Chenghao Xu , Guangtao Lyu , Qi Liu , Jiexi Yan , Muli Yang , Cheng Deng

Deep learning-based multi-view facial capture methods have shown impressive accuracy while being several orders of magnitude faster than a traditional mesh registration pipeline. However, the existing systems (e.g. TEMPEH) are strictly…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Jing Li , Di Kang , Zhenyu He

Grammar Detection, also referred to as Parts of Speech Tagging of raw text, is considered an underlying building block of the various Natural Language Processing pipelines like named entity recognition, question answering, and sentiment…

Computation and Language · Computer Science 2022-12-06 Surya Teja Chavali , Charan Tej Kandavalli , Sugash T M

Phase retrieval is a problem encountered not only in speech and audio processing, but in many other fields such as optics. Iterative algorithms based on non-convex set projections are effective and frequently used for retrieving the phase…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-10 Tal Peer , Simon Welker , Timo Gerkmann

In the field of document analysis and recognition using mobile devices for capturing, and the field of object recognition in a video stream, an important problem is determining the time when the capturing process should be stopped.…

Computer Vision and Pattern Recognition · Computer Science 2020-02-12 Konstantin Bulatov , Boris Savelyev , Vladimir V. Arlazarov

The auditory system of humanoid robots has gained increased attention in recent years. This system typically acquires the surrounding sound field by means of a microphone array. Signals acquired by the array are then processed using various…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-05 Vladimir Tourbabin , Boaz Rafaely

We introduce MoRAG, a novel multi-part fusion based retrieval-augmented generation strategy for text-based human motion generation. The method enhances motion diffusion models by leveraging additional knowledge obtained through an improved…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Sai Shashank Kalakonda , Shubh Maheshwari , Ravi Kiran Sarvadevabhatla

Talking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing methods mainly focus on…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Hao Zhu , Huaibo Huang , Yi Li , Aihua Zheng , Ran He

In video synthetic aperture radar (SAR) imaging mode, the polar format algorithm (PFA) is more computational effective than the backprojection algorithm (BPA). However, the two-dimensional (2-D) interpolation in PFA greatly affects its…

Systems and Control · Electrical Eng. & Systems 2022-09-20 Jiawei Jiang , Yinwei Li , Qibin Zheng

Multilingual pretrained representations generally rely on subword segmentation algorithms to create a shared multilingual vocabulary. However, standard heuristic algorithms often lead to sub-optimal segmentation, especially for languages…

Computation and Language · Computer Science 2021-04-07 Xinyi Wang , Sebastian Ruder , Graham Neubig