English
Related papers

Related papers: Capturing More: Learning Multi-Domain Representati…

200 papers

Smart contracts are increasingly targeted by adversaries employing obfuscation techniques such as bogus code injection and control flow manipulation to evade vulnerability detection. Existing multimodal methods often process semantic,…

Cryptography and Security · Computer Science 2026-04-06 Minh-Dai Tran-Duong , Nguyen Hai Phong , Nguyen Chi Thanh , Doan Minh Trung , Tram Truong-Huu , Van-Hau Pham , Phan The Duy

Optical diffraction tomography is an indispensable tool for studying objects in three-dimensions due to its ability to accurately reconstruct scattering objects. Until now this technique has been limited to coherent light because spatial…

Histopathology and transcriptomics are fundamental modalities in oncology, encapsulating the morphological and molecular aspects of the disease. Multi-modal self-supervised learning has demonstrated remarkable potential in learning…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Tianyi Wang , Jianan Fan , Dingxin Zhang , Dongnan Liu , Yong Xia , Heng Huang , Weidong Cai

Multi-frame methods improve monocular depth estimation over single-frame approaches by aggregating spatial-temporal information via feature matching. However, the spatial-temporal feature leads to accuracy degradation in dynamic scenes. To…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Jiquan Zhong , Xiaolin Huang , Xiao Yu

In this work we propose a hybrid NN/HMM model for online Arabic handwriting recognition. The proposed system is based on Hidden Markov Models (HMMs) and Multi Layer Perceptron Neural Networks (MLPNNs). The input signal is segmented to…

Computer Vision and Pattern Recognition · Computer Science 2014-01-03 Najiba Tagougui , Houcine Boubaker , Monji Kherallah , Adel M. ALIMI

Optical Coherence Tomography (OCT) enables the acquisition of high-resolution, three-dimensional fingerprint data, capturing rich subsurface structures for robust biometric recognition. However, the high cost and time-consuming nature of…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Qingran Miao , Haixia Wang , Haohao Sun , Yilong Zhang

Multi-beam selection is one of the crucial technologies in hybrid beamforming systems for frequency-selective fading channels. Addressing the problem in the frequency domain facilitates the procedure of acquiring observations for analog…

Information Theory · Computer Science 2018-11-30 Hsiao-Lan Chiang , Wolfgang Rave , Gerhard Fettweis

We present STRUM (Spectral Transcription and Rhythm Understanding Model), an audio-to-chart pipeline that converts raw recordings into playable Clone Hero / YARG charts for drums, guitar, bass, vocals, and keys without any oracle metadata.…

Sound · Computer Science 2026-05-13 Joshua Opria

Face signatures, including size, shape, texture, skin tone, eye color, appearance, and scars/marks, are widely used as discriminative, biometric information for access control. Despite recent advancements in facial recognition systems,…

Computer Vision and Pattern Recognition · Computer Science 2021-11-05 Jennifer Hamblin , Kshitij Nikhal , Benjamin S. Riggan

In this paper, we propose a novel stroke constrained attention network (SCAN) which treats stroke as the basic unit for encoder-decoder based online handwritten mathematical expression recognition (HMER). Unlike previous methods which use…

Computer Vision and Pattern Recognition · Computer Science 2020-02-21 Jiaming Wang , Jun Du , Jianshu Zhang

Linguistic information is encoded at varying timescales (subwords, phrases, etc.) and communicative levels, such as syntax and semantics. Contextualized embeddings have analogously been found to capture these phenomena at distinctive layers…

Computation and Language · Computer Science 2022-10-24 Max Müller-Eberstein , Rob van der Goot , Barbara Plank

With the rise of short video content, efficient video summarization techniques for extracting key information have become crucial. However, existing methods struggle to capture the global temporal dependencies and maintain the semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Wenrui Li , Wei Han , Liang-Jian Deng , Ruiqin Xiong , Xiaopeng Fan

A novel diverse domain (DCT-SVD & DWT-SVD) watermarking scheme is proposed in this paper. Here, the watermark is embedded simultaneously onto the two domains. It is shown that an audio signal watermarked using this scheme has better…

Multimedia · Computer Science 2017-07-07 Jerrin Thomas Panachakel , Anurenjan P. R

Pedestrian detection is a critical task in robot perception. Multispectral modalities (visible light and thermal) can boost pedestrian detection performance by providing complementary visual information. Several gaps remain with…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Asiegbu Miracle Kanu-Asiegbu , Nitin Jotwani , Xiaoxiao Du

In this paper, we study semi-supervised Handwritten Mathematical Expression Recognition (HMER) via exploring both labeled data and extra unlabeled data. We propose a novel consistency regularization framework, termed SemiHMER, which…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Kehua Chen , Haoyang Shen

Hand action recognition is essential. Communication, human-robot interactions, and gesture control are dependent on it. Skeleton-based action recognition traditionally includes hands, which belong to the classes which remain challenging to…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Katharina Prasse , Steffen Jung , Yuxuan Zhou , Margret Keuper

Multi-channel speech enhancement extracts speech using multiple microphones that capture spatial cues. Effectively utilizing directional information is key for multi-channel enhancement. Deep learning shows great potential on multi-channel…

Sound · Computer Science 2023-09-21 Jiahui Pan , Pengjie Shen , Hui Zhang , Xueliang Zhang

Inspired by the great success of language model (LM)-based pre-training, recent studies in visual document understanding have explored LM-based pre-training methods for modeling text within document images. Among them, pre-training that…

Computer Vision and Pattern Recognition · Computer Science 2023-09-25 Daehee Kim , Yoonsik Kim , DongHyun Kim , Yumin Lim , Geewook Kim , Taeho Kil

Audio fingerprinting is a technique used to identify and match audio recordings based on their unique characteristics. It involves creating a condensed representation of an audio signal that can be used to quickly compare and match against…

Sound · Computer Science 2023-05-03 Aarón López-García

This research introduces an innovative approach to explore the cognitive and biologically inspired underpinnings of feature vector splitting for analyzing the significance of different attributes in e-security biometric signature…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Marcos Faundez , Moises Diaz , Miguel Angel Ferrer