中文
相关论文

相关论文: Multimodal Real-Time Anomaly Detection and Industr…

200 篇论文

We present a multimodal traffic light state detection using vision and sound, from the viewpoint of a quadruped robot navigating in urban settings. This is a challenging problem because of the visual occlusions and noise from robot…

机器人学 · 计算机科学 2025-11-12 Sagar Gupta , Akansel Cosgun

This study investigates the performance of robust anomaly detection models in industrial inspection, focusing particularly on their ability to handle noisy data. We propose to leverage the adaptation ability of meta learning approaches to…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Muhammad Aqeel , Shakiba Sharifi , Marco Cristani , Francesco Setti

Surveillance systems often struggle with managing vast amounts of footage, much of which is irrelevant, leading to inefficient storage and challenges in event retrieval. This paper addresses these issues by proposing an optimized video…

计算机视觉与模式识别 · 计算机科学 2025-01-29 Youssef Elmir , Hayet Touati , Ouassila Melizou

Despite advances in human activity recognition (HAR) with different modalities, a precise, robust, and accurate daily log system is not yet available. Current solutions primarily rely on controlled, lab-based data collection, which limits…

人机交互 · 计算机科学 2026-01-07 Lixing He , Bufang Yang , Di Duan , Zhenyu Yan , Guoliang Xing

Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been…

声音 · 计算机科学 2022-07-08 Junwen Xiong , Yu Zhou , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Surface defect detection in industrial scenarios is both crucial and technically demanding due to the wide variability in defect types, irregular shapes and sizes, fine-grained requirements, and complex material textures. Although recent…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Jiawei Hu

Concurrent Speaker Detection (CSD), the task of identifying active speakers and their overlaps in an audio signal, is essential for various audio applications, including meeting transcription, speaker diarization, and speech separation.…

音频与语音处理 · 电气工程与系统科学 2025-01-16 Amit Eliav , Sharon Gannot

The core challenge in industrial equipment anoma lous sound detection (ASD) lies in modeling the time-frequency coupling characteristics of acoustic features. Existing modeling methods are limited by local receptive fields, making it…

声音 · 计算机科学 2025-09-03 Chengyuan Ma , Peng Jia , Hongyue Guo , Wenming Yang

In recent years, multimodal anomaly detection methods have demonstrated remarkable performance improvements over video-only models. However, real-world multimodal data is often corrupted due to unforeseen environmental distortions. In this…

This paper introduces the task of Auditory Referring Multi-Object Tracking (AR-MOT), which dynamically tracks specific objects in a video sequence based on audio expressions and appears as a challenging problem in autonomous driving. Due to…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Jiacheng Lin , Jiajun Chen , Kunyu Peng , Xuan He , Zhiyong Li , Rainer Stiefelhagen , Kailun Yang

Anomalous sound detection for machine condition monitoring has great potential in the development of Industry 4.0. However, these anomalous sounds of machines are usually unavailable in normal conditions. Therefore, the models employed have…

音频与语音处理 · 电气工程与系统科学 2023-02-17 Jisheng Bai , Jianfeng Chen , Mou Wang , Muhammad Saad Ayub , Qingli Yan

Multimodal time series (MTS) anomaly detection is crucial for maintaining the safety and stability of working devices (e.g., water treatment system and spacecraft), whose data are characterized by multivariate time series with diverse…

机器学习 · 计算机科学 2023-10-18 Chaoyue Ding , Shiliang Sun , Jing Zhao

Status prediction and anomaly detection are two fundamental tasks in automatic IT systems monitoring. In this paper, a joint model Predictor & Anomaly Detector (PAD) is proposed to address these two issues under one framework. In our…

机器学习 · 计算机科学 2021-04-23 Run-Qing Chen , Guang-Hui Shi , Wan-Lei Zhao , Chang-Hui Liang

The rapid advancement of artificial intelligence (AI) has enabled sophisticated audio generation and voice cloning technologies, posing significant security risks for applications reliant on voice authentication. While existing datasets and…

声音 · 计算机科学 2025-05-22 Kunyang Huang , Bin Hu

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovisual fusion by…

Deepfakes are synthetic media generated using deep generative algorithms and have posed a severe societal and political threat. Apart from facial manipulation and synthetic voice, recently, a novel kind of deepfakes has emerged with either…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Vinaya Sree Katamneni , Ajita Rattani

To ensure reliability and service availability, next-generation networks are expected to rely on automated anomaly detection systems powered by advanced machine learning methods with the capability of handling multi-dimensional data. Such…

机器学习 · 计算机科学 2026-01-07 Mahsa Raeiszadeh , Amin Ebrahimzadeh , Roch H. Glitho , Johan Eker , Raquel A. F. Mini

Multi-modal industrial anomaly detection typically relies on separate models for each product category, fundamentally limiting practical scalability. When shifting to a unified paradigm that handles diverse classes simultaneously, detection…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Yangchen Wu , Huiqiang Xie

Visual objects often have acoustic signatures that are naturally synchronized with them in audio-bearing video recordings. For this project, we explore the multimodal feature aggregation for video instance segmentation task, in which we…

计算机视觉与模式识别 · 计算机科学 2023-01-26 Kaihui Zheng , Yuqing Ren , Zixin Shen , Tianxu Qin

Multimodal emotion recognition (MER) aims to infer human affect by jointly modeling audio and visual cues; however, existing approaches often struggle with temporal misalignment, weakly discriminative feature representations, and suboptimal…

多媒体 · 计算机科学 2026-01-21 Joe Dhanith P R , Shravan Venkatraman , Vigya Sharma , Santhosh Malarvannan