English
Related papers

Related papers: MVAD: A Benchmark Dataset for Multimodal AI-Genera…

200 papers

In recent years, monitoring hate speech and offensive language on social media platforms has become paramount due to its widespread usage among all age groups, races, and ethnicities. Consequently, there have been substantial research…

Machine Learning · Computer Science 2022-02-15 Aneri Rana , Sonali Jha

Recent advances in Visual Anomaly Detection (VAD) have introduced sophisticated algorithms leveraging embeddings generated by pre-trained feature extractors. Inspired by these developments, we investigate the adaptation of such algorithms…

Generating Audio Description (AD) for movies is a challenging task that requires fine-grained visual understanding and an awareness of the characters and their names. Currently, visual language models for AD generation are limited by a lack…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Tengda Han , Max Bain , Arsha Nagrani , Gül Varol , Weidi Xie , Andrew Zisserman

Autonomous vehicles (AVs) are poised to redefine transportation by enhancing road safety, minimizing human error, and optimizing traffic efficiency. The success of AVs depends on their ability to interpret complex, dynamic environments…

Multimedia · Computer Science 2025-07-11 Abolfazl Zarghani , Amirhossein Ebrahimi , Amir Malekesfandiari

Machine-generated music (MGM) has emerged as a powerful tool with applications in music therapy, personalised editing, and creative inspiration for the music community. However, its unregulated use threatens the entertainment, education,…

Sound · Computer Science 2026-02-16 Yupei Li , Hanqian Li , Lucia Specia , Björn W. Schuller

The demand for realistic and versatile character animation has surged, driven by its wide-ranging applications in various domains. However, the animation generation algorithms modeling human pose with 2D or 3D structures all face various…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Tianyu Sun , Zhoujie Fu , Bang Zhang , Guosheng Lin

Audio-visual speech recognition (AVSR) gains increasing attention from researchers as an important part of human-computer interaction. However, the existing available Mandarin audio-visual datasets are limited and lack the depth…

Sound · Computer Science 2023-06-06 Jianrong Wang , Yuchen Huo , Li Liu , Tianyi Xu , Qi Li , Sen Li

Multimodal manipulations (also known as audio-visual deepfakes) make it difficult for unimodal deepfake detectors to detect forgeries in multimedia content. To avoid the spread of false propaganda and fake news, timely detection is crucial.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Sahibzada Adil Shahzad , Ammarah Hashmi , Yan-Tsung Peng , Yu Tsao , Hsin-Min Wang

Automatic Deception Detection has been a hot research topic for a long time, using machine learning and deep learning to automatically detect deception, brings new light to this old field. In this paper, we proposed a voting-based method…

Machine Learning · Computer Science 2024-03-18 Lana Touma , Mohammad Al Horani , Manar Tailouni , Anas Dahabiah , Khloud Al Jallad

Existing deepfake detection research has primarily focused on scenarios where the manipulated subject is actively speaking, i.e., generating fabricated content by altering the speaker's appearance or voice. However, in realistic interaction…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Miao Liu , Fangda Wei , Jing Wang , Xinyuan Qian

Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortions are often asymmetric, where one modality may be severely degraded while the other…

Multimedia · Computer Science 2026-05-05 Mayesha Maliha R. Mithila , Mylene C. Q. Farias

VAD is a critical field in machine learning focused on identifying deviations from normal patterns in images, often challenged by the scarcity of anomalous data and the need for unsupervised training. To accelerate research and deployment…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Manuel Barusco , Francesco Borsatti , Arianna Stropeni , Davide Dalle Pezze , Gian Antonio Susto

During the process of driving, humans usually rely on multiple senses to gather information and make decisions. Analogously, in order to achieve embodied intelligence in autonomous driving, it is essential to integrate multidimensional…

Anomaly detection and localization in visual data, including images and videos, are crucial in machine learning and real-world applications. Despite rapid advancements in visual anomaly detection (VAD), interpreting these often black-box…

Machine Learning · Computer Science 2025-08-19 Yizhou Wang , Dongliang Guo , Sheng Li , Octavia Camps , Yun Fu

Anomaly detection in videos has been attracting an increasing amount of attention. Despite the competitive performance of recent methods on benchmark datasets, they typically lack desirable features such as modularity, cross-domain…

Computer Vision and Pattern Recognition · Computer Science 2021-03-23 Keval Doshi , Yasin Yilmaz

Inspired by the fact that different modalities in videos carry complementary information, we propose a Multimodal Semantic Attention Network(MSAN), which is a new encoder-decoder framework incorporating multimodal semantic attributes for…

Computer Vision and Pattern Recognition · Computer Science 2019-05-09 Liang Sun , Bing Li , Chunfeng Yuan , Zhengjun Zha , Weiming Hu

As a multimodal medium combining images and text, memes frequently convey implicit harmful content through metaphors and humor, rendering the detection of harmful memes a complex and challenging task. Although recent studies have made…

Computation and Language · Computer Science 2026-04-02 Hexiang Gu , Qifan Yu , Yuan Liu , Zikang Li , Saihui Hou , Jian Zhao , Zhaofeng He

The increasing utilization of surveillance cameras in smart cities, coupled with the surge of online video applications, has heightened concerns regarding public security and privacy protection, which propelled automated Video Anomaly…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Jing Liu , Yang Liu , Jieyu Lin , Jielin Li , Liang Cao , Peng Sun , Bo Hu , Liang Song , Azzedine Boukerche , Victor C. M. Leung

Wide-angle video is favored for its wide viewing angle and ability to capture a large area of scenery, making it an ideal choice for sports and adventure recording. However, wide-angle video is prone to deformation, exposure and other…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Bo Hu , Wei Wang , Chunyi Li , Lihuo He , Leida Li , Xinbo Gao

Diverse promising datasets have been designed to hold back the development of fake audio detection, such as ASVspoof databases. However, previous datasets ignore an attacking situation, in which the hacker hides some small fake clips in…

Sound · Computer Science 2023-12-19 Jiangyan Yi , Ye Bai , Jianhua Tao , Haoxin Ma , Zhengkun Tian , Chenglong Wang , Tao Wang , Ruibo Fu
‹ Prev 1 8 9 10 Next ›