中文
相关论文

相关论文: Language identification as improvement for lip-bas…

200 篇论文

It is widely agreed that open-vocabulary-based approaches outperform classical closed-set training solutions for recognizing unseen objects in images for semantic segmentation. Existing open-vocabulary approaches leverage vision-language…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Huadong Tang , Youpeng Zhao , Yan Huang , Min Xu , Jun Wang , Qiang Wu

Language identification from speech is a common preprocessing step in many spoken language processing systems. In recent years, this field has seen fast progress, mostly due to the use of self-supervised models pretrained on multilingual…

音频与语音处理 · 电气工程与系统科学 2022-07-04 Kunnar Kukk , Tanel Alumäe

Today's Automatic Speech Recognition systems only rely on acoustic signals and often don't perform well under noisy conditions. Performing multi-modal speech recognition - processing acoustic speech signals and lip-reading video…

计算机视觉与模式识别 · 计算机科学 2018-03-14 Matthijs Van keirsbilck , Bert Moons , Marian Verhelst

Multimodal large language models (MLLMs) typically extract visual features from the final layers of a pretrained Vision Transformer (ViT). This widespread deep-layer bias, however, is largely driven by empirical convention rather than…

计算机视觉与模式识别 · 计算机科学 2025-10-13 Haoran Chen , Junyan Lin , Xinghao Chen , Yue Fan , Jianfeng Dong , Xin Jin , Hui Su , Jinlan Fu , Xiaoyu Shen

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived…

音频与语音处理 · 电气工程与系统科学 2023-12-19 Ziyi Ni , Minglun Han , Feilong Chen , Linghui Meng , Jing Shi , Pin Lv , Bo Xu

Despite the impressive performance of autoregressive Language Models (LM) it has been shown that due to reporting bias, LMs lack visual knowledge, i.e. they do not know much about the visual world and its properties. To augment LMs with…

计算与语言 · 计算机科学 2026-03-10 Paula Ontalvilla , Aitor Ormazabal , Gorka Azkune

Large vision-language models (LVLMs) achieve strong performance on multimodal tasks, yet they often default to their language prior (LP) -- memorized textual patterns from pre-training while under-utilizing visual evidence. Prior analyses…

机器学习 · 计算机科学 2026-02-12 Lin Long , Changdae Oh , Seongheon Park , Sharon Li

In the quest for greater computer lip-reading performance there are a number of tacit assumptions which are either present in the datasets (high resolution for example) or in the methods (recognition of spoken visual units called visemes…

计算机视觉与模式识别 · 计算机科学 2018-04-26 Helen L. Bear , Gari Owen , Richard Harvey , Barry-John Theobald

Language identification is the task of automatically determining the identity of a language conveyed by a spoken segment. It has a profound impact on the multilingual interoperability of an intelligent speech system. Despite language…

计算与语言 · 计算机科学 2025-01-14 Yuting Nie , Junhong Zhao , Wei-Qiang Zhang , Jinfeng Bai

The performance of automated lip reading using visemes as a classification schema has achieved less success compared with the use of ASCII characters and words largely due to the problem of different words sharing identical visemes. The…

计算与语言 · 计算机科学 2020-12-15 Souheil Fenghour , Daqing Chen , Kun Guo , Perry Xiao

The abilities of large language models (LLMs) have recently progressed to unprecedented levels, paving the way to novel applications in a wide variety of areas. In computer vision, LLMs can be used to prime vision-language tasks such image…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Théophane Vallaeys , Mustafa Shukor , Matthieu Cord , Jakob Verbeek

Numerous studies have investigated the effectiveness of audio-visual multimodal learning for speech enhancement (AVSE) tasks, seeking a solution that uses visual data as auxiliary and complementary input to reduce the noise of noisy speech…

音频与语音处理 · 电气工程与系统科学 2022-02-02 Shang-Yi Chuang , Hsin-Min Wang , Yu Tsao

Sign language visual recognition from continuous multi-modal streams is still one of the most challenging fields. Recent advances in human actions recognition are exploiting the ascension of GPU-based learning from massive data, and are…

计算机视觉与模式识别 · 计算机科学 2020-09-23 Bassem Seddik , Najoua Essoukri Ben Amara

Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared…

Technologies for recognizing facial attributes like race, gender, age, and emotion have several applications, such as surveillance, advertising content, sentiment analysis, and the study of demographic trends and social behaviors. Analyzing…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Nouar AlDahoul , Myles Joshua Toledo Tan , Harishwar Reddy Kasireddy , Yasir Zaki

Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such audio-centric systems inherently exclude individuals who are deaf or hard of hearing. Visual…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Jeong Hun Yeo , Hyeongseop Rha , Sungjune Park , Junil Won , Yong Man Ro

Vision-language models (VLMs) have recently emerged as powerful representation learning systems that align visual observations with natural language concepts, offering new opportunities for semantic reasoning in safety-critical autonomous…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Ross Greer , Maitrayee Keskar , Angel Martinez-Sanchez , Parthib Roy , Shashank Shriram , Mohan Trivedi

Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model. This is partially due to the difficulty of obtaining natural language labels for tactile data and the…

计算机视觉与模式识别 · 计算机科学 2024-02-21 Letian Fu , Gaurav Datta , Huang Huang , William Chung-Ho Panitch , Jaimyn Drake , Joseph Ortiz , Mustafa Mukadam , Mike Lambeta , Roberto Calandra , Ken Goldberg

Traditional object detection systems are typically constrained to predefined categories, limiting their applicability in dynamic environments. In contrast, open-vocabulary object detection (OVD) enables the identification of objects from…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Tianyi Zhang , Antoine Simoulin , Kai Li , Sana Lakdawala , Shiqing Yu , Arpit Mittal , Hongyu Fu , Yu Lin

Recent advances in biometric systems have significantly improved the detection and prevention of fraudulent activities. However, as detection methods improve, attack techniques become increasingly sophisticated. Attacks on face recognition…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Lazaro Janier Gonzalez-Soler , Maciej Salwowski , Christoph Busch