English
Related papers

Related papers: SwinLip: An Efficient Visual Speech Encoder for Li…

200 papers

The use of pretrained backbones with fine-tuning has been successful for 2D vision and natural language processing tasks, showing advantages over task-specific networks. In this work, we introduce a pretrained 3D backbone, called {\SST},…

Computer Vision and Pattern Recognition · Computer Science 2023-08-17 Yu-Qi Yang , Yu-Xiao Guo , Jian-Yu Xiong , Yang Liu , Hao Pan , Peng-Shuai Wang , Xin Tong , Baining Guo

The development of deep learning models in medical image analysis is majorly limited by the lack of large-sized and well-annotated datasets. Unsupervised learning does not require labels and is more suitable for solving medical image…

Computer Vision and Pattern Recognition · Computer Science 2023-01-06 Zi'an Xu , Yin Dai , Fayu Liu , Weibing Chen , Yue Liu , Lifu Shi , Sheng Liu , Yuhang Zhou

For streaming speech recognition, a Transformer-based encoder has been widely used with block processing. Although many studies addressed improving emission latency of transducers, little work has been explored for improving encoding…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-02 Emiru Tsunoo , Hayato Futami , Yosuke Kashiwagi , Siddhant Arora , Shinji Watanabe

Recent advancements in vision-language models have enhanced performance by increasing the length of visual tokens, making them much longer than text tokens and significantly raising computational costs. However, we observe that the visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Senqiao Yang , Yukang Chen , Zhuotao Tian , Chengyao Wang , Jingyao Li , Bei Yu , Jiaya Jia

In this paper, we present our approach to the Auto WCEBleedGen Challenge V2 2024. Our solution combines the Swin Transformer for the initial classification of bleeding frames and RT-DETR for further detection of bleeding in Wireless Capsule…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Sasidhar Alavala , Anil Kumar Vadde , Aparnamala Kancheti , Subrahmanyam Gorthi

Transformer network architecture has proven effective in speech enhancement. However, as its core module, self-attention suffers from quadratic complexity, making it infeasible for training on long speech utterances. In practical scenarios,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-10 Qiquan Zhang , Hongxu Zhu , Xinyuan Qian , Eliathamby Ambikairajah , Haizhou Li

Lipreading refers to understanding and further translating the speech of a speaker in the video into natural language. State-of-the-art lipreading methods excel in interpreting overlap speakers, i.e., speakers appear in both training and…

Computer Vision and Pattern Recognition · Computer Science 2023-02-07 Feng Xue , Yu Li , Deyin Liu , Yincen Xie , Lin Wu , Richang Hong

Understanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Munender Varshney , Ravindra Yadav , Vinay P. Namboodiri , Rajesh M Hegde

Transformer-based methods have achieved impressive image restoration performance due to their capacities to model long-range dependency compared to CNN-based methods. However, advances like SwinIR adopts the window-based and local attention…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Dafeng Zhang , Feiyu Huang , Shizhuo Liu , Xiaobing Wang , Zhezhu Jin

Direct speech-to-speech translation achieves high-quality results through the introduction of discrete units obtained from self-supervised learning. This approach circumvents delays and cascading errors associated with model cascading.…

Face super-resolution aims to recover high-quality facial images from severely degraded low-resolution inputs, but remains challenging due to the loss of fine structural details and identity-specific features. This work introduces SwinIFS,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Habiba Kausar , Saeed Anwar , Omar Jamal Hammad , Abdul Bais

Automated medical image segmentation can assist doctors to diagnose faster and more accurate. Deep learning based models for medical image segmentation have made great progress in recent years. However, the existing models fail to…

Image and Video Processing · Electrical Eng. & Systems 2023-04-26 Lei Shi , Tianyu Gao , Zheng Zhang , Junxing Zhang

Following the success in language domain, the self-attention mechanism (transformer) is adopted in the vision domain and achieving great success recently. Additionally, as another stream, multi-layer perceptron (MLP) is also explored in the…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Mocho Go , Hideyuki Tachibana

In this paper we present an end-to-end speech recognition model with Transformer encoders that can be used in a streaming speech recognition system. Transformer computation blocks based on self-attention are used to encode both audio and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-18 Qian Zhang , Han Lu , Hasim Sak , Anshuman Tripathi , Erik McDermott , Stephen Koo , Shankar Kumar

The large amount of audiovisual content being shared online today has drawn substantial attention to the prospect of audiovisual self-supervised learning. Recent works have focused on each of these modalities separately, while others have…

Machine Learning · Computer Science 2021-06-18 Pingchuan Ma , Rodrigo Mira , Stavros Petridis , Björn W. Schuller , Maja Pantic

Lipreading has witnessed a lot of progress due to the resurgence of neural networks. Recent works have placed emphasis on aspects such as improving performance by finding the optimal architecture or improving generalization. However, there…

Computer Vision and Pattern Recognition · Computer Science 2021-06-03 Pingchuan Ma , Brais Martinez , Stavros Petridis , Maja Pantic

Large-scale NLP models have been shown to significantly improve the performance on language tasks with no signs of saturation. They also demonstrate amazing few-shot capabilities like that of human beings. This paper aims to explore…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Ze Liu , Han Hu , Yutong Lin , Zhuliang Yao , Zhenda Xie , Yixuan Wei , Jia Ning , Yue Cao , Zheng Zhang , Li Dong , Furu Wei , Baining Guo

In this paper, we show that a simple self-supervised pre-trained audio model can achieve comparable inference efficiency to more complicated pre-trained models with speech transformer encoders. These speech transformers rely on mixing…

Sound · Computer Science 2024-02-09 Sungho Jeon , Ching-Feng Yeh , Hakan Inan , Wei-Ning Hsu , Rashi Rungta , Yashar Mehdad , Daniel Bikel

In recent years, vision transformers have been introduced into face recognition and analysis and have achieved performance breakthroughs. However, most previous methods generally train a single model or an ensemble of models to perform the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Lixiong Qin , Mei Wang , Chao Deng , Ke Wang , Xi Chen , Jiani Hu , Weihong Deng

Large datasets as required for deep learning of lip reading do not exist in many languages. In this paper we present the dataset GLips (German Lips) consisting of 250,000 publicly available videos of the faces of speakers of the Hessian…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Gerald Schwiebert , Cornelius Weber , Leyuan Qu , Henrique Siqueira , Stefan Wermter