English
Related papers

Related papers: Exploring the Potential of Multimodal LLM with Kno…

200 papers

As multimodal large language models (MLLMs) grow increasingly capable, fixed benchmarks are gradually losing their effectiveness in evaluating high-level scientific understanding. In this paper, we introduce the Multimodal Academic Cover…

Computation and Language · Computer Science 2025-08-25 Mohan Jiang , Jin Gao , Jiahao Zhan , Dequan Wang

Continuous sign language recognition (SLR) deals with unaligned video-text pair and uses the word error rate (WER), i.e., edit distance, as the main evaluation metric. Since it is not differentiable, we usually instead optimize the learning…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Junfu Pu , Wengang Zhou , Hezhen Hu , Houqiang Li

Automatic Speech Recognition (ASR) is a technology that converts spoken words into text, facilitating interaction between humans and machines. One of the most common applications of ASR is Speech-To-Text (STT) technology, which simplifies…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-02 Jaeyoung Huh , Sangjoon Park , Jeong Eun Lee , Jong Chul Ye

Multi-Modal automatic speech recognition (ASR) techniques aim to leverage additional modalities to improve the performance of speech recognition systems. While existing approaches primarily focus on video or contextual information, the…

Sound · Computer Science 2023-12-27 Haoxu Wang , Fan Yu , Xian Shi , Yuezhang Wang , Shiliang Zhang , Ming Li

Recent advances in multimodal large language models (LLMs) have highlighted their potential for medical and surgical applications. However, existing surgical datasets predominantly adopt a Visual Question Answering (VQA) format with…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Tae-Min Choi , Tae Kyeong Jeong , Garam Kim , Jaemin Lee , Yeongyoon Koh , In Cheul Choi , Jae-Ho Chung , Jong Woong Park , Juyoun Park

This paper presents Seewo's systems for both tracks of the Multilingual Conversational Speech Language Model Challenge (MLC-SLM), addressing automatic speech recognition (ASR) and speaker diarization with ASR (SD-ASR). We introduce a…

Computation and Language · Computer Science 2025-06-19 Bo Li , Chengben Xu , Wufeng Zhang

Visual reasoning in multimodal large language models (MLLMs) has primarily been studied in static, fully observable settings, limiting their effectiveness in real-world environments where information is often incomplete due to occlusion or…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Weijie Zhou , Xuantang Xiong , Yi Peng , Manli Tao , Chaoyang Zhao , Honghui Dong , Ming Tang , Jinqiao Wang

Multimodal fusion of remote sensing images serves as a core technology for overcoming the limitations of single-source data and improving the accuracy of surface information extraction, which exhibits significant application value in fields…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Siyu Zhang , Lianlei Shan , Runhe Qiu

Current multilingual vision-language models either require a large number of additional parameters for each supported language, or suffer performance degradation as languages are added. In this paper, we propose a Scalable Multilingual…

Computer Vision and Pattern Recognition · Computer Science 2020-08-31 Andrea Burns , Donghyun Kim , Derry Wijaya , Kate Saenko , Bryan A. Plummer

Large language models (LLMs) have demonstrated immense capabilities in understanding textual data and are increasingly being adopted to help researchers accelerate scientific discovery through knowledge extraction (information retrieval),…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Robinson Umeike , Neil Getty , Fangfang Xia , Rick Stevens

Generative large language models (LLMs) exhibit impressive capabilities, which can be further augmented by integrating a pre-trained vision model into the original LLM to create a multimodal LLM (MLLM). However, this integration often…

Computation and Language · Computer Science 2025-08-14 Shikhar Srivastava , Md Yousuf Harun , Robik Shrestha , Christopher Kanan

Vision is often used as a complementary modality for audio speech recognition (ASR), especially in the noisy environment where performance of solo audio modality significantly deteriorates. After combining visual modality, ASR is upgraded…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Bo Xu , Cheng Lu , Yandong Guo , Jacob Wang

Large Language Models (LLMs) are currently under exploration for various tasks, including Automatic Speech Recognition (ASR), Machine Translation (MT), and even End-to-End Speech Translation (ST). In this paper, we present KIT's offline…

Computation and Language · Computer Science 2024-06-25 Sai Koneru , Thai-Binh Nguyen , Ngoc-Quan Pham , Danni Liu , Zhaolin Li , Alexander Waibel , Jan Niehues

State-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems primarily rely on acoustic information while disregarding additional multi-modal context. However, visual information are essential in disambiguation and adaptation. While…

Artificial Intelligence · Computer Science 2025-10-17 Supriti Sinhamahapatra , Jan Niehues

The integration of Artificial Intelligence (AI), particularly Large Language Model (LLM)-based systems, in education has shown promise in enhancing teaching and learning experiences. However, the advent of Multimodal Large Language Models…

Despite significant progress in Unified Multimodal Retrieval (UMR) powered by Large Multimodal Models (LMMs), existing embedding methods primarily focus on sample-level objectives via contrastive learning while overlooking the crucial…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Guosheng Zhang , Linkai Liu , Keyao Wang , Haixiao Yue , Zhiwen Tan , Xiao Tan

This paper presents Audio-Visual LLM, a Multimodal Large Language Model that takes both visual and auditory inputs for holistic video understanding. A key design is the modality-augmented training, which involves the integration of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Fangxun Shu , Lei Zhang , Hao Jiang , Cihang Xie

Automatic Speech Recognition (ASR) has achieved remarkable success with deep learning, driving advancements in conversational artificial intelligence, media transcription, and assistive technologies. However, ASR systems still struggle in…

Sound · Computer Science 2026-03-17 Haoyuan Yang , Yue Zhang , Liqiang Jing , John H. L. Hansen

Recent advancements in multimodal large language models (MLLMs) have achieved significant multimodal generation capabilities, akin to GPT-4. These models predominantly map visual information into language representation space, leveraging…

Computation and Language · Computer Science 2025-12-30 Yunxin Li , Zhenyu Liu , Baotian Hu , Wei Wang , Yuxin Ding , Xiaochun Cao , Min Zhang

6G networks promise revolutionary immersive communication experiences including augmented reality (AR), virtual reality (VR), and holographic communications. These applications demand high-dimensional multimodal data transmission and…

Machine Learning · Computer Science 2025-07-08 Yusong Zhang , Yuxuan Sun , Lei Guo , Wei Chen , Bo Ai , Deniz Gunduz