English
Related papers

Related papers: GPT-4o System Card

200 papers

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-22 Ryandhimas E. Zezario , Sabato M. Siniscalchi , Hsin-Min Wang , Yu Tsao

OpenAI's multimodal GPT-4o has demonstrated remarkable capabilities in image generation and editing, yet its ability to achieve world knowledge-informed semantic synthesis--seamlessly integrating domain knowledge, contextual reasoning, and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Ning Li , Jingran Zhang , Justin Cui

In this paper, we evaluate different abilities of GPT-4V including visual understanding, language understanding, visual puzzle solving, and understanding of other modalities such as depth, thermal, video, and audio. To estimate GPT-4V's…

Computation and Language · Computer Science 2023-10-26 Yang Wu , Shilong Wang , Hao Yang , Tian Zheng , Hongbo Zhang , Yanyan Zhao , Bing Qin

Large language models (LLMs) can extract information from veterinary electronic health records (EHRs), but performance differences between models, the effect of temperature settings, and the influence of text ambiguity have not been…

Recent advances in multimodal generative models have unlocked photorealistic, instruction-aligned image generation, yet leading systems like GPT-4o-Image remain proprietary and inaccessible. To democratize these capabilities, we present…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Junying Chen , Zhenyang Cai , Pengcheng Chen , Shunian Chen , Ke Ji , Xidong Wang , Yunjin Yang , Benyou Wang

Recent advances in large language models (LLMs) have enabled general-purpose systems to perform increasingly complex domain-specific reasoning without extensive fine-tuning. In the medical domain, decision-making often requires integrating…

Computation and Language · Computer Science 2025-08-14 Shansong Wang , Mingzhe Hu , Qiang Li , Mojtaba Safari , Xiaofeng Yang

While the recent advances in Multimodal Large Language Models (MLLMs) constitute a significant leap forward in the field, these models are predominantly confined to the realm of input-side multimodal comprehension, lacking the capacity for…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Zhanyu Wang , Longyue Wang , Zhen Zhao , Minghao Wu , Chenyang Lyu , Huayang Li , Deng Cai , Luping Zhou , Shuming Shi , Zhaopeng Tu

Animal ethology is an crucial aspect of animal research, and animal behavior labeling is the foundation for studying animal behavior. This process typically involves labeling video clips with behavioral semantic tags, a task that is…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Yiqi Wu , Xiaodan Hu , Ziming Fu , Siling Zhou , Jiangong Li

Recent advances in GPT-4o like multi-modality models have demonstrated remarkable progress for direct speech-to-speech conversation, with real-time speech interaction experience and strong speech understanding ability. However, current…

Sound · Computer Science 2024-12-09 Ze Yuan , Yanqing Liu , Shujie Liu , Sheng Zhao

End-to-end spoken dialogue models such as GPT-4o-audio have recently garnered significant attention in the speech domain. However, the evaluation of spoken dialogue models' conversational performance has largely been overlooked. This is…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Shengpeng Ji , Tianle Liang , Yangzhuo Li , Jialong Zuo , Minghui Fang , Jinzheng He , Yifu Chen , Zhengqing Liu , Ziyue Jiang , Xize Cheng , Siqi Zheng , Jin Xu , Junyang Lin , Zhou Zhao

We investigate the multilingual and multimodal performance of a large language model-based artificial intelligence (AI) system, GPT-4o, using a diverse set of physics concept inventories spanning multiple languages and subject categories.…

Physics Education · Physics 2025-07-14 Gerd Kortemeyer , Marina Babayeva , Giulia Polverini , Ralf Widenhorn , Bor Gregorcic

In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. While the emergence of large language…

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and…

Computation and Language · Computer Science 2025-01-06 Qinglin Zhang , Luyao Cheng , Chong Deng , Qian Chen , Wen Wang , Siqi Zheng , Jiaqing Liu , Hai Yu , Chaohong Tan , Zhihao Du , Shiliang Zhang

We present a vision and language model named MultiModal-GPT to conduct multi-round dialogue with humans. MultiModal-GPT can follow various instructions from humans, such as generating a detailed caption, counting the number of interested…

Computer Vision and Pattern Recognition · Computer Science 2023-06-14 Tao Gong , Chengqi Lyu , Shilong Zhang , Yudong Wang , Miao Zheng , Qian Zhao , Kuikun Liu , Wenwei Zhang , Ping Luo , Kai Chen

Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for audio-and-text large language models (LLMs). Previous…

MiniMind-O is an open 0.1B-scale omni model built on the MiniMind language model. It accepts text, speech, and image inputs, and returns both text and streaming speech. The release includes model code, checkpoints, and the main Parquet…

Sound · Computer Science 2026-05-06 Jingyao Gong

In this research short, we examine the potential of using GPT-4o, a state-of-the-art large language model (LLM) to undertake evidence synthesis and systematic assessment tasks. Traditional workflows for such tasks involve large groups of…

Computation and Language · Computer Science 2024-07-19 Elphin Tom Joe , Sai Dileep Koneru , Christine J Kirchhoff

As large language models (LLMs) advance, their role in higher education, particularly in free-response problem-solving, requires careful examination. This study assesses the performance of GPT-4o and o1-preview under realistic educational…

Computers and Society · Computer Science 2025-05-21 Ming Ding , Rasmus Kyng , Federico Solda , Weixuan Yuan

With the growing requirement for natural human-computer interaction, speech-based systems receive increasing attention as speech is one of the most common forms of daily communication. However, the existing speech models still experience…

Computation and Language · Computer Science 2025-10-22 Zuwei Long , Yunhang Shen , Chaoyou Fu , Heting Gao , Lijiang Li , Peixian Chen , Mengdan Zhang , Hang Shao , Jian Li , Jinlong Peng , Haoyu Cao , Ke Li , Rongrong Ji , Xing Sun

Multimodal conversational agents are highly desirable because they offer natural and human-like interaction. However, there is a lack of comprehensive end-to-end solutions to support collaborative development and benchmarking. While…

Human-Computer Interaction · Computer Science 2024-11-19 Qiang Sun , Yuanyi Luo , Sirui Li , Wenxiao Zhang , Wei Liu