English
Related papers

Related papers: M$^{2}$Chat: Empowering VLM for Multimodal LLM Int…

200 papers

Advances in large language models (LLMs) and real-time speech recognition now make it possible to issue any graphical user interface (GUI) action through natural language and receive the corresponding system response directly through the…

Human-Computer Interaction · Computer Science 2025-10-10 Hans G. W. van Dam

The increasing availability of multimodal data across text, tables, and images presents new challenges for developing models capable of complex cross-modal reasoning. Existing methods for Multimodal Multi-hop Question Answering (MMQA) often…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Qi Zhi Lim , Chin Poo Lee , Kian Ming Lim , Kalaiarasi Sonai Muthu Anbananthen

Large Language Models (LLMs), such as ChatGPT, have demonstrated the capability to generate human like, natural responses across a range of tasks, including task oriented dialogue and question answering. However, their application in real…

Computation and Language · Computer Science 2025-05-01 Naheed Rayhan , Md. Ashrafuzzaman

In this technical report, we target generating anthropomorphized personas for LLM-based characters in an online manner, including visual appearance, personality and tones, with only text descriptions. To achieve this, we first leverage the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Yilin Zhao , Xinbin Yuan , Shanghua Gao , Zhijie Lin , Qibin Hou , Jiashi Feng , Daquan Zhou

Although LLMs demonstrate proficiency in several text-based reasoning and planning tasks, their implementation in robotics control is constrained by significant deficiencies: (1) LLM agents are designed to work mainly with textual inputs…

Artificial Intelligence · Computer Science 2025-10-17 Shuang Ao , Flora D. Salim , Simon Khan

We report the development of Alter3, a humanoid robot capable of generating spontaneous motion using a Large Language Model (LLM), specifically GPT-4. This achievement was realized by integrating GPT-4 into our proprietary android, Alter3,…

Robotics · Computer Science 2023-12-12 Takahide Yoshida , Atsushi Masumori , Takashi Ikegami

In a rapidly evolving digital landscape autonomous tools and robots are becoming commonplace. Recognizing the significance of this development, this paper explores the integration of Large Language Models (LLMs) like Generative pre-trained…

Human-Computer Interaction · Computer Science 2024-03-22 Younes Lakhnati , Max Pascher , Jens Gerken

Generative Large Language Models (LLMs) show potential in data analysis, yet their full capabilities remain uncharted. Our work explores the capabilities of LLMs for creating and refining visualizations via conversational interfaces. We…

Human-Computer Interaction · Computer Science 2023-11-10 Matt-Heun Hong , Anamaria Crisan

Advancements in Multimodal Large Language Models (MLLMs) have improved human motion understanding. However, these models remain constrained by their "instruct-only" nature, lacking interactivity and adaptability for diverse analytical…

Artificial Intelligence · Computer Science 2025-02-28 Lei Li , Sen Jia , Jianhao Wang , Zhaochong An , Jiaang Li , Jenq-Neng Hwang , Serge Belongie

Image-text matching (ITM) aims to address the fundamental challenge of aligning visual and textual modalities, which inherently differ in their representations, continuous, high-dimensional image features vs. discrete, structured text. We…

Multimedia · Computer Science 2025-07-14 Junyu Chen , Yihua Gao , Mingyong Li

Generating human portraits is a hot topic in the image generation area, e.g. mask-to-face generation and text-to-face generation. However, these unimodal generation methods lack controllability in image generation. Controllability can be…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Debin Meng , Christos Tzelepis , Ioannis Patras , Georgios Tzimiropoulos

Large language models (LMs) are typically adapted to improve performance on new contexts (\eg text prompts that define new tasks or domains) through fine-tuning or prompting. However, there is an accuracy compute tradeoff -- fine-tuning…

Machine Learning · Computer Science 2024-11-12 Tong Chen , Hao Fang , Patrick Xia , Xiaodong Liu , Benjamin Van Durme , Luke Zettlemoyer , Jianfeng Gao , Hao Cheng

The development of large language models (LLMs) has greatly advanced the field of multimodal understanding, leading to the emergence of large multimodal models (LMMs). In order to enhance the level of visual comprehension, recent studies…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Ao Zhang , Yuan Yao , Wei Ji , Zhiyuan Liu , Tat-Seng Chua

Conversational systems must be robust to user interactions that naturally exhibit diverse conversational traits. Capturing and simulating these diverse traits coherently and efficiently presents a complex challenge. This paper introduces…

Computation and Language · Computer Science 2024-10-29 Rafael Ferreira , David Semedo , João Magalhães

Multimodal Vision-Language Models (VLMs) enable powerful applications from their fused understanding of images and language, but many perform poorly on UI tasks due to the lack of UI training data. In this paper, we adapt a recipe for…

Human-Computer Interaction · Computer Science 2023-10-10 Yue Jiang , Eldon Schoop , Amanda Swearngin , Jeffrey Nichols

The revolution of artificial intelligence content generation has been rapidly accelerated with the booming text-to-image (T2I) diffusion models. Within just two years of development, it was unprecedentedly of high-quality, diversity, and…

Artificial Intelligence · Computer Science 2023-10-16 Zeqiang Lai , Xizhou Zhu , Jifeng Dai , Yu Qiao , Wenhai Wang

Conventional Voice Assistants (VAs) rely on traditional language models to discern user intent and respond to their queries, leading to interactions that often lack a broader contextual understanding, an area in which Large Language Models…

Human-Computer Interaction · Computer Science 2024-12-02 Amama Mahmood , Junxiang Wang , Bingsheng Yao , Dakuo Wang , Chien-Ming Huang

Recent text-to-video (T2V) generation methods have seen significant advancements. However, the majority of these works focus on producing short video clips of a single event (i.e., single-scene videos). Meanwhile, recent large language…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Han Lin , Abhay Zala , Jaemin Cho , Mohit Bansal

Perception and understanding are two pillars of computer vision. While multimodal large language models (MLLM) have demonstrated remarkable visual understanding capabilities, they arguably lack accurate perception abilities, e.g. the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Qing Jiang , Gen Luo , Yuqin Yang , Yuda Xiong , Yihao Chen , Zhaoyang Zeng , Tianhe Ren , Lei Zhang

Addressing the issues of who saying what to whom in multi-party conversations (MPCs) has recently attracted a lot of research attention. However, existing methods on MPC understanding typically embed interlocutors and utterances into…

Computation and Language · Computer Science 2023-07-19 Jia-Chen Gu , Zhen-Hua Ling , Quan Liu , Cong Liu , Guoping Hu