English
Related papers

Related papers: Training-Free Dense Hand Contact Estimation with M…

200 papers

We explore the human motion knowledge of Large Language Models (LLMs) through 3D avatar control. Given a motion instruction, we prompt LLMs to first generate a high-level movement plan with consecutive steps (High-level Planning), then…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Kunhang Li , Jason Naradowsky , Yansong Feng , Yusuke Miyao

Large vision-language models (LVLMs), such as the Generative Pre-trained Transformer 4-omni (GPT-4o), are emerging multi-modal foundation models which have great potential as powerful artificial-intelligence (AI) assistance tools for a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Keshav Bimbraw , Ye Wang , Jing Liu , Toshiaki Koike-Akino

Image-text matching (ITM) aims to address the fundamental challenge of aligning visual and textual modalities, which inherently differ in their representations, continuous, high-dimensional image features vs. discrete, structured text. We…

Multimedia · Computer Science 2025-07-14 Junyu Chen , Yihua Gao , Mingyong Li

How can large language models (LLMs) process and translate endangered languages? Many languages lack a large corpus to train a decent LLM; therefore existing LLMs rarely perform well in unseen, endangered languages. On the contrary, we…

Computation and Language · Computer Science 2024-11-13 Kexun Zhang , Yee Man Choi , Zhenqiao Song , Taiqi He , William Yang Wang , Lei Li

Affect recognition, encompassing emotions, moods, and feelings, plays a pivotal role in human communication. In the realm of conversational artificial intelligence, the ability to discern and respond to human affective cues is a critical…

Computation and Language · Computer Science 2024-08-06 Shutong Feng , Guangzhi Sun , Nurul Lubis , Wen Wu , Chao Zhang , Milica Gašić

Recently, Large Language Models (LLMs) have emerged as an alternative to training task-specific dialog agents, due to their broad reasoning capabilities and performance in zero-shot learning scenarios. However, many LLM-based dialog systems…

Computation and Language · Computer Science 2025-03-05 Dirk Väth , Ngoc Thang Vu

Handwritten mathematical expression recognition is a challenging problem due to the complicated two-dimensional structures, ambiguous handwriting input and variant scales of handwritten math symbols. To settle this problem, we utilize the…

Computer Vision and Pattern Recognition · Computer Science 2018-02-01 Jianshu Zhang , Jun Du , Lirong Dai

Multimodal large language models (MLLMs) have shown remarkable performance in vision-language tasks. However, existing MLLMs are primarily trained on generic datasets, limiting their ability to reason on domain-specific visual cues such as…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Hatef Otroshi Shahreza , Sébastien Marcel

Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack the ground locations and relations for visual entities.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Xiangtai Li , Tao Zhang , Yanwei Li , Haobo Yuan , Shihao Chen , Yikang Zhou , Jiahao Meng , Yueyi Sun , Shilin Xu , Lu Qi , Tianheng Cheng , Yi Lin , Zilong Huang , Wenhao Huang , Jiashi Feng , Guang Shi

Pretrained language models like BERT and T5 serve as crucial backbone encoders for dense retrieval. However, these models often exhibit limited generalization capabilities and face challenges in improving in domain accuracy. Recent research…

Computation and Language · Computer Science 2024-08-26 Kun Luo , Minghao Qin , Zheng Liu , Shitao Xiao , Jun Zhao , Kang Liu

We propose a real-time DNN-based technique to segment hand and object of interacting motions from depth inputs. Our model is called DenseAttentionSeg, which contains a dense attention mechanism to fuse information in different scales and…

Computer Vision and Pattern Recognition · Computer Science 2020-12-08 Zihao Bo , Hao Zhang , Junhai Yong , Feng Xu

Vision-language-action models (VLAs) have shown generalization capabilities in robotic manipulation tasks by inheriting from vision-language models (VLMs) and learning action generation. Most VLA models focus on interpreting vision and…

Traditional metasurface design is limited by the computational cost of full-wave simulations, preventing thorough exploration of complex configurations. Data-driven approaches have emerged as a solution to this bottleneck, replacing costly…

Optics · Physics 2026-01-28 Huanshu Zhang , Lei Kang , Sawyer D. Campbell , Douglas H. Werner

We empirically investigate proper pre-training methods to build good visual tokenizers, making Large Language Models (LLMs) powerful Multimodal Large Language Models (MLLMs). In our benchmark, which is curated to evaluate MLLMs visual…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Guangzhi Wang , Yixiao Ge , Xiaohan Ding , Mohan Kankanhalli , Ying Shan

We propose LangHOPS, the first Multimodal Large Language Model (MLLM) based framework for open-vocabulary object-part instance segmentation. Given an image, LangHOPS can jointly detect and segment hierarchical object and part instances from…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Yang Miao , Jan-Nico Zaech , Xi Wang , Fabien Despinoy , Danda Pani Paudel , Luc Van Gool

Multimodal Large Language Models (MLLMs) have shown success in various general image processing tasks, yet their application in medical imaging is nascent, lacking tailored models. This study investigates the potential of MLLMs in improving…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Ling Yang , Zhanyu Wang , Zhenghao Chen , Xinyu Liang , Luping Zhou

Recent Multi-Modal Large Language Models (MLLMs) have demonstrated strong capabilities in learning joint representations from text and images. However, their spatial reasoning remains limited. We introduce 3DFroMLLM, a novel framework that…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Noor Ahmed , Cameron Braunstein , Steffen Eger , Eddy Ilg

While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored. In this work, we…

Sound · Computer Science 2026-01-05 Akanksha Chuchra , Shukesh Reddy , Sudeepta Mishra , Abhijit Das , Abhinav Dhall

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, which requires multi-round dialogue spatial reasoning and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Xunyi Zhao , Gengze Zhou , Qi Wu

Lightweight Large Language Models (LwLLMs) are reduced-parameter, optimized models designed to run efficiently on consumer-grade hardware, offering significant advantages in resource efficiency, cost-effectiveness, and data privacy.…

Computation and Language · Computer Science 2025-06-10 Hongming Yang , Shi Lin , Jun Shao , Changting Lin , Donghai Zhu , Meng Han , Qinglei Kong