English
Related papers

Related papers: MIMo grows! Simulating body and sensory developmen…

200 papers

GPT-4o, an all-encompassing model, represents a milestone in the development of large multi-modal language models. It can understand visual, auditory, and textual modalities, directly output audio, and support flexible duplex interaction.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-06 Zhifei Xie , Changqiao Wu

Task-oriented object grasping and rearrangement are critical skills for robots to accomplish different real-world manipulation tasks. However, they remain challenging due to partial observations of the objects and shape variations in…

Robotics · Computer Science 2026-03-06 Yichen Cai , Jianfeng Gao , Christoph Pohl , Tamim Asfour

This paper presents an innovative method for humanoid robots to acquire a comprehensive set of motor skills through reinforcement learning. The approach utilizes an achievement-triggered multi-path reward function rooted in developmental…

Robotics · Computer Science 2023-11-14 Fanxing Meng , Jing Xiao

Current human motion synthesis frameworks rely on global action descriptions, creating a modality gap that limits both motion understanding and generation capabilities. A single coarse description, such as run, fails to capture details such…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Pengfei Zhang , Pinxin Liu , Pablo Garrido , Hyeongwoo Kim , Bindita Chaudhuri

The promise of multimodal models for real-world applications has inspired research in visualizing and understanding their internal mechanics with the end goal of empowering stakeholders to visualize model behavior, perform model debugging,…

Self-recognition -- the ability to maintain an internal representation of one's own body within the environment -- underpins intelligent, autonomous behavior. As a foundational component of the minimal self, self-recognition provides the…

Animals can accomplish many incredible behavioral feats across a wide range of operational environments and scales that current robots struggle to match. One explanation for this performance gap is the extraordinary properties of the…

Robotics · Computer Science 2024-08-30 Saul Schaffer , Hima Hrithik Pamu , Victoria A. Webster-Wood

Existing video avatar models can produce fluid human animations, yet they struggle to move beyond mere physical likeness to capture a character's authentic essence. Their motions typically synchronize with low-level cues like audio rhythm,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Jianwen Jiang , Weihong Zeng , Zerong Zheng , Jiaqi Yang , Chao Liang , Wang Liao , Han Liang , Yuan Zhang , Mingyuan Gao

We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal reasoning. MiMo-VL-7B-RL outperforms Qwen2.5-VL-7B on 35 out of…

Learning robust and scalable visual representations from massive multi-view video data remains a challenge in computer vision and autonomous driving. Existing pre-training methods either rely on expensive supervised learning with 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Jialv Zou , Bencheng Liao , Qian Zhang , Wenyu Liu , Xinggang Wang

Bedside caregivers assess infants' pain at constant intervals by observing specific behavioral and physiological signs of pain. This standard has two main limitations. The first limitation is the intermittent assessment of pain, which might…

Computer Vision and Pattern Recognition · Computer Science 2019-01-17 Ghada Zamzmi , Dmitry Goldgof , Rangachar Kasturi , Yu Sun , Terri Ashmeade

We present a novel 4.5B parameter small language model that can handle multiple input and output modalities, including text, images, videos, and audio. Despite its small size, the model achieves near state-of-the-art performance on a…

Machine Learning · Computer Science 2024-11-12 Ben Koska , Mojmír Horváth

We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-V2-Flash adopts a hybrid attention architecture that…

Computation and Language · Computer Science 2026-01-09 Core Team , Bangjun Xiao , Bingquan Xia , Bo Yang , Bofei Gao , Bowen Shen , Chen Zhang , Chenhong He , Chiheng Lou , Fuli Luo , Gang Wang , Gang Xie , Hailin Zhang , Hanglong Lv , Hanyu Li , Heyu Chen , Hongshen Xu , Houbin Zhang , Huaqiu Liu , Jiangshan Duo , Jianyu Wei , Jiebao Xiao , Jinhao Dong , Jun Shi , Junhao Hu , Kainan Bao , Kang Zhou , Lei Li , Liang Zhao , Linghao Zhang , Peidian Li , Qianli Chen , Shaohui Liu , Shihua Yu , Shijie Cao , Shimao Chen , Shouqiu Yu , Shuo Liu , Tianling Zhou , Weijiang Su , Weikun Wang , Wenhan Ma , Xiangwei Deng , Bohan Mao , Bowen Ye , Can Cai , Chenghua Wang , Chengxuan Zhu , Chong Ma , Chun Chen , Chunan Li , Dawei Zhu , Deshan Xiao , Dong Zhang , Duo Zhang , Fangyue Liu , Feiyu Yang , Fengyuan Shi , Guoan Wang , Hao Tian , Hao Wu , Heng Qu , Hongfei Yi , Hongxu An , Hongyi Guan , Xing Zhang , Yifan Song , Yihan Yan , Yihao Zhao , Yingchun Lai , Yizhao Gao , Yu Cheng , Yuanyuan Tian , Yudong Wang , Zhen Tang , Zhengju Tang , Zhengtao Wen , Zhichao Song , Zhixian Zheng , Zihan Jiang , Jian Wen , Jiarui Sun , Jiawei Li , Jinlong Xue , Jun Xia , Kai Fang , Menghang Zhu , Nuo Chen , Qian Tu , Qihao Zhang , Qiying Wang , Rang Li , Rui Ma , Shaolei Zhang , Shengfan Wang , Shicheng Li , Shuhao Gu , Shuhuai Ren , Sirui Deng , Tao Guo , Tianyang Lu , Weiji Zhuang , Weikang Zhang , Weimin Xiong , Wenshan Huang , Wenyu Yang , Xin Zhang , Xing Yong , Xu Wang , Xueyang Xie , Yilin Jiang , Yixin Yang , Yongzhe He , Yu Tu , Yuanliang Dong , Yuchen Liu , Yue Ma , Yue Yu , Yuxing Xiang , Zhaojun Huang , Zhenru Lin , Zhipeng Xu , Zhiyang Chen , Zhonghua Deng , Zihan Zhang , Zihao Yue

Dynamic facial emotion is essential for believable AI-generated avatars, yet most systems remain visually static, limiting their use in simulations like virtual training for investigative interviews with abused children. We present a…

Human-Computer Interaction · Computer Science 2025-07-09 Pegah Salehi , Sajad Amouei Sheshkal , Vajira Thambawita , Michael A. Riegler , Pål Halvorsen

By the recent spread of machine learning in the robotics field, a humanoid that can act, perceive, and learn in the real world through contact with the environment needs to be developed. In this study, as one of the choices, we propose a…

Multimodal large language models (MLLMs) have shown remarkable potential in various domains, yet their application in the medical field is hindered by several challenges. General-purpose MLLMs often lack the specialized knowledge required…

Artificial Intelligence · Computer Science 2025-09-29 Guanghao Zhu , Zhitian Hou , Zeyu Liu , Zhijie Sang , Congkai Xie , Hongxia Yang

The remarkable progress of Multimodal Large Language Models (MLLMs) has attracted increasing attention to extend them to physical entities like legged robot. This typically requires MLLMs to not only grasp multimodal understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Gen Luo , Ganlin Yang , Ziyang Gong , Guanzhou Chen , Haonan Duan , Erfei Cui , Ronglei Tong , Zhi Hou , Tianyi Zhang , Zhe Chen , Shenglong Ye , Lewei Lu , Jingbo Wang , Wenhai Wang , Jifeng Dai , Yu Qiao , Rongrong Ji , Xizhou Zhu

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture,…

Reconstructing human dynamic vision from brain activity is a challenging task with great scientific significance. Although prior video reconstruction methods have made substantial progress, they still suffer from several limitations,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-20 Yizhuo Lu , Changde Du , Chong Wang , Xuanliu Zhu , Liuyun Jiang , Xujin Li , Huiguang He

Automated biomechanical testing has great potential for the development of VR applications, as initial insights into user behaviour can be gained in silico early in the design process. In particular, it allows prediction of user movements…