中文
相关论文

相关论文: MIMo grows! Simulating body and sensory developmen…

200 篇论文

GPT-4o, an all-encompassing model, represents a milestone in the development of large multi-modal language models. It can understand visual, auditory, and textual modalities, directly output audio, and support flexible duplex interaction.…

音频与语音处理 · 电气工程与系统科学 2024-11-06 Zhifei Xie , Changqiao Wu

Task-oriented object grasping and rearrangement are critical skills for robots to accomplish different real-world manipulation tasks. However, they remain challenging due to partial observations of the objects and shape variations in…

机器人学 · 计算机科学 2026-03-06 Yichen Cai , Jianfeng Gao , Christoph Pohl , Tamim Asfour

This paper presents an innovative method for humanoid robots to acquire a comprehensive set of motor skills through reinforcement learning. The approach utilizes an achievement-triggered multi-path reward function rooted in developmental…

机器人学 · 计算机科学 2023-11-14 Fanxing Meng , Jing Xiao

Current human motion synthesis frameworks rely on global action descriptions, creating a modality gap that limits both motion understanding and generation capabilities. A single coarse description, such as run, fails to capture details such…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Pengfei Zhang , Pinxin Liu , Pablo Garrido , Hyeongwoo Kim , Bindita Chaudhuri

The promise of multimodal models for real-world applications has inspired research in visualizing and understanding their internal mechanics with the end goal of empowering stakeholders to visualize model behavior, perform model debugging,…

Self-recognition -- the ability to maintain an internal representation of one's own body within the environment -- underpins intelligent, autonomous behavior. As a foundational component of the minimal self, self-recognition provides the…

Animals can accomplish many incredible behavioral feats across a wide range of operational environments and scales that current robots struggle to match. One explanation for this performance gap is the extraordinary properties of the…

机器人学 · 计算机科学 2024-08-30 Saul Schaffer , Hima Hrithik Pamu , Victoria A. Webster-Wood

Existing video avatar models can produce fluid human animations, yet they struggle to move beyond mere physical likeness to capture a character's authentic essence. Their motions typically synchronize with low-level cues like audio rhythm,…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Jianwen Jiang , Weihong Zeng , Zerong Zheng , Jiaqi Yang , Chao Liang , Wang Liao , Han Liang , Yuan Zhang , Mingyuan Gao

We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal reasoning. MiMo-VL-7B-RL outperforms Qwen2.5-VL-7B on 35 out of…

Learning robust and scalable visual representations from massive multi-view video data remains a challenge in computer vision and autonomous driving. Existing pre-training methods either rely on expensive supervised learning with 3D…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Jialv Zou , Bencheng Liao , Qian Zhang , Wenyu Liu , Xinggang Wang

Bedside caregivers assess infants' pain at constant intervals by observing specific behavioral and physiological signs of pain. This standard has two main limitations. The first limitation is the intermittent assessment of pain, which might…

计算机视觉与模式识别 · 计算机科学 2019-01-17 Ghada Zamzmi , Dmitry Goldgof , Rangachar Kasturi , Yu Sun , Terri Ashmeade

We present a novel 4.5B parameter small language model that can handle multiple input and output modalities, including text, images, videos, and audio. Despite its small size, the model achieves near state-of-the-art performance on a…

机器学习 · 计算机科学 2024-11-12 Ben Koska , Mojmír Horváth

We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-V2-Flash adopts a hybrid attention architecture that…

计算与语言 · 计算机科学 2026-01-09 Core Team , Bangjun Xiao , Bingquan Xia , Bo Yang , Bofei Gao , Bowen Shen , Chen Zhang , Chenhong He , Chiheng Lou , Fuli Luo , Gang Wang , Gang Xie , Hailin Zhang , Hanglong Lv , Hanyu Li , Heyu Chen , Hongshen Xu , Houbin Zhang , Huaqiu Liu , Jiangshan Duo , Jianyu Wei , Jiebao Xiao , Jinhao Dong , Jun Shi , Junhao Hu , Kainan Bao , Kang Zhou , Lei Li , Liang Zhao , Linghao Zhang , Peidian Li , Qianli Chen , Shaohui Liu , Shihua Yu , Shijie Cao , Shimao Chen , Shouqiu Yu , Shuo Liu , Tianling Zhou , Weijiang Su , Weikun Wang , Wenhan Ma , Xiangwei Deng , Bohan Mao , Bowen Ye , Can Cai , Chenghua Wang , Chengxuan Zhu , Chong Ma , Chun Chen , Chunan Li , Dawei Zhu , Deshan Xiao , Dong Zhang , Duo Zhang , Fangyue Liu , Feiyu Yang , Fengyuan Shi , Guoan Wang , Hao Tian , Hao Wu , Heng Qu , Hongfei Yi , Hongxu An , Hongyi Guan , Xing Zhang , Yifan Song , Yihan Yan , Yihao Zhao , Yingchun Lai , Yizhao Gao , Yu Cheng , Yuanyuan Tian , Yudong Wang , Zhen Tang , Zhengju Tang , Zhengtao Wen , Zhichao Song , Zhixian Zheng , Zihan Jiang , Jian Wen , Jiarui Sun , Jiawei Li , Jinlong Xue , Jun Xia , Kai Fang , Menghang Zhu , Nuo Chen , Qian Tu , Qihao Zhang , Qiying Wang , Rang Li , Rui Ma , Shaolei Zhang , Shengfan Wang , Shicheng Li , Shuhao Gu , Shuhuai Ren , Sirui Deng , Tao Guo , Tianyang Lu , Weiji Zhuang , Weikang Zhang , Weimin Xiong , Wenshan Huang , Wenyu Yang , Xin Zhang , Xing Yong , Xu Wang , Xueyang Xie , Yilin Jiang , Yixin Yang , Yongzhe He , Yu Tu , Yuanliang Dong , Yuchen Liu , Yue Ma , Yue Yu , Yuxing Xiang , Zhaojun Huang , Zhenru Lin , Zhipeng Xu , Zhiyang Chen , Zhonghua Deng , Zihan Zhang , Zihao Yue

Dynamic facial emotion is essential for believable AI-generated avatars, yet most systems remain visually static, limiting their use in simulations like virtual training for investigative interviews with abused children. We present a…

人机交互 · 计算机科学 2025-07-09 Pegah Salehi , Sajad Amouei Sheshkal , Vajira Thambawita , Michael A. Riegler , Pål Halvorsen

By the recent spread of machine learning in the robotics field, a humanoid that can act, perceive, and learn in the real world through contact with the environment needs to be developed. In this study, as one of the choices, we propose a…

Multimodal large language models (MLLMs) have shown remarkable potential in various domains, yet their application in the medical field is hindered by several challenges. General-purpose MLLMs often lack the specialized knowledge required…

人工智能 · 计算机科学 2025-09-29 Guanghao Zhu , Zhitian Hou , Zeyu Liu , Zhijie Sang , Congkai Xie , Hongxia Yang

The remarkable progress of Multimodal Large Language Models (MLLMs) has attracted increasing attention to extend them to physical entities like legged robot. This typically requires MLLMs to not only grasp multimodal understanding…

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture,…

Reconstructing human dynamic vision from brain activity is a challenging task with great scientific significance. Although prior video reconstruction methods have made substantial progress, they still suffer from several limitations,…

计算机视觉与模式识别 · 计算机科学 2025-02-20 Yizhuo Lu , Changde Du , Chong Wang , Xuanliu Zhu , Liuyun Jiang , Xujin Li , Huiguang He

Automated biomechanical testing has great potential for the development of VR applications, as initial insights into user behaviour can be gained in silico early in the design process. In particular, it allows prediction of user movements…