中文
相关论文

相关论文: OmniFysics: Towards Physical Intelligence Evolutio…

200 篇论文

Understanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act…

We introduce OmniSource, a novel framework for leveraging web data to train video recognition models. OmniSource overcomes the barriers between data formats, such as images, short videos, and long untrimmed videos for webly-supervised…

计算机视觉与模式识别 · 计算机科学 2020-08-26 Haodong Duan , Yue Zhao , Yuanjun Xiong , Wentao Liu , Dahua Lin

The convergence of cross-modal adversarial learning and physics-driven methods represents a cutting-edge direction for tackling challenges in complex multi-modal tasks and scientific computing. This review focuses on systematically…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Hana Satou , Alan Mitkiy

The diversity and complementarity of sensors available for Earth Observations (EO) calls for developing bespoke self-supervised multimodal learning approaches. However, current multimodal EO datasets and models typically focus on a single…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Guillaume Astruc , Nicolas Gonthier , Clement Mallet , Loic Landrieu

Multimodal physiological signals, such as EEG, ECG, EOG, and EMG, are crucial for healthcare and brain-computer interfaces. While existing methods rely on specialized architectures and dataset-specific fusion strategies, they struggle to…

信号处理 · 电气工程与系统科学 2026-03-18 Wei-Bang Jiang , Xi Fu , Yi Ding , Cuntai Guan

Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-modal large language models. Existing benchmarks fall short in…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Ruixiang Zhao , Jie Yang , Zijie Xin , Tianyi Wang , Fengyun Rao , Jing LYU , Xirong Li

The field of 4D world modeling - aiming to jointly capture spatial geometry and temporal dynamics - has witnessed remarkable progress in recent years, driven by advances in large-scale generative models and multimodal learning. However, the…

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specific encoders with a \emph{video-coarse, audio-dense} design --…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Detao Bai , Shimin Yao , Weixuan Chen , Chengen Lai , Yuanming Li , Zhiheng Ma , Xihan Wei

Autonomous driving has seen remarkable advancements, largely driven by extensive real-world data collection. However, acquiring diverse and corner-case data remains costly and inefficient. Generative models have emerged as a promising…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Tao Tang , Enhui Ma , xia zhou , Letian Wang , Tianyi Yan , Xueyang Zhang , Kun Zhan , Peng Jia , XianPeng Lang , Jia-Wang Bian , Kaicheng Yu , Xiaodan Liang

We are perceiving and communicating with the world in a multisensory manner, where different information sources are sophisticatedly processed and interpreted by separate parts of the human brain to constitute a complex, yet harmonious and…

计算机视觉与模式识别 · 计算机科学 2024-06-12 Ye Zhu , Yu Wu , Nicu Sebe , Yan Yan

Scalable and generalizable physics-aware deep learning has long been considered a significant challenge with various applications across diverse domains ranging from robotics to molecular dynamics. Central to almost all physical systems are…

Learning motion tracking from rich human motion data is a foundational task for achieving general control in humanoid robots, enabling them to perform diverse behaviors. However, discrepancies in morphology and dynamics between humans and…

机器人学 · 计算机科学 2026-03-02 Yuhan Li , Peiyuan Zhi , Yunshen Wang , Tengyu Liu , Sixu Yan , Wenyu Liu , Xinggang Wang , Baoxiong Jia , Siyuan Huang

Camera control, which achieves diverse visual effects by changing camera position and pose, has attracted widespread attention. However, existing methods face challenges such as complex interaction and limited control capabilities. To…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Xiaoda Yang , Jiayang Xu , Kaixuan Luan , Xinyu Zhan , Hongshun Qiu , Shijun Shi , Hao Li , Shuai Yang , Li Zhang , Checheng Yu , Cewu Lu , Lixin Yang

Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision. Whether similar principles apply to modeling brain activity remains unclear. Here we leveraged a dataset of 3.1 million neurons…

Physical neural networks are artificial neural networks that mimic synapses and neurons using physical systems or materials. These networks harness the distinctive characteristics of physical systems to carry out computations effectively,…

应用物理 · 物理学 2024-08-13 Weichao Yu , Hangwen Guo , Jiang Xiao , Jian Shen

Data centers (DCs) as mission-critical infrastructures are pivotal in powering the growth of artificial intelligence (AI) and the digital economy. The evolution from Internet DC to AI DC has introduced new challenges in operating and…

人工智能 · 计算机科学 2025-04-16 Zhiwei Cao , Minghao Li , Feng Lin , Jimin Jia , Yonggang Wen , Jianxiong Yin , Simon See

We study the problem of learning physical object representations for robot manipulation. Understanding object physics is critical for successful object manipulation, but also challenging because physical object properties can rarely be…

机器人学 · 计算机科学 2019-06-13 Zhenjia Xu , Jiajun Wu , Andy Zeng , Joshua B. Tenenbaum , Shuran Song

Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are constrained by their single-modality scope, simplified manipulations, or unrealistic…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Tianxiao Li , Zhenglin Huang , Haiquan Wen , Yiwei He , Xinze Li , Bingyu Zhu , Wuhui Duan , Congang Chen , Zeyu Fu , Yi Dong , Baoyuan Wu , Jason Li , Guangliang Cheng

We present OmniBooth, an image generation framework that enables spatial control with instance-level multi-modal customization. For all instances, the multimodal instruction can be described through text prompts or image references. Given a…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Leheng Li , Weichao Qiu , Xu Yan , Jing He , Kaiqiang Zhou , Yingjie Cai , Qing Lian , Bingbing Liu , Ying-Cong Chen

Whole-body humanoid teleoperation enables humans to remotely control humanoid robots, serving as both a real-time operational tool and a scalable engine for collecting demonstrations for autonomous learning. Despite recent advances,…

机器人学 · 计算机科学 2026-03-17 Yixuan Li , Le Ma , Yutang Lin , Yushi Du , Mengya Liu , Kaizhe Hu , Jieming Cui , Yixin Zhu , Wei Liang , Baoxiong Jia , Siyuan Huang