中文
相关论文

相关论文: TORA: Topological Representation Alignment for 3D …

200 篇论文

Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language learning mitigates this limitation by leveraging a small set of labeled image-text pairs…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Junwon You , Mihyun Jang , Sangwoo Mo , Jae-Hun Jung

We introduce a latent 3D representation that models 3D surfaces as probability density functions in 3D, i.e., p(x,y,z), with flow-matching. Our representation is specifically designed for consumption by machine learning models, offering…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Jen-Hao Rick Chang , Yuyang Wang , Miguel Angel Bautista Martin , Jiatao Gu , Xiaoming Zhao , Josh Susskind , Oncel Tuzel

Model merging offers a scalable alternative to multi-task learning but often yields suboptimal performance on classification tasks. We attribute this degradation to a geometric misalignment between the merged encoder and static…

机器学习 · 计算机科学 2026-02-03 Fanshuang Kong , Richong Zhang , Zhijie Nie , Hang Zhou , Ziqiao Wang , Qiang Sun , Chunming Hu

We present ROCA, a novel end-to-end approach that retrieves and aligns 3D CAD models from a shape database to a single input image. This enables 3D perception of an observed scene from a 2D RGB observation, characterized as a lightweight,…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Can Gümeli , Angela Dai , Matthias Nießner

Modern diffusion models encounter a fundamental trade-off between training efficiency and generation quality. While existing representation alignment methods, such as REPA, accelerate convergence through patch-wise alignment, they often…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Hesen Chen , Junyan Wang , Zhiyu Tan , Hao Li

Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions specified in the prompt. Representation-alignment objectives such as VideoREPA and MoAlign…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Jiesong Lian , Zixiang Zhou , Ruizhe Zhong , Yuan Zhou , Qinglin Lu , Rui Wang , Long Hu , Yixue Hao , Baoru Huang

Reinforcement learning (RL) has become a critical paradigm for LLM post-training, yet the rollout phase -- accounting for 50--80% of total step time -- is bottlenecked by skewed generation: long-tailed trajectories indispensable for model…

Recent healthcare foundation models have achieved strong predictive performance through large scale self supervised learning, yet their latent representations frequently entangle physiologic severity, intervention intensity, observational…

机器学习 · 计算机科学 2026-05-19 Yuanyun Zhang , Shi Li

Neural networks encode inputs as high-dimensional vectors, known as representations, that capture how models process data by encoding task-relevant structure and semantics. Representation alignment refers to the degree to which different…

计算几何 · 计算机科学 2026-05-26 Xinyuan Yan , Rita Sevastjanova , Mennatallah El-Assady , Bei Wang

Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible motion-sound relations remains challenging. Existing methods often produce object motions…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Junchao Liao , Zhenghao Zhang , Xiangyu Meng , Litao Li , Ziying Zhang , Siyu Zhu , Long Qin , Weizhi Wang

Robotic middleware serves as the foundational infrastructure, enabling complex robotic systems to operate in a coordinated and modular manner. In data-intensive robotic applications, especially in industrial scenarios, communication…

机器人学 · 计算机科学 2026-02-17 Xiaodong Zhang , Baorui Lv , Xavier Tao , Xiong Wang , Jie Bao , Yong He , Yue Chen , Zijiang Yang

The majority of existing large 3D shape datasets contain meshes that lend themselves extremely well to visual applications such as rendering, yet tend to be topologically invalid (i.e, contain non-manifold edges and vertices, disconnected…

图形学 · 计算机科学 2023-07-04 Amir Barda , Yotam Erel , Yoni Kasten , Amit H. Bermano

Parametric human body models are foundational to human reconstruction, animation, and simulation, yet they remain mutually incompatible: SMPL, SMPL-X, MHR, Anny, and related models each diverge in mesh topology, skeletal structure, shape…

There is often variation in the shape and size of input data used for deep learning. In many cases, such data can be represented using tensors with non-uniform shapes, or ragged tensors. Due to limited and non-portable support for efficient…

机器学习 · 计算机科学 2022-03-23 Pratik Fegade , Tianqi Chen , Phillip B. Gibbons , Todd C. Mowry

In this paper a new asymmetric 3-translational (3T) parallel manipulator, i.e., RPa(3R) 2R+RPa, with zero coupling degree and decoupled motion is firstly proposed according to topology design theory of parallel mechanism (PM) based on…

机器人学 · 计算机科学 2018-05-24 Huiping Shen , Chengqi Wu , Damien Chablat , Guang Wu , Ting-Li Yang

Visual object recognition in unseen and cluttered indoor environments is a challenging problem for mobile robots. This study presents a 3D shape and color-based descriptor, TOPS2, for point clouds generated from RGB-D images and an…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Ekta U. Samani , Ashis G. Banerjee

Multimodal alignment is commonly learned from isolated image-text pairs via CLIP-style dual encoders, leaving the relational context among entities largely unused. Multimodal attributed graphs (MAGs), where nodes carry multimodal attributes…

机器学习 · 计算机科学 2026-05-18 Xu Wang , Xunkai Li , Yinlin Zhu , Rong-Hua Li , Guoren Wang

Representation learning of geospatial locations remains a core challenge in achieving general geospatial intelligence, with increasingly diverging philosophies and techniques. While Earth observation paradigms excel at depicting locations…

人工智能 · 计算机科学 2026-01-28 Ya Wen , Jixuan Cai , Qiyao Ma , Linyan Li , Xinhua Chen , Chris Webster , Yulun Zhou

Topology optimization is used for the design of high-performance structures but remains fundamentally limited by its iterative nature, requiring repeated finite element analyses that prevent real-time deployment and large-scale design…

计算工程、金融与科学 · 计算机科学 2026-04-07 Aaron Lutheran , Srijan Das , Alireza Tabarraei

Recent advances in Diffusion Transformers (DiTs) demonstrate that aligning noisy latent states with well-trained semantic features-as pioneered by Representation Alignment (REPA)-can substantially accelerate training and improve generation…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Shaodong Xu , Zhendong Wang , Litong Gong , Zexian Li , Wengang Zhou , Tiezheng Ge , Houqiang Li
‹ 上一页 1 2 3 10 下一页 ›