中文
相关论文

相关论文: Learnable Cost Volume Using the Cayley Representat…

200 篇论文

This paper introduces a novel transformer-based network architecture, FlowFormer, along with the Masked Cost Volume AutoEncoding (MCVA) for pretraining it to tackle the problem of optical flow estimation. FlowFormer tokenizes the 4D…

计算机视觉与模式识别 · 计算机科学 2023-06-12 Zhaoyang Huang , Xiaoyu Shi , Chao Zhang , Qiang Wang , Yijin Li , Hongwei Qin , Jifeng Dai , Xiaogang Wang , Hongsheng Li

Learning-based multi-view stereo (MVS) has by far centered around 3D convolution on cost volumes. Due to the high computation and memory consumption of 3D CNN, the resolution of output depth is often considerably limited. Different from…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Junhua Xi , Yifei Shi , Yijie Wang , Yulan Guo , Kai Xu

Identifiability, or recovery of the true latent representations from which the observed data originates, is de facto a fundamental goal of representation learning. Yet, most deep generative models do not address the question of…

机器学习 · 计算机科学 2020-04-28 Shen Li , Bryan Hooi , Gim Hee Lee

The remarkable natural language understanding, reasoning, and generation capabilities of large language models (LLMs) have made them attractive for application to video understanding, utilizing video tokens as contextual input. However,…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Jiaqi Xu , Cuiling Lan , Wenxuan Xie , Xuejin Chen , Yan Lu

Fine-grained skill representations, commonly referred to as knowledge components (KCs), are fundamental to many approaches in student modeling and learning analytics. However, KC-level correctness labels are rarely available in real-world…

计算与语言 · 计算机科学 2026-03-31 Zhangqi Duan , Arnav Kankaria , Dhruv Kartik , Andrew Lan

Interpretability is crucial for ensuring RL systems align with human values. However, it remains challenging to achieve in complex decision making domains. Existing methods frequently attempt interpretability at the level of fundamental…

机器学习 · 计算机科学 2025-06-03 Anna Soligo , Pietro Ferraro , David Boyle

It is often desirable to be able to recognize when inputs to a recognition function learned in a supervised manner correspond to classes unseen at training time. With this ability, new class labels could be assigned to these inputs by a…

机器学习 · 计算机科学 2017-05-23 Ethan M. Rudd , Lalit P. Jain , Walter J. Scheirer , Terrance E. Boult

Despite the increase in calculation power in the last decades, the estimation of brain connectivity is still a tedious task. The high computational cost of the algorithms escalates with the square of the number of signals evaluated, usually…

信号处理 · 电气工程与系统科学 2018-06-29 Ricardo Bruña , Fernando Maestú , Ernesto Pereda

In spite of considerable progress, computing curvature in Volume of Fluid (VOF) methods continues to be a challenge. The goal is to develop a function or a subroutine that returns the curvature in computational cells containing an interface…

计算物理 · 物理学 2018-11-14 Yinghe Qi , Jiacai Lu , Ruben Scardovelli , Stephane Zaleski , Gretar Tryggvason

Channel capacity bounds are derived for a point-to-point indoor visible light communications (VLC) system with signal-dependent Gaussian noise. Considering both illumination and communication, the non-negative input of VLC is constrained by…

信息论 · 计算机科学 2020-11-03 Jin-Yuan Wang , Xian-Tao Fu , Rong-Rong Lu , Jun-Bo Wang , Min Lin , Julian Cheng

Contrastive Learning (CL) has shown promising performance in collaborative filtering. The key idea is to generate augmentation-invariant embeddings by maximizing the Mutual Information between different augmented views of the same instance.…

信息检索 · 计算机科学 2024-01-01 Huiyuan Chen , Vivian Lai , Hongye Jin , Zhimeng Jiang , Mahashweta Das , Xia Hu

Implicit neural representation (INR) has emerged as a promising solution for encoding volumetric data, offering continuous representations and seamless compatibility with the volume rendering pipeline. However, optimizing an INR network…

计算机视觉与模式识别 · 计算机科学 2025-02-17 Maizhe Yang , Kaiyuan Tang , Chaoli Wang

Recent advances on Multi-modal Large Language Models have demonstrated that high-resolution image input is crucial for model capabilities, especially for fine-grained tasks. However, high-resolution images lead to a quadratic increase in…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Yuke Zhu , Chi Xie , Shuang Liang , Bo Zheng , Sheng Guo

Guided policy search is a popular approach for training controllers for high-dimensional systems, but it has a number of pitfalls. Non-convex trajectory optimization has local minima, and non-uniqueness in the optimal policy itself can mean…

机器人学 · 计算机科学 2018-09-18 Robin Deits , Twan Koolen , Russ Tedrake

Particle Image Velocimetry (PIV) is fundamental to fluid dynamics, yet deep learning applications face significant hurdles. A critical gap exists: the lack of comprehensive evaluation of how diverse optical flow models perform specifically…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Zicheng Lin , Xiaoqiang Li , Yichao Wang , Chuang Zhu

While deep neural networks have succeeded in several visual applications, such as object recognition, detection, and localization, by reaching very high classification accuracies, it is important to note that many real-world applications…

计算机视觉与模式识别 · 计算机科学 2020-10-06 Yu-An Chung , Shao-Wen Yang , Hsuan-Tien Lin

We study the task of extending the large language model (LLM) into a vision-language instruction-following model. This task is crucial but challenging since the LLM is trained on text modality only, making it hard to effectively digest the…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Lizhao Liu , Xinyu Sun , Tianhang Xiang , Zhuangwei Zhuang , Liuren Yin , Mingkui Tan

Multimodal large language models (MLLMs) have recently demonstrated strong capabilities in understanding and generating responses from diverse visual inputs, including high-resolution images and long video sequences. As these models scale…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Junwan Kim , Hyunkyung Bae

Our goal is to estimate the characteristic exponent of the input to a L\'evy-driven storage system from a sample of equispaced workload observations. The estimator relies on an approximate moment equation associated with the…

概率论 · 数学 2024-08-29 Dennis Nieman , Michel Mandjes , Liron Ravner

Transformer-based large language models (LLMs) have demonstrated remarkable potential across a wide range of practical applications. However, long-context inference remains a significant challenge due to the substantial memory requirements…

分布式、并行与集群计算 · 计算机科学 2026-01-09 Bo Jiang , Taolue Yang , Youyuan Liu , Xubin He , Sheng Di , Sian Jin