中文
相关论文

相关论文: ChatStitch: Visualizing Through Structures via Sur…

200 篇论文

Vehicle-road collaboration is a promising approach for enhancing the safety and efficiency of autonomous driving by extending the intelligence of onboard systems to smart roadside infrastructures. The introduction of digital twins (DTs),…

系统与控制 · 电气工程与系统科学 2024-10-21 Kui Wang , Kazuma Nonomura , Zongdian Li , Tao Yu , Kei Sakaguchi , Omar Hashash , Walid Saad , Changyang She , Yonghui Li

While vision-language models (VLMs) have exhibited multi-turn visual reasoning capabilities, their reasoning trajectories remain relatively shallow and are dominated by a text-centric paradigm, limiting their applicability to complex visual…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Zhiwei Ning , Wenwen Tong , Xiangli Kong , Shengnan Ma , Ziyi Shang , Jingcheng Ni , Tao Hu , Yong Xien Chng , Jixuan Ying , Zehuan Wu , Hanming Deng , Jie Yang , Yuanjie Zheng , Wei Liu , Lewei Lu

Humans make extensive use of vision and touch as complementary senses, with vision providing global information about the scene and touch measuring local information during manipulation without suffering from occlusions. While prior work…

机器人学 · 计算机科学 2023-08-01 Justin Kerr , Huang Huang , Albert Wilcox , Ryan Hoque , Jeffrey Ichnowski , Roberto Calandra , Ken Goldberg

Advancements at the intersection of computer vision and natural language processing are crucial for applications like assistive tech, multimedia querying, and robotics. This dissertation proposes novel architectures to improve intelligent…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Van Quang Nguyen

With the development of advanced communication technology, connected vehicles become increasingly popular in our transportation systems, which can conduct cooperative maneuvers with each other as well as road entities through…

人机交互 · 计算机科学 2020-09-01 Ziran Wang , Kyungtae Han , Prashant Tiwari

Person re-identification (Re-ID) is a crucial task in computer vision, aiming to recognize individuals across non-overlapping camera views. While recent advanced vision-language models (VLMs) excel in logical reasoning and multi-task…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Ke Niu , Haiyang Yu , Mengyang Zhao , Teng Fu , Siyang Yi , Wei Lu , Bin Li , Xuelin Qian , Xiangyang Xue

The concept of large intelligent surface (LIS)-based communication has recently raised research attention, in which a LIS is regarded as an antenna array whose entire surface area can be used for radio signal transmission and reception. To…

信息论 · 计算机科学 2020-07-13 Jide Yuan , Hien Quoc Ngo , Michail Matthaiou

We present \textit{RopStitch}, an unsupervised deep image stitching framework with both robustness and naturalness. To ensure the robustness of \textit{RopStitch}, we propose to incorporate the universal prior of content perception into the…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Lang Nie , Yuan Mei , Kang Liao , Yunqiu Xu , Chunyu Lin , Bin Xiao

Self-supervised Multi-view stereo (MVS) with a pretext task of image reconstruction has achieved significant progress recently. However, previous methods are built upon intuitions, lacking comprehensive explanations about the effectiveness…

计算机视觉与模式识别 · 计算机科学 2021-09-09 Hongbin Xu , Zhipeng Zhou , Yali Wang , Wenxiong Kang , Baigui Sun , Hao Li , Yu Qiao

Dynamic urban environments are often captured by cameras placed at spatially separated locations with little or no view overlap. However, most existing 4D reconstruction methods assume densely overlapping views. When applied to such sparse…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Hina Kogure , Kei Katsumata , Taiki Miyanishi , Komei Sugiura

Semantic scene completion (SSC) has recently gained popularity because it can provide both semantic and geometric information that can be used directly for autonomous vehicle navigation. However, there are still challenges to overcome. SSC…

计算机视觉与模式识别 · 计算机科学 2024-02-08 Yuanfang Zhang , Junxuan Li , Kaiqing Luo , Yiying Yang , Jiayi Han , Nian Liu , Denghui Qin , Peng Han , Chengpei Xu

Retrieving range information in three-dimensional (3D) radio imaging is particularly challenging due to the limited communication bandwidth and pilot resources. To address this issue, we consider a reconfigurable intelligent surface…

信息论 · 计算机科学 2024-03-19 Yixuan Huang , Jie Yang , Chao-Kai Wen , Shi Jin

We introduce a novel robotic system for improving unseen object instance segmentation in the real world by leveraging long-term robot interaction with objects. Previous approaches either grasp or push an object and then obtain the…

Communication enables the expansion of human visual perception beyond the limitations of time and distance, while computational imaging overcomes the constraints of depth and breadth. Although impressive achievements have been witnessed…

图像与视频处理 · 电气工程与系统科学 2024-10-30 Zhenming Yu , Liming Cheng , Hongyu Huang , Wei Zhang , Liang Lin , Kun Xu

Most of the existing multi-modal models, hindered by their incapacity to adeptly manage interleaved image-and-text inputs in multi-image, multi-round dialogues, face substantial constraints in resource allocation for training and data…

计算机视觉与模式识别 · 计算机科学 2023-11-30 Zhewei Yao , Xiaoxia Wu , Conglong Li , Minjia Zhang , Heyang Qin , Olatunji Ruwase , Ammar Ahmad Awan , Samyam Rajbhandari , Yuxiong He

This paper considers a hybrid reconfigurable environment comprising a UAV-mounted reflecting RIS, an outdoor STAR-RIS enabling simultaneous transmission and reflection, and an indoor holographic RIS (H-RIS), jointly enhancing secure…

网络与互联网体系结构 · 计算机科学 2026-02-02 Elhadj Moustapha Diallo , Mamadou Aliou Diallo , Abusaeed B. M. Adam , Muhammad Naeem Shah

Recent work in vision-and-language demonstrates that large-scale pretraining can learn generalizable models that are efficiently transferable to downstream tasks. While this may improve dataset-scale aggregate metrics, analyzing performance…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Eric Slyman , Minsuk Kahng , Stefan Lee

Surround-view perception is increasingly important for robotic navigation and loco-manipulation, especially in human-in-the-loop settings such as teleoperation, data collection, and emergency takeover. However, current robotic visual…

The RGB-D camera maintains a limited range for working and is hard to accurately measure the depth information in a far distance. Besides, the RGB-D camera will easily be influenced by strong lighting and other external factors, which will…

计算机视觉与模式识别 · 计算机科学 2019-01-23 Mingyang Geng , Suning Shang , Bo Ding , Huaimin Wang , Pengfei Zhang , Lei Zhang

We present a framework to translate between 2D image views and 3D object shapes. Recent progress in deep learning enabled us to learn structure-aware representations from a scene. However, the existing literature assumes that pairs of…

计算机视觉与模式识别 · 计算机科学 2021-02-02 Berk Kaya , Radu Timofte