English
Related papers

Related papers: OmiEmbed: a unified multi-task deep learning frame…

200 papers

Multimodal Large Language Models demonstrate strong performance on natural image understanding, yet exhibit limited capability in interpreting scientific images, including but not limited to schematic diagrams, experimental…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Haoyi Tao , Chaozheng Huang , Nan Wang , Han Lyu , Linfeng Zhang , Guolin Ke , Xi Fang

Developing large-scale foundational datasets is a critical milestone in advancing artificial intelligence (AI)-driven scientific innovation. However, unlike AI-mature fields such as natural language processing, materials science,…

Chemical Physics · Physics 2025-11-18 Ryo Yoshida , Yoshihiro Hayashi , Hidemine Furuya , Ryohei Hosoya , Kazuyoshi Kaneko , Hiroki Sugisawa , Yu Kaneko , Aiko Takahashi , Yoh Noguchi , Shun Nanjo , Keiko Shinoda , Tomu Hamakawa , Mitsuru Ohno , Takuya Kitamura , Misaki Yonekawa , Stephen Wu , Masato Ohnishi , Chang Liu , Teruki Tsurimoto , Arifin , Araki Wakiuchi , Kohei Noda , Junko Morikawa , Teruaki Hayakawa , Junichiro Shiomi , Masanobu Naito , Kazuya Shiratori , Tomoki Nagai , Norio Tomotsu , Hiroto Inoue , Ryuichi Sakashita , Masashi Ishii , Isao Kuwajima , Kenji Furuichi , Norihiko Hiroi , Yuki Takemoto , Takahiro Ohkuma , Keita Yamamoto , Naoya Kowatari , Masato Suzuki , Naoya Matsumoto , Seiryu Umetani , Hisaki Ikebata , Yasuyuki Shudo , Mayu Nagao , Shinya Kamada , Kazunori Kamio , Taichi Shomura , Kensaku Nakamura , Yudai Iwamizu , Atsutoshi Abe , Koki Yoshitomi , Yuki Horie , Katsuhiko Koike , Koichi Iwakabe , Shinya Gima , Kota Usui , Gikyo Usuki , Takuro Tsutsumi , Keitaro Matsuoka , Kazuki Sada , Masahiro Kitabata , Takuma Kikutsuji , Akitaka Kamauchi , Yusuke Iijima , Tsubasa Suzuki , Takenori Goda , Yuki Takabayashi , Kazuko Imai , Yuji Mochizuki , Hideo Doi , Koji Okuwaki , Hiroya Nitta , Taku Ozawa , Hitoshi Kamijima , Toshiaki Shintani , Takuma Mitamura , Massimiliano Zamengo , Yuitsu Sugami , Seiji Akiyama , Yoshinari Murakami , Atsushi Betto , Naoya Matsuo , Satoru Kagao , Tetsuya Kobayashi , Norie Matsubara , Shosei Kubo , Yuki Ishiyama , Yuri Ichioka , Mamoru Usami , Satoru Yoshizaki , Seigo Mizutani , Yosuke Hanawa , Shogo Kunieda , Mitsuru Yambe , Takeru Nakamura , Hiromori Murashima , Kenji Takahashi , Naoki Wada , Masahiro Kawano , Yosuke Harada , Takehiro Fujita , Erina Fujita , Ryoji Himeno , Hiori Kino , Kenji Fukumizu

Occlusion Boundary Estimation (OBE) identifies boundaries arising from both inter-object occlusions and self-occlusion within individual objects. This task is closely related to Monocular Depth Estimation (MDE), which infers depth from a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Lintao Xu , Yinghao Wang , Chaohui Wang

We propose a unified cross-domain transfer learning framework that leverages knowledge from multiple heterogeneous medical imaging datasets to improve performance across segmentation, classification, and object detection tasks. Our approach…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Ceausescu Ciprian-Mihai , Anghelina Ion-Marian , Alexe Dumitru-Bogdan

Universal multimodal embedding models have achieved great success in capturing semantic relevance between queries and candidates. However, current methods either condense queries and candidates into a single vector, potentially limiting the…

Information Retrieval · Computer Science 2026-04-08 Zilin Xiao , Qi Ma , Mengting Gu , Chun-cheng Jason Chen , Xintao Chen , Vicente Ordonez , Vijai Mohan

Embodied navigation presents a core challenge for intelligent robots, requiring the comprehension of visual environments, natural language instructions, and autonomous exploration. Existing models often fall short in offering a unified…

Robotics · Computer Science 2026-01-08 Xinda Xue , Junjun Hu , Minghua Luo , Shichao Xie , Jintao Chen , Zixun Xie , Kuichen Quan , Wei Guo , Mu Xu , Zedong Chu

Mitotic activity is an important feature for grading several cancer types. Counting mitotic figures (MFs) is a time-consuming, laborious task prone to inter-observer variation. Inaccurate recognition of MFs can lead to incorrect grading and…

Clinical cystoscopy, the current standard for bladder cancer diagnosis, suffers from significant reliance on physician expertise, leading to variability and subjectivity in diagnostic outcomes. There is an urgent need for objective,…

Image and Video Processing · Electrical Eng. & Systems 2025-08-22 Jinliang Yu , Mingduo Xie , Yue Wang , Tianfan Fu , Xianglai Xu , Jiajun Wang

Natural images exhibit label diversity (clean vs. noisy) in noisy-labeled image classification and prevalence diversity (abundant vs. sparse) in long-tailed image classification. Similarly, medical images in universal lesion detection (ULD)…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Han Li , Hu Han , S. Kevin Zhou

Existing machine learning methods for molecular (e.g., gene) embeddings are restricted to specific tasks or data modalities, limiting their effectiveness within narrow domains. As a result, they fail to capture the full breadth of gene…

Manifold learning (ML) aims to seek low-dimensional embedding from high-dimensional data. The problem is challenging on real-world datasets, especially with under-sampling data, and we find that previous methods perform poorly in this case.…

Machine Learning · Computer Science 2022-07-27 Zelin Zang , Siyuan Li , Di Wu , Ge Wang , Lei Shang , Baigui Sun , Hao Li , Stan Z. Li

Missing input sequences are common in medical imaging data, posing a challenge for deep learning models reliant on complete input data. In this work, inspired by MultiMAE [2], we develop a masked autoencoder (MAE) paradigm for multi-modal,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Ayhan Can Erdur , Christian Beischl , Daniel Scholz , Jiazhen Pan , Benedikt Wiestler , Daniel Rueckert , Jan C Peeken

Learning unified text embeddings that excel across diverse downstream tasks is a central goal in representation learning, yet negative transfer remains a persistent obstacle. This challenge is particularly pronounced when jointly training a…

Computation and Language · Computer Science 2025-09-30 Bowen Zhang , Zixin Song , Chunquan Chen , Qian-Wen Zhang , Di Yin , Xing Sun

The integration of AI-assisted biomedical image analysis into clinical practice demands AI-generated findings that are not only accurate but also interpretable to clinicians. However, existing biomedical AI models generally lack the ability…

Data-driven approaches such as deep learning can result in predictive models for material properties with exceptional accuracy and efficiency. However, in many applications, data is sparse, severely limiting their accuracy and…

Machine Learning · Computer Science 2025-10-29 Robert J Appleton , Brian C Barnes , Alejandro Strachan

Deep learning models, such as the fully convolutional network (FCN), have been widely used in 3D biomedical segmentation and achieved state-of-the-art performance. Multiple modalities are often used for disease diagnosis and quantification.…

Image and Video Processing · Electrical Eng. & Systems 2019-08-23 Yu Chen , Jiawei Chen , Dong Wei , Yuexiang Li , Yefeng Zheng

Developing computational tools for integrative analysis across multiple types of omics data has been of immense importance in cancer molecular biology and precision medicine research. While recent advancements have yielded integrative…

The practical deployment of medical vision-language models (Med-VLMs) necessitates seamless integration of textual data with diverse visual modalities, including 2D/3D images and videos, yet existing models typically employ separate…

Computation and Language · Computer Science 2025-04-22 Songtao Jiang , Yuan Wang , Sibo Song , Yan Zhang , Zijie Meng , Bohan Lei , Jian Wu , Jimeng Sun , Zuozhu Liu

Multimodal pretraining is an effective strategy for the trinity of goals of representation learning in autonomous robots: 1) extracting both local and global task progressions; 2) enforcing temporal consistency of visual representation; 3)…

A large-scale labeled dataset is a key factor for the success of supervised deep learning in computer vision. However, a limited number of annotated data is very common, especially in ophthalmic image analysis, since manual annotation is…

Image and Video Processing · Electrical Eng. & Systems 2022-03-15 Zhiyuan Cai , Li Lin , Huaqing He , Xiaoying Tang
‹ Prev 1 4 5 6 7 8 10 Next ›