English
Related papers

Related papers: An Extensible Multimodal Multi-task Object Dataset…

200 papers

Does a machine learning model actually gain an understanding of the material space? We answer this question in the affirmative on the example of the OptiMate model, a graph attention network trained to predict the optical properties of…

Materials Science · Physics 2026-01-19 Malte Grunert , Max Großmann , Erich Runge

Robotic manipulation and navigation are fundamental capabilities of embodied intelligence, enabling effective robot interactions with the physical world. Achieving these capabilities requires a cohesive understanding of the environment,…

Robotics · Computer Science 2025-11-18 Xiaoshuai Hao , Yingbo Tang , Lingfeng Zhang , Yanbiao Ma , Yunfeng Diao , Ziyu Jia , Wenbo Ding , Hangjun Ye , Long Chen

The ability to recognize various food-items in a generic food plate is a key determinant for an automated diet assessment system. This study motivates the need for automated diet assessment and proposes a framework to achieve this. Within…

Computer Vision and Pattern Recognition · Computer Science 2022-10-26 Rameez Ismail , Zhaorui Yuan

Extreme multi-label classification (XMC) refers to supervised multi-label learning involving hundreds of thousand or even millions of labels. In this paper, we develop a suite of algorithms, called Bonsai, which generalizes the notion of…

Machine Learning · Computer Science 2019-08-13 Sujay Khandagale , Han Xiao , Rohit Babbar

We propose an approach for annotating object classes using free-form text written by undirected and untrained annotators. Free-form labeling is natural for annotators, they intuitively provide very specific and exhaustive labels, and no…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Jordi Pont-Tuset , Michael Gygli , Vittorio Ferrari

This paper explores the usage of multimodal image-to-text models to enhance text-based item retrieval. We propose utilizing pre-trained image captioning and tagging models, such as instructBLIP and CLIP, to generate text-based product…

Information Retrieval · Computer Science 2024-02-14 Jason Tang , Garrin McGoldrick , Marie Al-Ghossein , Ching-Wei Chen

Product embedding serves as a cornerstone for a wide range of applications in eCommerce. The product embedding learned from multiple modalities shows significant improvement over that from a single modality, since different modalities…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Baohao Liao , Michael Kozielski , Sanjika Hewavitharana , Jiangbo Yuan , Shahram Khadivi , Tomer Lancewicki

We present the HANDAL dataset for category-level object pose estimation and affordance prediction. Unlike previous datasets, ours is focused on robotics-ready manipulable objects that are of the proper size and shape for functional grasping…

Robotics · Computer Science 2023-08-04 Andrew Guo , Bowen Wen , Jianhe Yuan , Jonathan Tremblay , Stephen Tyree , Jeffrey Smith , Stan Birchfield

Product attribute values are essential in many e-commerce scenarios, such as customer service robots, product recommendations, and product retrieval. While in the real world, the attribute values of a product are usually incomplete and vary…

Computation and Language · Computer Science 2020-09-16 Tiangang Zhu , Yue Wang , Haoran Li , Youzheng Wu , Xiaodong He , Bowen Zhou

Multispectral oriented object detection faces challenges due to both inter-modal and intra-modal discrepancies. Recent studies often rely on transformer-based models to address these issues and achieve cross-modal fusion detection. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Minghang Zhou , Tianyu Li , Chaofan Qiao , Dongyu Xie , Guoqing Wang , Ningjuan Ruan , Lin Mei , Yang Yang

We present exa-AMD, an open-source, high-performance framework designed for accelerated materials discovery on modern supercomputers. exa-AMD overcomes key computational bottlenecks in large-scale structure prediction through task-based…

Materials Science · Physics 2025-12-11 Weiyi Xia , Maxim Moraru , Ying Wai Li , Cai-Zhuang Wang

Three-dimensional (3D) objects have wide applications. Despite the growing interest in 3D modeling in academia and industries, designing and/or creating 3D objects from scratch remains time-consuming and challenging. With the development of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 XiuYu Zhang , Xiaolei Ye , Jui-Che Chang , Yue Fang

Multimodal electronic health record (EHR) data provide richer, complementary insights into patient health compared to single-modality data. However, effectively integrating diverse data modalities for clinical prediction modeling remains…

We introduce Meta-Album, an image classification meta-dataset designed to facilitate few-shot learning, transfer learning, meta-learning, among other tasks. It includes 40 open datasets, each having at least 20 classes with 40 examples per…

Computer Vision and Pattern Recognition · Computer Science 2023-02-20 Ihsan Ullah , Dustin Carrión-Ojeda , Sergio Escalera , Isabelle Guyon , Mike Huisman , Felix Mohr , Jan N van Rijn , Haozhe Sun , Joaquin Vanschoren , Phan Anh Vu

Human-designed visual manuals are crucial components in shape assembly activities. They provide step-by-step guidance on how we should move and connect different parts in a convenient and physically-realizable way. While there has been an…

Computer Vision and Pattern Recognition · Computer Science 2023-02-06 Ruocheng Wang , Yunzhi Zhang , Jiayuan Mao , Ran Zhang , Chin-Yi Cheng , Jiajun Wu

Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly impractical at scale. Existing evaluation protocols rely on average accuracy over fixed item sets,…

Computation and Language · Computer Science 2026-02-03 Peiyu Li , Xiuxiu Tang , Si Chen , Ying Cheng , Ronald Metoyer , Ting Hua , Nitesh V. Chawla

Current multi-category Multiple Object Tracking (MOT) metrics use class labels to group tracking results for per-class evaluation. Similarly, MOT methods typically only associate objects with the same class predictions. These two prevalent…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Siyuan Li , Martin Danelljan , Henghui Ding , Thomas E. Huang , Fisher Yu

We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufacturing and quality inspection. IMDD-1M contains high-resolution…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 TsaiChing Ni , ZhenQi Chen , YuanFu Yang

To reliably navigate ever-shifting real-world environments, agents must grapple with incomplete knowledge and adapt their behavior through experience. However, current evaluations largely focus on tasks that leave no ambiguity, and do not…

Machine Learning · Computer Science 2025-12-01 Gilbert Yang , Yaqin Chen , Thomson Yen , Hongseok Namkoong

Multimodal learning, a rapidly evolving field in artificial intelligence, seeks to construct more versatile and robust systems by integrating and analyzing diverse types of data, including text, images, audio, and video. Inspired by the…

Artificial Intelligence · Computer Science 2024-12-24 Priyaranjan Pattnayak , Hitesh Laxmichand Patel , Bhargava Kumar , Amit Agarwal , Ishan Banerjee , Srikant Panda , Tejaswini Kumar