ClaimCompare:用于评估新颖度破坏性专利对的数据管道
摘要
专利申请过程中的一个基本步骤是确定是否存在先前的专利会破坏新颖性。这一步骤通常由申请人和审查员执行,以 Assess proposed inventions 的新颖性,其中涉及数百万项年度提交的专利申请。然而,conducting this search 是 时间和劳动密集型的,因为 searchers 必须 navigate complex legal and technical jargon while covering a large amount of legal claims。使用信息检索和机器学习方法的自动化方法来检测 newness-destroying patents presents a promising avenue to streamline this process, yet research focusing on this space remains limited。在本文中,我们引入了一个新的数据管道,ClaimCompare,旨在生成适合训练 IR 和 ML 模型的标记化专利 claim 数据集,以解决 newness-destruction assessment 的挑战。 To the best of our knowledge, ClaimCompare is the first pipeline that can generate multiple novelty destroying patent datasets. 为说明此管道的实际相关性,我们利用它构建了一个由超过 27K 专利组成的数据集,涉及电化学领域:1,045 项 base patents 来自 USPTO,每个 base patent 都与 25 项相关专利相 关联,并根据其对 base patent 的 newness destruction 进行标记。随后,我们进行了初步实验,展示了该数据集在 fine-tuning transformer 模型以识别 newness-destroying patents 方面的效 用, demonstrating 29.2% 和 32.7% 的绝对提升 in MRR 和 P@1。
关键词
引用
@article{arxiv.2407.12193,
title = {ClaimCompare: A Data Pipeline for Evaluation of Novelty Destroying Patent Pairs},
author = {Arav Parikh and Shiri Dori-Hacohen},
journal= {arXiv preprint arXiv:2407.12193},
year = {2024}
}