English

Can Large Language Models Generate High-quality Patent Claims?

Computation and Language 2025-05-27 v3

Abstract

Large language models (LLMs) have shown exceptional performance across various text generation tasks but remain under-explored in the patent domain, which offers highly structured and precise language. This paper constructs a dataset to investigate the performance of current LLMs in patent claim generation. Our results demonstrate that generating claims based on patent descriptions outperforms previous research relying on abstracts. Interestingly, current patent-specific LLMs perform much worse than state-of-the-art general LLMs, highlighting the necessity for future research on in-domain LLMs. We also find that LLMs can produce high-quality first independent claims, but their performances markedly decrease for subsequent dependent claims. Moreover, fine-tuning can enhance the completeness of inventions' features, conceptual clarity, and feature linkage. Among the tested LLMs, GPT-4 demonstrates the best performance in comprehensive human evaluations by patent experts, with better feature coverage, conceptual clarity, and technical coherence. Despite these capabilities, comprehensive revision and modification are still necessary to pass rigorous patent scrutiny and ensure legal robustness.

Keywords

Cite

@article{arxiv.2406.19465,
  title  = {Can Large Language Models Generate High-quality Patent Claims?},
  author = {Lekang Jiang and Caiqi Zhang and Pascal A Scherz and Stephan Goetz},
  journal= {arXiv preprint arXiv:2406.19465},
  year   = {2025}
}

Comments

Accepted to NAACL 2025. 16 pages, 2 figures, 12 tables

R2 v1 2026-06-28T17:21:53.531Z