Domain-specific complex nature of patent text, unique drafting styles of patent applicants, and mammoth volume of patent data makes classification a challenging task. To become a helping hand, in the recent time, Google has released pre-trained BERT model trained over 100 million patent documents. However, to the best of our knowledge, there has not been any testament about prediction capabilities and performance of the BERT-for-Patents model over any patent tasks on standard benchmarks. Our work addresses this problem, investigates BERT-for-patents in multi-label patent classification at both CPC and IPC sub-class level. Evidence from experiments enables us to claim that, this work outperformed SOTA by an absolute 2% on micro-F1 with a newly proposed USPTO 2.8M dataset. In order to introduce robustness to the classification process, our collaborative Machine Learning models including NB and SVM uplifted the micro-F1 measures to 70%. This work stands as a corroboration to promote development of patent-specific language models and also claims, robustness in patent analysis tasks can be achieved by not forgetting plain old Machine Learning models. The contributions of this work including code, models, and a novel dataset of the size 2.8M with patent claims are released to the public1, in order to nurture the patent community in developing AI solutions.
Paper
Full text
Patent Classification Using BERT-for-Patents on USPTO
Semantic Scholar · Computer Science · 2022
Abstract
Domain-specific complex nature of patent text, unique drafting styles of patent applicants, and mammoth volume of patent data makes classification a challenging task. To become a helping hand, in the recent time, Google has released pre-trained BERT model trained over 100 million patent documents. However, to the best of our knowledge, there has not been any testament about prediction capabilities and performance of the BERT-for-Patents model over any patent tasks on standard benchmarks. Our work addresses this problem, investigates BERT-for-patents in multi-label patent classification at both CPC and IPC sub-class level. Evidence from experiments enables us to claim that, this work outperformed SOTA by an absolute 2% on micro-F1 with a newly proposed USPTO 2.8M dataset. In order to introduce robustness to the classification process, our collaborative Machine Learning models including NB and SVM uplifted the micro-F1 measures to 70%. This work stands as a corroboration to promote development of patent-specific language models and also claims, robustness in patent analysis tasks can be achieved by not forgetting plain old Machine Learning models. The contributions of this work including code, models, and a novel dataset of the size 2.8M with patent claims are released to the public1, in order to nurture the patent community in developing AI solutions.