PatentME: A Dataset and Reference-Free Post-OCR Verification Task for Printed Mathematical Expression Recognition

Published in International Conference on Document Analysis and Recognition (ICDAR 2026), 2026

This paper introduces PatentME, a dataset and reference-free post-OCR verification task for printed mathematical expression recognition. The work addresses the difficulty of evaluating mathematical expression recognition systems from scanned patent documents, where noisy inputs, multiple font styles, and missing reference markup make evaluation incomplete.

The proposed benchmark includes two datasets:

  • PatentME-OCR, with roughly 41k images of mathematical expressions extracted from patent documents and paired with MathML annotations.
  • PatentME-Siamese, built from OCR erroneous predictions and designed for a post-OCR verification task where a judge model determines whether two rendered expressions are semantically equivalent.

Read on HAL

Recommended citation: François Wieckowiak, Véronique Eglin, Tony Bonnet, Stéphane Bres, and Laetitia Rousseau. (2026). "PatentME: A Dataset and Reference-Free Post-OCR Verification Task for Printed Mathematical Expression Recognition." International Conference on Document Analysis and Recognition (ICDAR 2026).
Download Paper