International Conference on Document Analysis and Recognition (ICDAR2015)

ICDAR 2015  Nancy, France  -   August 23-26, 2015

Hervé Dejean : Extracting Structured Data from Unstructured Documents with Incomplete Resources

Abstract: We present a method for extracting structured elements of information, called structured data (sdata), from ocr’ed pages. The method first analyzes the layout of the page, building several concurrent layout structures. Then a tagging step is performed in order to tag textual elements based on their content. Combining the layout structures and the tagged elements, layout models for representing the structured data are inferred for the current page. These models are used to correct or tag some elements missed by the tagging step. The final set of structured data is extracted. An evaluation is presented.Abstract:

Published on : Tuesday 11 August 2015