Belfort Birth Records Transcription: Preprocessing, and Structured Data Generation

Wissam AlKendi, Franck Gechter, Franck Gechter, Laurent Heyberger, Christophe Guyeux

2024

Abstract

Historical documents are invaluable windows into the past. They play a critical role in shaping our perception of the world and its rich tapestry of stories. This paper presents techniques to facilitate the transcription of the French Belfort Civil Registers of Births, which are valuable historical resources spanning from 1807 to 1919. The methodology focuses on preprocessing steps such as binarization, skew correction, and text line segmentation, tailored to address the challenges posed by these documents including various text styles, marginal annotations, and a hybrid mix of printed and handwritten text. The paper also introduces this archive as a new database by developing a structured strategy for the components of the documents using XML tags, ensuring accurate formatting and alignment of transcriptions with image components at both the paragraph and text line levels for further enhancements to handwritten text recognition models. The results of the preprocessing phase show an accuracy rate of 96%, facilitating the preservation and study of this rich cultural heritage.

Download


Paper Citation


in Harvard Style

AlKendi W., Gechter F., Heyberger L. and Guyeux C. (2024). Belfort Birth Records Transcription: Preprocessing, and Structured Data Generation. In Proceedings of the 4th International Conference on Image Processing and Vision Engineering - Volume 1: IMPROVE; ISBN 978-989-758-693-4, SciTePress, pages 32-43. DOI: 10.5220/0012715600003720


in Bibtex Style

@conference{improve24,
author={Wissam AlKendi and Franck Gechter and Laurent Heyberger and Christophe Guyeux},
title={Belfort Birth Records Transcription: Preprocessing, and Structured Data Generation},
booktitle={Proceedings of the 4th International Conference on Image Processing and Vision Engineering - Volume 1: IMPROVE},
year={2024},
pages={32-43},
publisher={SciTePress},
organization={INSTICC},
doi={10.5220/0012715600003720},
isbn={978-989-758-693-4},
}


in EndNote Style

TY - CONF

JO - Proceedings of the 4th International Conference on Image Processing and Vision Engineering - Volume 1: IMPROVE
TI - Belfort Birth Records Transcription: Preprocessing, and Structured Data Generation
SN - 978-989-758-693-4
AU - AlKendi W.
AU - Gechter F.
AU - Heyberger L.
AU - Guyeux C.
PY - 2024
SP - 32
EP - 43
DO - 10.5220/0012715600003720
PB - SciTePress