Belfort Birth Records Transcription: Preprocessing, and Structured Data Generation
Wissam AlKendi, Franck Gechter, Franck Gechter, Laurent Heyberger, Christophe Guyeux
2024
Abstract
Historical documents are invaluable windows into the past. They play a critical role in shaping our perception of the world and its rich tapestry of stories. This paper presents techniques to facilitate the transcription of the French Belfort Civil Registers of Births, which are valuable historical resources spanning from 1807 to 1919. The methodology focuses on preprocessing steps such as binarization, skew correction, and text line segmentation, tailored to address the challenges posed by these documents including various text styles, marginal annotations, and a hybrid mix of printed and handwritten text. The paper also introduces this archive as a new database by developing a structured strategy for the components of the documents using XML tags, ensuring accurate formatting and alignment of transcriptions with image components at both the paragraph and text line levels for further enhancements to handwritten text recognition models. The results of the preprocessing phase show an accuracy rate of 96%, facilitating the preservation and study of this rich cultural heritage.
DownloadPaper Citation
in Harvard Style
AlKendi W., Gechter F., Heyberger L. and Guyeux C. (2024). Belfort Birth Records Transcription: Preprocessing, and Structured Data Generation. In Proceedings of the 4th International Conference on Image Processing and Vision Engineering - Volume 1: IMPROVE; ISBN 978-989-758-693-4, SciTePress, pages 32-43. DOI: 10.5220/0012715600003720
in Bibtex Style
@conference{improve24,
author={Wissam AlKendi and Franck Gechter and Laurent Heyberger and Christophe Guyeux},
title={Belfort Birth Records Transcription: Preprocessing, and Structured Data Generation},
booktitle={Proceedings of the 4th International Conference on Image Processing and Vision Engineering - Volume 1: IMPROVE},
year={2024},
pages={32-43},
publisher={SciTePress},
organization={INSTICC},
doi={10.5220/0012715600003720},
isbn={978-989-758-693-4},
}
in EndNote Style
TY - CONF
JO - Proceedings of the 4th International Conference on Image Processing and Vision Engineering - Volume 1: IMPROVE
TI - Belfort Birth Records Transcription: Preprocessing, and Structured Data Generation
SN - 978-989-758-693-4
AU - AlKendi W.
AU - Gechter F.
AU - Heyberger L.
AU - Guyeux C.
PY - 2024
SP - 32
EP - 43
DO - 10.5220/0012715600003720
PB - SciTePress