Authors:
Elvis Koci
1
;
Maik Thiele
1
;
Oscar Romero
2
and
Wolfgang Lehner
1
Affiliations:
1
Technische Universität Dresden, Germany
;
2
Universitat Politecnica de Catalunya (UPC-BarcelonaTech), Spain
Keyword(s):
Speadsheets, Tabular, Layout, Structure, Machine Learning, Knowledge Discovery.
Related
Ontology
Subjects/Areas/Topics:
Artificial Intelligence
;
Computational Intelligence
;
Evolutionary Computing
;
Information Extraction
;
Knowledge Discovery and Information Retrieval
;
Knowledge-Based Systems
;
Machine Learning
;
Soft Computing
;
Symbolic Systems
;
Web Mining
Abstract:
Spreadsheet applications are one of the most used tools for content generation and presentation in industry and the Web. In spite of this success, there does not exist a comprehensive approach to automatically extract and reuse the richness of data maintained in this format. The biggest obstacle is the lack of awareness about the structure of the data in spreadsheets, which otherwise could provide the means to automatically understand and extract knowledge from these files. In this paper, we propose a classification approach to discover the layout of tables in spreadsheets. Therefore, we focus on the cell level, considering a wide range of features not covered before by related work. We evaluated the performance of our classifiers on a large dataset covering three different corpora from various domains. Finally, our work includes a novel technique for detecting and repairing incorrectly classified cells in a post-processing step. The experimental results show that our approach delive
rs very high accuracy bringing us a crucial step closer towards automatic table extraction.
(More)