A Machine Learning Approach for Layout Inference in Spreadsheets

Elvis Koci; Maik Thiele; Oscar Romero; Wolfgang Lehner

Research.Publish.Connect.

*Please fill out at least one Field. *Value must be an number!

Title:
ISBN:
Year:
Acronym:
Subject:

Advanced Search Proceedings Search

If you're looking for an exact phrase use quotation marks on text fields.

*Please fill out at least one Field.

Title:
Author:
Affiliation:
Subject:

Advanced Search Papers Search

If you're looking for an exact phrase use quotation marks on text fields.

*Please fill out at least one Field.

Name:
Affiliation:
Country:
Conference:
Subject:

Advanced Search Authors Search

If you're looking for an exact phrase use quotation marks on text fields.

*Please fill out at least one Field.

Name:
Country:
Subject:

Advanced Search Affiliations Search

If you're looking for an exact phrase use quotation marks on text fields.

Proceedings

Proceedings Search *Please fill out at least one Field. *Value must be an number!

Title:
ISBN:
Year:
Acronym:
Subject:

Advanced Search Proceedings Search

If you're looking for an exact phrase use quotation marks on text fields.

Papers

Papers Search *Please fill out at least one Field.

Title:
Author:
Affiliation:
Subject:

Advanced Search Papers Search

If you're looking for an exact phrase use quotation marks on text fields.

Authors

Authors Search *Please fill out at least one Field.

Name:
Affiliation:
Country:
Conference:
Subject:

Advanced Search Authors Search

If you're looking for an exact phrase use quotation marks on text fields.

Advanced Search

Paper

A Machine Learning Approach for Layout Inference in Spreadsheets

Topics: Information Extraction; Machine Learning; Web Mining

In Proceedings of the 8th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management - Volume 0IC3K, 77-88, 2016 , Porto, Portugal

Authors: Elvis Koci ¹ ; Maik Thiele ¹ ; Oscar Romero ² and Wolfgang Lehner ¹

Affiliations: ¹ Technische Universität Dresden, Germany ; ² Universitat Politecnica de Catalunya (UPC-BarcelonaTech), Spain

Keyword(s): Speadsheets, Tabular, Layout, Structure, Machine Learning, Knowledge Discovery.

Related Ontology Subjects/Areas/Topics: Artificial Intelligence ; Computational Intelligence ; Evolutionary Computing ; Information Extraction ; Knowledge Discovery and Information Retrieval ; Knowledge-Based Systems ; Machine Learning ; Soft Computing ; Symbolic Systems ; Web Mining

Abstract: Spreadsheet applications are one of the most used tools for content generation and presentation in industry and the Web. In spite of this success, there does not exist a comprehensive approach to automatically extract and reuse the richness of data maintained in this format. The biggest obstacle is the lack of awareness about the structure of the data in spreadsheets, which otherwise could provide the means to automatically understand and extract knowledge from these files. In this paper, we propose a classification approach to discover the layout of tables in spreadsheets. Therefore, we focus on the cell level, considering a wide range of features not covered before by related work. We evaluated the performance of our classifiers on a large dataset covering three different corpora from various domains. Finally, our work includes a novel technique for detecting and repairing incorrectly classified cells in a post-processing step. The experimental results show that our approach delive rs very high accuracy bringing us a crucial step closer towards automatic table extraction. (More)

CC BY-NC-ND 4.0

Guest: Register as new SciTePress user now for free.

SciTePress user: please login.

My Papers

You are not signed in, therefore limits apply to your IP address 216.73.216.219

In the current month:

Recent papers: 100 available of 100 total

2⁺ years older papers: 200 available of 200 total

Paper citation in several formats:

Koci, E., Thiele, M., Romero, O. and Lehner, W. (2016). A Machine Learning Approach for Layout Inference in Spreadsheets. In Proceedings of the 8th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K 2016) - KDIR; ISBN 978-989-758-203-5; ISSN 2184-3228, SciTePress, pages 77-88. DOI: 10.5220/0006052200770088

@conference{kdir16,
author={Elvis Koci and Maik Thiele and Oscar Romero and Wolfgang Lehner},
title={A Machine Learning Approach for Layout Inference in Spreadsheets},
booktitle={Proceedings of the 8th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K 2016) - KDIR},
year={2016},
pages={77-88},
publisher={SciTePress},
organization={INSTICC},
doi={10.5220/0006052200770088},
isbn={978-989-758-203-5},
issn={2184-3228},
}

TY - CONF

JO - Proceedings of the 8th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K 2016) - KDIR
TI - A Machine Learning Approach for Layout Inference in Spreadsheets
SN - 978-989-758-203-5
IS - 2184-3228
AU - Koci, E.
AU - Thiele, M.
AU - Romero, O.
AU - Lehner, W.
PY - 2016
SP - 77
EP - 88
DO - 10.5220/0006052200770088
PB - SciTePress