Extraction and segmentation of tables from Chinese ink documents based on a matrix model

作者:

Highlights:

摘要

This paper presents an approach for extracting and segmenting tables from Chinese ink documents based on a matrix model. An ink document is first modeled as a matrix containing ink rows, including writing and drawing ones. Each row consists of collinear ink lines containing ink characters. Together with their associated drawing rows, adjacent writing rows having an identical distribution of writing lines and⧹or the same associated drawing rows if available are extracted to form a table. Row and column headers, nested sub-headers and cells are identified. Experiments demonstrate that the proposed approach is more effective and robust.

论文关键词:Chinese ink document,Digital ink,Handwriting,Table extraction,Table segmentation

论文评审过程:Received 10 October 2005, Revised 12 May 2006, Accepted 25 May 2006, Available online 30 March 2007.

论文官网地址:https://doi.org/10.1016/j.patcog.2006.05.029