Abstract
This paper proposes an advanced handwritten
document processing system, which is designed to enhance the
handwritten text recognition by jointly using a Convolutional
Neural Networks and Auto-Encoding Transformer models with
BERT-based spelling correction and GECToR-based
grammatical correction, and is further compared to an existing
OCR-based approach for handwritten recognition. Materials
and Methods: The results of this study consist of two groups:
Group 1 consists of an experimental group that used the already-
existing OCR system, which was tested against ten test samples,
averaging about 78% accuracy in recognition. Group 2 presents
the proposed CNN + Auto-Encoding Transformer model
integrated with BERT and GECToR. Testing criteria included
accuracy, error rate, and processing time. Sample size in both
groups was determined through prior studies. It was expected to
have a test power of 80%, with a significance level of 0.05 and a
95% confidence interval. Statistical analysis was done using
SPSS 26.0 software; for comparing two independent samples, the
appropriate statistical tool is the independent samples t-test.
Result: The results show clearly that the proposed approach of
utilizing the CNN Auto Encoding Transformer performs much
better than a traditional approach of utilizing an Optical
Character Recognition (OCR) system. The proposed approach
achieves high accuracy of 93.6%, reduces errors to 0.11, and also
speeds up recognizing handwritten text in less than 0.58 seconds
while maintaining statistically significant improvements to
confidence level of (p < 0.001). Conclusion: Overall, the
combined effect of the CNN + Auto-Encoding Transformer with
the use of BERT and GECToR makes the recognition of the
handwritten text much more precise and even easier to read. The
system is faster compared to the traditional OCR methods.
info
Full Text Preview
The full text of this article is currently available via the PDF download. We are working on bringing full HTML accessibility to all our research articles.
picture_as_pdfView Full Manuscript