ISSN : 2663-2187

A Frame for domain based identification and insertion of missing pages using Document Classification

Main Article Content

V.Nirmala , Dr.B.Lavanya
» doi: 10.48047/AFJBS.6.14.2024.4983-4987

Abstract

There are numerous challenges associated with document classification; here we focus on the identification of missing pages in the given set of documents. The missing pages are identified, then, the attempt to classify them to a specific domain, using BOW of the missing pages with the BOW of the domain using classification algorithm is done. We the next phase of the proposed work is to search through all the domain specific documents available to identify the incomplete document. Appending of the identified missing pages into the incomplete document and it is made complete. Even if the inserted original document is complete, an incorrect page arrangement could prevent the file from being comprehensive. So, the next phase of internal sorting of pages is carried out, to create a comprehensive complete document. Missing pages are a common occurrence in the process of classifying documents , which can disrupt the accurate classification of domain-specific information. The missing page results in a loss of specific time and file usage. System errors and manual processing errors often leave many datasets incomplete, leading to confusion during data submission. To address this issue, this study proposes a framework that identifies missing pages within a file, classifies these missing pages within a specified domain using a Mis-BOW Naïve Bayes algorithm and then arranges the pages in sequential order using internal sorting and creates a complete document.

Article Details