Doctracer: Intelligent Document Information Extraction and Processing Toolkit for Sri Lankan Extraordinary Gazettes.

Show simple item record

dc.contributor.author Amarasinghe, A.U.
dc.contributor.author De Zoysa, R.N.C.
dc.contributor.author De Silva, P.K.H.D.
dc.contributor.author Ranaweera, K.D.
dc.contributor.author Sudheera, K.L.K.
dc.contributor.author Abeykoon, V.
dc.date.accessioned 2026-09-11T06:56:32Z
dc.date.available 2026-09-11T06:56:32Z
dc.date.issued 2026-03-04
dc.identifier.citation Amarasinghe, A. U., De Zoysa, R. N. C., De Silva, P. K. H. D., Ranaweera, K. D., Sudheera, K. L. K. & Abeykoon, V. (2026). Doctracer: Intelligent Document Information Extraction and Processing Toolkit for Sri Lankan Extraordinary Gazettes. 23rd Academic Sessions & Vice – Chancellor’s Awards, Faculty of Engineering, University of Ruhuna, Sri Lanka. 105. en_US
dc.identifier.issn 2362-0412
dc.identifier.uri http://ir.lib.ruh.ac.lk/handle/iruor/21758
dc.description.abstract Sri Lankan Extraordinary Gazettes serve as the primary legal mechanism for publishing government structures, including ministries, their allocated departments, and the Acts and functions assigned to each ministry. Following national elections and subsequent administrative revisions, these structures are frequently amended through complex legal documents involving transfers of departments, reassignment of laws, and structural reorganizations. Due to the unstructured nature, dense layouts, and heavy numerical formatting of gazette documents, manual analysis and automated extraction remain highly challenging. This research introduces Doctracer, an intelligent document information extraction and processing toolkit designed specifically for Sri Lankan Extraordinary Gazettes. This system contains customized layout aware parsing pipeline specially optimized for these gazette-specific multi column and tabular structures. This pipeline is closely aligned with controlled, few shot learning optimized Large Language Model (LLM) prompting for accurate information extraction. Extracted content is converted into structured format and modeled using a domain-specific graph schema consisting of key entities such as Gazette, Minister, Department, Law, and Function, along with their semantic relationships. Then this structured data is stored in Neo4j Graph Database to preserve temporal dependencies and enable interactive visualization. Doctracer also introduces an automated amendment comparison module that accurately detects insertions, updates, deletions, and renumbering operations between base and amendment gazettes. Comparative experimental evaluation against widely used LLMs demonstrates that Doctracer achieves significantly higher accuracy in both base gazette extraction and amendment change detection tasks. For base gazette processing, Doctracer achieved 95% extraction accuracy, outperforming ChatGPT (79%), DeepSeek (87%), and Gemini (56%), particularly in handling complex multi-column layouts and maintaining structural consistency. In amendment change detection, Doctracer achieved 100% accuracy in identifying all modification operations, matching the performance of Gemini and DeepSeek while ChatGPT showed no successful detections in this task. en_US
dc.language.iso en en_US
dc.publisher Faculty of Engineering , University of Ruhuna, Sri Lanka. en_US
dc.subject Information extraction en_US
dc.subject Knowledge graph en_US
dc.subject Layout-aware parsing en_US
dc.subject Legal document analysis en_US
dc.subject Natural language processing en_US
dc.title Doctracer: Intelligent Document Information Extraction and Processing Toolkit for Sri Lankan Extraordinary Gazettes. en_US
dc.type Article en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record

Search DSpace


Browse

My Account