A Multimodal AI Framework for Automated Grading of Digital and Text-Based Handwritten Exams.

Show simple item record

dc.contributor.author Virajani, M.Y.
dc.contributor.author Udayanthika, K.D.S.
dc.contributor.author Wahalathanthri, W.A.S.
dc.contributor.author Poornima, K.N.
dc.contributor.author Sandamali, G.G.N.
dc.date.accessioned 2026-09-08T04:41:38Z
dc.date.available 2026-09-08T04:41:38Z
dc.date.issued 2026-03-04
dc.identifier.citation Virajani, M. Y., Udayanthika, K. D. S., Wahalathanthri, W. A. S., Poornima, K. N. & Sandamali, G. G. N. (2026). A Multimodal AI Framework for Automated Grading of Digital and Text-Based Handwritten Exams. 23rd Academic Sessions & Vice – Chancellor’s Awards, Faculty of Engineering, University of Ruhuna, Sri Lanka. 89. en_US
dc.identifier.issn 2362-0412
dc.identifier.uri http://ir.lib.ruh.ac.lk/handle/iruor/21733
dc.description.abstract Manual examination grading is time-consuming, inconsistent, and difficult to scale. This study presents a multimodal Intelligent Exam Evaluation System for automated grading of digital and handwritten scripts across various question types, including short answers, lists, essays, equations, diagrams, graphs, and multiple-choice questions (MCQs). The framework integrates Transformer-based OCR (TrOCR), Retrieval-Augmented Generation, and multiple large language models in a context-grounded evaluation pipeline. Handwritten text is extracted using TrOCR, achieving 93.24% character-level and 76.59% word-level accuracy. Grading is based on rubric-aligned semantic evaluation, grounded in lecture materials and model answers stored as vector embeddings rather than exact word matching. Text-based grading was evaluated using OpenAI GPT, Google Gemini, Anthropic Claude, DeepSeek, and a fine-tuned DeepSeek-R1 model. Zeroshot and few-shot prompting were applied strictly as inference-time strategies to guide rubric-aligned evaluation. Fine-tuning was performed through continued pre-training on domain-specific academic text in structured JSON format to enhance subject-domain understanding. Evaluation included 500 text answers, 350 equation responses, 250 diagram answers, 250 graph answers, and 200 handwritten responses. Temperature sensitivity analysis at 0.0 and 0.2 showed that few-shot prompting at 0.0 provided the most stable and reproducible grading, suitable for high-stakes academic assessment. At this setting, the system achieved the highest matching accuracy of 72% for short answers using Gemini, 80% for list-type using OpenAI, 60% for essays using Claude, and 63.43% for equations using OpenAI. Multimodal evaluation achieved a strong correlation with human scores, with Pearson’s correlation coefficient of 0.784 for diagrams using Claude and 0.787 for graphs using Gemini. Compared to prior automated grading systems focused primarily on MCQs or text-only responses, the proposed framework supports handwritten text, equations, and visual content within a single scalable pipeline. en_US
dc.language.iso en en_US
dc.publisher Faculty of Engineering , University of Ruhuna, Sri Lanka. en_US
dc.subject Automated Grading en_US
dc.subject Handwritten Text Recognition en_US
dc.subject Large Language Models en_US
dc.subject Multimodal Evaluation en_US
dc.subject Retrieval-Augmented Generation en_US
dc.title A Multimodal AI Framework for Automated Grading of Digital and Text-Based Handwritten Exams. en_US
dc.type Article en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record

Search DSpace


Browse

My Account