Media Ethics Violation Detection in Sinhala News Headlines Using Natural Language Processing Techniques.

Show simple item record

dc.contributor.author Vihangani, E. W. V
dc.contributor.author Dilhan, A.W.A.T.
dc.date.accessioned 2026-09-25T09:35:49Z
dc.date.available 2026-09-25T09:35:49Z
dc.date.issued 2024-11-01
dc.identifier.citation A en_US
dc.identifier.issn 3021-6834
dc.identifier.uri http://ir.lib.ruh.ac.lk/handle/iruor/21877
dc.description.abstract The Sri Lankan media industry has undergone significant transformations in recent years. These changes are often driven by technological advancements and shifting audience preferences. As a result, upholding media ethics in news reporting has become more challenging. Over the years, the majority of machine learning-based research has been primarily devoted only to detecting hate speech, fake news, and offensive statements in developing semantic analysis systems. To tackle this pressing gap the study will aim to develop an automatic solution to uncover unethical reporting practices by identifying offensive patterns in Sinhala news headlines using natural language processing techniques. It will also focus on developing a solution to prevent its detrimental effects from thwarting them. Solution development involved several steps, and the data collection process was done by gathering over 2500 Sinhala news headlines from main digital media portals. Those data were annotated as "violated" and "non-violated" based on the code of ethics introduced by Verité Research Institute. Data preprocessing steps included tokenization, stop word removal, punctuation removal, spelling correction, numeric values removal, non-Sinhala values removal, and encoding. Two classification algorithms, Support Vector Machine, and Logistic Regression were used to train different models considering five feature extraction techniques. They were TF-IDF, N-Gram, Count Vectorizer, Word2Vec, and FastText. Based on the performance evaluated on the testing data, the combination of Support Vector Machine (SVM) with TF-IDF was chosen as the primary model due to its robust performance and adaptability across the vast data set and complex ethical scenarios in news headlines. This selected model outperformed other approaches achieving an accuracy of 91% on the tested dataset. Media ethics encompass a broad and complex field. This research focuses only on detecting violations in selected media ethics practices, specifically marginalization, prejudiced reporting, reporting on suicides, and cases of women abuse. Future research can expand on this work by addressing other ethical concerns in headlines, adding an automatic suggestion feature to correct violated headlines and extending the analysis to broader news content. en_US
dc.language.iso en en_US
dc.publisher Faculty of Technology, University of Ruhuna, Sri Lanka. en_US
dc.subject Sinhala Natural Language Processing en_US
dc.subject Media Ethic Violation en_US
dc.subject Sinhala News Headlines en_US
dc.subject Violation Detection en_US
dc.subject Supervised Machine Learning en_US
dc.title Media Ethics Violation Detection in Sinhala News Headlines Using Natural Language Processing Techniques. en_US
dc.type Article en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record

Search DSpace


Browse

My Account