Mitigating Factual Hallucination in Low-Resource Indic Summarization: A Generate–Verify–Correct Approach for Telugu

Authors

  • Dr. Sheeja Sudheer
  • Boina Uma Devi

Abstract

While the ability of Large Language Models (LLMs) to produce abstractive summaries has proven to be very good, they do have a tendency to hallucinate facts that are missing or incorrect from the input text. This issue is not well studied for low-resource Indian languages like Telugu where standard factual-consistency measures (like FactCC, QAGS) from English do not map well because of Telugu's morphologically complex, agglutinative structure. In this paper, the focus is laid on the factual consistency in Telugu abstractive summarization system using text extracted from the 10th class Telugu medium Social History textbook (AP/Telangana SCERT). We build a dataset that consists of 183 paragraph-level text units from two chapters titled "The Rise of Nationalism in Europe" and "Nationalism in India", and evaluate in detail 28 representative text units. We automatically generate summaries with our fact-verification pipeline and these are compared with baseline (plain-prompted) summaries as well as with manually produced factual-accuracy evaluations. Results of our analysis indicate that 93% of the baseline summaries include at least one factual omission or error: the most frequent omission or error is a missing named entity (32%) and a missing date or numeric information (21%). The fact-correction pipeline removes all the errors detected and corrected and enhances the ROUGE-1 score from 0.199 to 0.500. The findings suggest that factually correcting Telugu summarization systems via verification is a viable approach for enhancing the factual correctness of such systems, especially in the context of educational content where factual accuracy is paramount.

Downloads

Published

2026-09-14

How to Cite

Sudheer, D. S., & Devi, B. U. (2026). Mitigating Factual Hallucination in Low-Resource Indic Summarization: A Generate–Verify–Correct Approach for Telugu . International Journal of Artificial Intelligence and Machine Learning, 6(10s), 1723–1732. Retrieved from https://mail.svedbergopen.com/index.php/ijaiml/article/view/2004