Comparative Study on Automated Reference Summary Generation using BERT Models and ROUGE Score Assessment
Abstract
Automatic text summarization is a sub-area of text mining in which a computer system determines the most informative content in the original text to produce a summary for a given task and users. In developing such systems, one of the most important tasks is evaluating the quality of the summaries produced, which is generally laborious, time-consuming, and expensive because it requires significant human annotation effort to manually create reference summaries. Being able to generate automatic reference summaries would speed up the development and evaluation of summarization systems. In this paper, we propose an Auto-Ref Summary Generation framework for automatically generating reference summaries used in a generic text-summarization evaluation task, the Sliced Summary. Given a set of clusters from a ground-truth label dataset, variants of BERT models were used to create cluster representations, and automatic reference summaries were generated through a centroid-based summarization approach. DistilBERT, RoBERTa, and SBERT played crucial roles in the automatic summary generation, achieving a highest ROUGE-1 score of 0.47060, though this does not yet meet expectations on text coherence and readability. While the generated summaries could not replace manually written summaries, this study sheds new light on the acquisition of automatic reference summaries from a ground-truth label dataset.
Keywords: automatic summarization, document clustering, k-means, centroid-based summarization, natural language processing, text mining
References
Download paper (PDF) Download BibTeX
BibTeX entry@article{2024_Sanchan,
title = {Comparative Study on Automated Reference Summary Generation using BERT Models and ROUGE Score Assessment},
author = {Sanchan, Nattapong},
journal = {Journal of Current Science and Technology},
volume = {14},
number = {2},
year = {2024},
month = {May},
doi = {10.59796/jcst.V14N2.2024.26}
}Rich-text citation (copy & paste)Sanchan, N. (2024). Comparative Study on Automated Reference Summary Generation using BERT Models and ROUGE Score Assessment. Journal of Current Science and Technology, 14(2). https://doi.org/10.59796/jcst.V14N2.2024.26