Domain-Focused Summarization of Polarized Debates
Abstract
Due to the exponential growth of Internet use, textual content is increasingly published in online media. Every day, more news content, blog posts, and scientific articles are published to the online volumes, opening doors for the text summarization research community to conduct research in those areas. Whilst there are freely accessible repositories for such content, online debates, which have recently become popular, have remained largely unexplored. This thesis addresses the challenge of applying text summarization to online debates. We view that the task of summarizing online debates should not only focus on summarization techniques but also look further into presenting the summaries in formats favored by users. In this thesis, we present how a summarization system is developed to generate online debate summaries in accordance with a designed output, called Combination 2 — the combination of two summaries. The primary objective of the first, Chart Summary, is to visualize the debate summary as a bar chart in a high-level view: bars conveying clusters of salient sentences, labels showing short descriptions of the bars, and numbers of salient sentences conveyed on the two opposing sides. The second, Side-By-Side Summary, linked to the Chart Summary, shows a more detailed summary of an online debate related to a bar clicked by a user. The development of the summarization system is divided into three processes. First, we create a gold standard dataset of online debates, annotated subjectively by 5 judges, and develop a summarization system to identify salient sentences; the system outperforms the baseline. Second, we generate Chart Summary from the selected salient sentences using a framework with two branches — term-based clustering with term-based labeling, or X-means based clustering with MI labeling — and find the X-means approach preferable. Finally, we treat the generation of Side-By-Side Summary as a contradiction detection task, creating two debate entailment datasets from the two clustering approaches, annotated with Contradiction and Non-Contradiction relations, and develop a classifier investigating feature combinations that maximize F1 scores.
Keywords: online debate summarization, text summarization, X-means clustering, term-based clustering, ontology, summary design, summary representation, text mining, information extraction, sentence extraction, semantic similarity, inter-annotator agreement
References
Download thesis (PDF) Download BibTeX
BibTeX entry@phdthesis{SanchanBA18,
author = {Sanchan, Nattapong},
title = {Domain-Focused Summarization of Polarized Debates},
school = {The University of Sheffield, UK},
year = {2018},
month = {5},
note = {http://etheses.whiterose.ac.uk/20878/}
}Rich-text citation (copy & paste)Sanchan, N. (2018). Domain-Focused Summarization of Polarized Debates. PhD thesis, The University of Sheffield, UK. http://etheses.whiterose.ac.uk/20878/.
More information
This thesis was completed under the supervision of Prof. Dr. Kalina Bontcheva and Dr. Ahmet Aker.
Related chapters and papers from this thesis:
- Understanding Human Preferences for Summary Designs in Online Debates Domain
- Gold Standard Online Debates Summaries and First Experiments Towards Automatic Summarization of Online Debate Data
- Automatic Summarization of Online Debates
- An Adoption of a Contradiction Detection Task to Assist the Summarization of Online Debates