Semantic-Aware Hybrid Text Summarization Using Supervised Sentence Scoring and Redundancy Control
Abstract
The rapid growth of digital textual content has intensified the need for automatic text summarization sys- tems that are both effective and reliable. While extractive summarization methods are interpretable and preserve factual content, they often suffer from redundancy and limited coherence. In contrast, abstractive approaches based on large pretrained transformer models improve fluency and readability but are prone to factual inconsistencies and hallucination. To address these limitations, this paper proposes a semantic- aware hybrid text summarization framework that integrates supervised extractive sentence scoring with constrained abstractive generation. The proposed approach employs an interpretable sentence importance model based on lexical, positional, and semantic features, learned using a Gradient Boosting Regressor. Semantic redundancy among candidate sentences is explicitly controlled using Word Mover’s Distance, enabling improved content diversity without increasing training complexity. The selected sentences are subsequently refined using a transformer-based abstractive model to enhance coherence and linguistic quality while preserving factual consistency. The framework is evaluated on a benchmark news summa- rization dataset using ROUGE-1, ROUGE-2, and ROUGE-L metrics. Experimental results demonstrate consistent improvements over a purely extractive baseline and competitive performance compared to repre- sentative extractive, abstractive, and hybrid approaches. Ablation studies further confirm the contribution of supervised sentence scoring, semantic redundancy control, and abstractive refinement to performance stability. Overall, the results indicate that combining interpretable sentence selection with semantic simi- larity modeling within a hybrid architecture provides a balanced and practical solution for automatic text summarization.References
Nenkova, A., McKeown, K.: A survey of text summarization techniques. In: Ag-
garwal, C., Zhai, C. (eds.) Mining Text Data, pp. 43–76. Springer, Boston (2012)
2. Allahyari, M., Pouriyeh, S., Assefi, M., Safaei, S., Trippe, E., Gutierrez, J., Kochut,
K.: Text summarization techniques: A brief survey. Int. J. Adv. Comput. Sci. Appl.
8(10), 397–405 (2017)
3. Radev, D.R.: Introduction to the Special Issue on Summarization. Comput. Lin-
guist. 28(4), 399–408 (2002)
4. Luhn, H.P.: The automatic creation of literature abstracts. IBM J. Res. Dev. 2(2),
159–165 (1958)
5. Edmundson, H.P.: New methods in automatic extracting. J. ACM 16(2), 264–285
(1969)
6. Carbonell, J., Goldstein, J.: The use of MMR, diversity-based reranking. In: Proc.
SIGIR, pp. 335–336 (1998)
7. Vaswani, A. et al.: Attention is all you need. In: NeurIPS, pp. 5998–6008 (2017)
8. See, A., Liu, P.J., Manning, C.D.: Get to the point. In: ACL, pp. 1073–1083 (2017)
9. Lewis, M. et al.: BART. In: ACL, pp. 7871–7880 (2020)
10. Zhang, J. et al.: PEGASUS. In: ICML, pp. 11328–11339 (2020)
11. Maynez, J. et al.: On faithfulness in abstractive summarization. In: ACL, pp. 1906–
1919 (2020)
12. Gehrmann, S. et al.: The hallucinations problem. Commun. ACM 66(2), 38–45
(2023)
13. Cao, Z., Li, W., Li, S., Wei, F.: Ranking sentences for extractive summarization
with reinforcement learning. In: Proc. NAACL-HLT, pp. 1747–1759 (2018)
14. Chen, Y., Bansal, M.: Fast abstractive summarization with reinforce-selected sen-
tence rewriting. In: Proc. ACL, pp. 5195–5205 (2020)
15. Paulus, R., Xiong, C., Socher, R.: Deep reinforced summarization. In: ICLR (2018)
16. Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., Meger, D.: Deep
reinforcement learning that matters. In: Proc. AAAI, vol. 32(1) (2018)
17. Li, C., Zhu, F., Yang, Y., Li, J.: Determinantal point processes for document
summarization. In: Proc. EMNLP, pp. 373–378 (2018)
18. Chen, Y., Bansal, M.: Fast abstractive summarization with reinforce-selected sen-
tence rewriting. In: Proc. ACL, pp. 5195–5205 (2018)
19. Narayan, S., Cohen, S.B., Lapata, M.: Ranking sentences for extractive summariza-
tion using reinforcement learning. In: Proc. NAACL-HLT, pp. 1747–1759 (2018)
12 K. Belila et al.
20. Hsu, W.-T., Lin, C.-K., Lee, M., Min, K., Tang, J., Sun, M.: A unified model for
extractive and abstractive summarization using inconsistency loss. In: Proc. ACL,
pp. 132–141 (2018)
21. Narayan, S., Cohen, S.B., Lapata, M.: Ranking sentences for extractive summariza-
tion using reinforcement learning. In: Proc. NAACL-HLT, pp. 1747–1759 (2018)
22. Pineau, J., Vincent-Lamarre, P., Sinha, K., Larivière, V., Beygelzimer, A., d’Alché-
Buc, F., Fox, E., Larochelle, H.: Improving reproducibility in machine learning
research. J. Mach. Learn. Res. 22(164), 1–20 (2021)
23. Kusner, M. et al.: From word embeddings to document distances. In: ICML, pp.
957–966 (2015)
24. Lin, C.-Y.: ROUGE. In: ACL Workshop, pp. 74–81 (2004)
25. NLP Progress: Summarization. https://nlpprogress.com/english/summarization.html,
accessed 2024
26. Zhou, Q., Yang, N., Wei, F., Zhou, M.: Neural document summarization by jointly
learning to score and select sentences. In: Proc. ACL, pp. 654–663 (2018)
27. Xu, J., Gan, Z., Cheng, Y., Liu, J.: Discourse-aware neural extractive summariza-
tion. In: Proc. ACL, pp. 5021–5031 (2019)
28. Liu, Y., Lapata, M.: Text summarization with pretrained encoders. In: Proc.
EMNLP-IJCNLP, pp. 3730–3740 (2019)
29. Zhong, J., Liu, Y., Chen, D.: Extractive summarization as text matching. In: Proc.
ACL, pp. 6197–6208 (2020)
30. Raffel, C. et al.: Exploring the limits of transfer learning with a unified text-to-text
transformer. J. Mach. Learn. Res. 21(140), 1–67 (2020)
31. Liu, Y., Zhong, M., Lee, S., Liu, J., Zettlemoyer, L.: SimCLS: A simple framework
for contrastive learning of abstractive summarization. In: Proc. ACL, pp. 1062–
1072 (2021)
32. Liu, T. et al.: BRIO: Bringing order to abstractive summarization. In: Proc.
EMNLP, pp. 2890–2901 (2022)
DOI:
https://doi.org/10.31449/inf.v50i15.14051Keywords:
main author is belila khaoula, the main developper is bedida with sekhri thamer under supervision of belila, all authors reading the manuscriptDownloads
Published
Issue
Section
License
Authors retain copyright in their work. By submitting to and publishing with Informatica, authors grant the publisher (Slovene Society Informatika) the non-exclusive right to publish, reproduce, and distribute the article and to identify itself as the original publisher.
All articles are published under the Creative Commons Attribution license CC BY 3.0. Under this license, others may share and adapt the work for any purpose, provided appropriate credit is given and changes (if any) are indicated.
Authors may deposit and share the submitted version, accepted manuscript, and published version, provided the original publication in Informatica is properly cited.







