An Empirical Benchmarking of Traditional Machine Learning and DistilBERT-Based Zero-Shot Hierarchical Sentiment Analysis on Large-Scale Twitter Data

Abstract

Text sentiment analysis of the social media text faces challenges posed by unstructured data and labori- ous human labeling for intent-driven, hierarchical classification. This work compares conventional ML models (SVM, Naïve Bayes, Logistic Regression) with contextual DL models (DistilBERT) in terms of their performance on Sentiment140 dataset (1.6 million tweets) where a balanced 300,000 tweets were selected (200,000 training, 50,000 validation and 50,000 test). As a solution to the bottleneck of human labeling for detailed topic classification, a zero-shot classification pipeline that uses Natural Language Inference (NLI) for classification of 1,000 tweets into a two-layer taxonomy of 30 parent topics and 330 subtopics has been created without any human-labeled samples. SVM is able to achieve 81.52% accuracy, while DistilBERT scores 84.44% on 50,000 tweet test set and 83.0% on a small 1,000 tweets sample.

References

I. Goodfellow, Y. Bengio, A. Courville (2023) Deep Learning (Updated Edition), MIT Press.

Y. Zhang (2026) Comparative evaluation of deep neural networks and classical ML for sentiment analysis, Informatica, pp. 305–320. https://doi.org/10.31449/inf.v50i1.9539

H. Liu (2025) Multi-task BERT-BiLSTM-CNN framework for sentiment analysis, Informatica, pp. 1–15. https://doi.org/10.31449/inf.v49i5.9541

M. Li (2025) BERT-based consumer sentiment analysis for marketing, Informatica, pp. 1–12. https://doi.org/10.31449/inf.v49i21.8232

L. Ye (2025) Multi-modal fusion for sentiment analysis, Informatica, pp. 111–122. https://doi.org/10.31449/inf.v49i24.8315

M. Hossain et al. (2025) Transformer-based sentiment and emotion analysis on Twitter, Informatica, pp. 1–10. https://doi.org/10.31449/inf.v49i18.6499

A. Chauhan et al. (2025) RoBERTa-based aspect sentiment analysis model, Informatica, pp. 193–202. https://doi.org/10.31449/inf.v49i14.5423

K. Yang (2025) Tourism sentiment analysis, Informatica, pp. 19–32. https://doi.org/10.31449/inf.v49i24.8119

B. Madhurika et al. (2026) Optimized RNN-based sentiment classification model, Informatica, pp. 77–98.

V. Tandon (2023) Integrated approach for sentiment analysis on social media, Informatica, pp. 1–15.

B. Panchal (2025) Context-aware NLP applications for sentiment systems, Informatica, pp. 1–12.

J. Novak, M. Kranjc (2024) Hybrid neural models for sentiment classification, Informatica, pp. 25–40.

A. Zupan, B. Brumen (2023) Machine learning approaches in text mining, Informatica, pp. 41–55.

R. Kenda, D. Mladenić (2024) Transformer models in NLP applications, Informatica, pp. 56–70.

S. Džeroski (2023) Data mining for social media analytics, Informatica, pp. 71–85.

N. Lavrač (2023) Machine learning techniques for text classification, Informatica, pp. 86–100.

M. Robnik-Šikonja (2024) Explainable AI in NLP and sentiment tasks, Informatica, pp. 101–115.

K. Grčar (2023) Feature engineering and vectorization in NLP, Informatica, pp. 116–130.

B. Sluban (2024) Network-based sentiment classification, Informatica, pp. 131–145.

A. Žnidaršič (2023) Hybrid classification systems for NLP, Informatica, pp. 146–160.

T. Curk (2024) Feature extraction for large-scale text datasets, Informatica, pp. 161–175.

J. Brank (2024) Zero-shot learning for hierarchical text classification, Informatica, pp. 176–190.

M. Možina (2023) Evaluation metrics in machine learning systems, Informatica, pp. 191– 205.

Z. Chai et al. (2025) Hybrid BERT-based sentiment analysis on Twitter, Scientific Reports, pp. 1–12. https://doi.org/10.1038/s41598-025-32399-8

N. Jonnala et al. (2025) Hybrid model for sentiment classification, Scientific Reports, pp. 1–10. https://doi.org/10.1038/s41598-025-09794-2

U. Ishfaq et al. (2025) Sentiment analysis using network features, Social Network Analysis and Mining, pp. 1–15. https://doi.org/10.1007/s13278-024-01396-6

A. Gopalakrishnan et al. (2024) BERT-based classification systems, Neural Computing and Applications, pp. 112–128. https://doi.org/10.1007/s00521-023-09123-y

P. Guleria (2025) Clinical sentiment analysis using transformers, Neural Computing and Applications, pp. 341–366. https://doi.org/10.1007/s00521-024-10482-x

S. Aljameel et al. (2025) CNN-LSTM sentiment model, Scientific Reports, pp. 1–10. https://doi.org/10.1038/s41598-024-81234-x

R. Kumar, A. Singh (2025) Context-aware sentiment analysis using deep models, Scientific Reports, pp. 1–11. https://doi.org/10.1038/s41598-025-22929-w

W. Xu et al. (2024) Contrastive learning for hierarchical classification, Applied Intelligence, pp. 1–15. https://doi.org/10.1007/s10489-023-05123-x

M. Xia et al. (2025) Zero-shot hierarchical classification with LLMs, EMNLP Proceedings, ACL, USA, pp. 18189–18208. https://doi.org/10.18653/v1/2025.emnlp-main.1001

Q. Zhang et al. (2025) Hierarchical prompting for classification, Findings of ACL, ACL, USA, pp. 3846–3859. https://doi.org/10.18653/v1/2025.findings-emnlp.1002

L. Paletto et al. (2024) Label augmentation for zero-shot classification, ACL Proceedings, ACL, USA, pp. 7697–7706. https://doi.org/10.18653/v1/2024.acl-long.415

W. Zhang et al. (2023) Sentiment analysis in era of LLMs, arXiv / NLP Conference, USA, pp. 1–15.

Authors

  • Bhumit Peshavariya Vivekanand Education Society's Institute of Technology, Chembur
  • Sunny Nahar Vivekanand Education Society's Institute of Technology, Chembur

DOI:

https://doi.org/10.31449/inf.v50i15.14996

Keywords:

BERT, Sentiment Analysis, Zero-Shot Classification, Deep Learning, Machine learning, Twitter Data Mining, Hierarchal Text Classification

Downloads

Published

09/09/2026

How to Cite

Peshavariya, B., & Nahar, S. (2026). An Empirical Benchmarking of Traditional Machine Learning and DistilBERT-Based Zero-Shot Hierarchical Sentiment Analysis on Large-Scale Twitter Data. Informatica, 50(15). https://doi.org/10.31449/inf.v50i15.14996