ANALISIS KOMPARATIF KINERJA LOGISTIC REGRESSION DAN RANDOM FOREST BERBASIS TF-IDF UNTUK KLASIFIKASI SENTIMEN KOMENTAR TIKTOK TERHADAP PROGRAM MAKAN BERGIZI GRATIS

  • Febriza Evan Nugraha Universitas Dinamika Bangsa
  • Ibni Faiq Athallah Universitas Dinamika Bangsa
  • Marrylinteri Istoningtyas Universitas Dinamika Bangsa

Abstract

Sentiment analysis of public policy on social media has become increasingly important as public participation in digital spaces continues to grow. This study compares the performance of Logistic Regression and Random Forest algorithms based on TF-IDF for classifying the sentiment of TikTok comments regarding the Free Nutritious Meal Program into three classes: positive, neutral, and negative. The dataset consists of 4,811 public comments collected from six TikTok videos between January and June 2026. After preprocessing and manual labeling, 3,884 valid comments were obtained with a sentiment distribution of 39.73% negative, 37.97% positive, and 22.30% neutral. Class imbalance was addressed using SMOTE on the training data, and the dataset was split using an 80:20 stratified split. Evaluation results show that Logistic Regression outperformed Random Forest across all metrics, achieving an accuracy of 0.76 and a macro F1-score of 0.74 compared to Random Forest's accuracy of 0.73 and macro F1-score of 0.71. In both models, the neutral class consistently showed the lowest performance, indicating semantic ambiguity that cannot be optimally captured by frequency-based feature representations. This study provides empirical evidence that Logistic Regression is more suitable for Indonesian social media text sentiment classification with TF-IDF representation, and recommends exploring context-based models such as IndoBERT for future research.

Downloads

Download data is not yet available.

References

[1] D. A. Sulistyo and E. Setiadi, “Analisis Sentimen Kebijakan Makan Bergizi Gratis Menggunakan IndoBERT dan Machine Learning,” Jurnal FASILKOM (Teknologi Informasi dan Ilmu Komputer), vol. 15, no. 3, pp. 507–513, Dec. 2025, doi: 10.37859/jf.v15i3.10546.
[2] “Digital 2026: Top digital and social media trends in Indonesia,” We Are Social. [Online]. Available: https://wearesocial.com/id/blog/2025/11/digital-2026-top-digital-and-social-media-trends-in-indonesia/
[3] S. Yusra and R. A. Putri, “Klasifikasi Stance Opini Publik Komentar TikTok Kasus Guru Menampar Murid memakai TF-IDF dan SVM,” Jurnal Sistem Komputer dan Informatika (JSON), vol. 7, no. 3, pp. 910–921, 2026, doi: 10.30865/json.v7i3.9544.
[4] A. Sulistyawati and D. W. Prabowo, “Analisis Sentimen Komentar Pengguna Tiktok terhadap Konten Fashion Brand Clotiva dengan Random Forest, Logistic Regression, dan Support Vector Machine,” Journal of Community Service and Research (JURAGAN), vol. 3, no. 1, pp. 2860–2872, Mar. 2026, doi: 10.62710/v6exn887.
[5] “Pemerintah Salurkan Makan Bergizi Gratis (MBG), Ini Sasaran Utama Penerimanya,” Media Keuangan. [Online]. Available: https://mediakeuangan.kemenkeu.go.id/article/show/pemerintah-salurkan-makan-bergizi-gratis-mbg-ini-sasaran-utama-penerimanya
[6] “Anggaran Makan Bergizi Gratis Rp 335 Triliun, untuk Intervensi Gizi hingga Digitalisasi,” Tempo. [Online]. Available: https://www.tempo.co/politik/anggaran-makan-bergizi-gratis-rp-335-triliun-untuk-intervensi-gizi-hingga-digitalisasi-2060952
[7] R. A. Munir, “Analisis Sentimen Cuitan di Media Sosial X tentang Program Makan Bergizi Gratis dengan Metode NLP,” Jurnal Informatika Dan Teknik Elektro Terapan, vol. 13, no. 3, pp. 589–595, Jul. 2025, doi: 10.23960/jitet.v13i3.6912.
[8] M. A. Hasan and N. P. Bimby, “Analisis Sentimen Publik Terhadap Kenaikan Pajak PPN di Indonesia Tahun 2024 Menggunakan Algoritma Machine Learning,” Jurnal FASILKOM (Teknologi Informasi dan Ilmu Komputer), vol. 15, no. 1, pp. 179–184, Apr. 2025, doi: 10.37859/jf.v15i1.8556.
[9] M. Samantri and Afiyati, “Perbandingan Algoritma Support Vector Machine dan Random Forest untuk Analisis Sentimen Terhadap Kebijakan Pemerintah Indonesia Terkait Kenaikan Harga BBM Tahun 2022,” Jurnal Teknologi Informasi Dan Komunikasi (JTIK), vol. 8, no. 1, pp. 1–9, Jan. 2024, doi: 10.35870/jtik.v8i1.1202.
[10] S. Butsianto and A. M. Rifa’i, “Analisis Sentimen Ulasan Aplikasi Jamsostek dengan SVM, Random Forest, dan Logistic Regression,” Jurnal Informatika Ekonomi Bisnis, vol. 7, no. 3, pp. 700–706, Sep. 2025, doi: 10.37034/infeb.v7i3.1266.
[11] L. N. Afifah, S. Rahayu, and Purwadi, “Sentiment Analysis of TikTok User Comments on The Free Nutritious Meal Program Using Support Vector Machine,” Journal of Artificial Intelligence and Engineering Applications (JAIEA), vol. 5, no. 2, pp. 2366–2371, Feb. 2026, doi: 10.59934/jaiea.v5i2.1879.
[12] W. O. Simanjuntak, A. B. P. Negara, and R. Septriana, “Perbandingan Algoritma Logistic Regression dan Random Forest (Studi Kasus : Klasifikasi Emosi Tweet),” Jurnal Aplikasi dan Riset Informatika (JUARA), vol. 2, no. 1, pp. 160–164, Aug. 2023, doi: 10.26418/juara.v2i1.69682.
[13] H. Junianto, R. E. Saputro, B. A. Kusuma, and D. I. S. Saputra, “Comparison of Logistic Regression and Random Forest in Sentiment Analysis of Disdukcapil Application Reviews,” Jurnal Teknik Informatika (JUTIF), vol. 5, no. 6, pp. 1539–1547, Feb. 2024, doi: 10.52436/1.jutif.2024.5.6.1802.
[14] A. M. Wahid, Turino, K. A. Nugroho, T. Safitri, and F. S. Utomo, “Optimasi Logistic Regression dan Random Forest untuk Deteksi Berita Hoax Berbasis TF-IDF,” Jurnal Pendidikan dan Teknologi Indonesia (JPTI), vol. 4, no. 8, pp. 381–392, Aug. 2024, doi: 10.52436/1.jpti.602.
[15] R. Couronné, P. Probst, and A.-L. Boulesteix, “Random forest versus logistic regression: a large-scale benchmark experiment,” BMC Bioinformatics, vol. 19, no. 1, p. 270, Jul. 2018, doi: 10.1186/s12859-018-2264-5.
[16] A. Saepudin, A. Faqih, and G. Dwilestari, “Perbandingan Algoritma Klasifikasi Support Vector Machine, Random Forest dan Logistic Regression Pada Ulasan Shopee,” Jurnal TEKNO KOMPAK, vol. 18, no. 1, pp. 178–192, Feb. 2024, doi: 10.33365/jtk.v18i1.3764.
[17] S. E. Prianto, Berlilana, and R. E. Saputro, “Analisis Sentimen Program Makan Bergizi Gratis Menggunakan Random Forest dan Support Vector Machine dengan Penyeimbangan Data Berbasis SMOTE,” Jurnal Pendidikan dan Teknologi Indonesia, vol. 6, no. 2, pp. 353–365, Mar. 2026.
[18] J. Han, M. Kamber, and J. Pei, Data Mining: Concepts and Techniques, 3rd ed. Morgan Kaufmann, 2012.
[19] B. Liu, Sentiment Analysis and Opinion Mining. Morgan & Claypool Publishers, 2012.
[20] P. A. Permatasari, L. Linawati, and L. Jasa, “Survei Tentang Analisis Sentimen Pada Media Sosial,” Majalah Ilmiah Teknologi Elektro, vol. 20, no. 2, pp. 177–186, Dec. 2021, doi: 10.24843/mite.2021.v20i02.p01.
[21] M. Siino, I. Tinnirello, and M. La Cascia, “Is text preprocessing still worth the time? A comparative survey on the influence of popular preprocessing methods on Transformers and traditional classifiers,” Inf. Syst., vol. 121, Mar. 2024, doi: 10.1016/j.is.2023.102342.
[22] V. W. D. Thomas and F. Rumaisa, “Analisis Sentimen Ulasan Hotel Bahasa Indonesia Menggunakan Support Vector Machine dan TF-IDF,” Jurnal Media Informatika Budidarma, vol. 6, no. 3, pp. 1767–1774, Jul. 2022, doi: 10.30865/mib.v6i3.4218.
[23] B. Ramadhani and Suryono, “Komparasi Algoritma Naïve Bayes dan Logistic Regression Untuk Analisis Sentimen Metaverse,” Jurnal Media Informatika Budidarma, vol. 8, no. 2, pp. 714–725, Apr. 2024, doi: 10.30865/mib.v8i2.7458.
[24] M. U. Albab, Y. K. P., and M. N. Fawaiq, “Optimization of the Stemming Technique on Text Preprocessing President 3 Periods Topic,” Jurnal Transformatika, vol. 20, no. 2, pp. 1–12, Jan. 2023, doi: 10.26623/transformatika.v20i2.5374.
[25] R. Wati and H. Rachmi, “Pembobotan TF-IDF Menggunakan Naïve Bayes Pada Sentimen Masyarakat Mengenai Isu Kenaikan BIPIH,” Jurnal Manajemen Informatika (JAMIKA), vol. 13, no. 1, pp. 84–93, Apr. 2023, doi: 10.34010/jamika.v13i1.9424.
[26] M. H. Mahendra, D. T. Murdiansyah, and K. M. Lhaksamana, “Analisis Sentimen Tweet COVID-19 Menggunakan Metode K-Nearest Neighbors dengan Ekstraksi Fitur TF-IDF dan CountVectorizer,” Jurnal Ilmu Multidisiplin, vol. 1, no. 2, pp. 37–43, Aug. 2023, doi: 10.69688/dike.v1i2.35.
[27] N. V Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic Minority Over-sampling Technique,” Journal of Artificial Intelligence Research, vol. 16, pp. 321–357, 2002.
[28] J. L. Leevy, T. M. Khoshgoftaar, R. A. Bauder, and N. Seliya, “A survey on addressing high-class imbalance in big data,” J. Big Data, vol. 5, no. 1, Dec. 2018, doi: 10.1186/s40537-018-0151-6.
[29] Ash Shiddicky and S. Agustian, “Analisis Sentimen Masyarakat Terhadap Kebijakan Vaksinasi Covid-19 pada Media Sosial Twitter menggunakan Metode Logistic Regression,” Jurnal Computer Science and Information Technology (CoSciTech), vol. 3, no. 2, pp. 99–106, Aug. 2022, doi: 10.37859/coscitech.v3i2.3836.
[30] A. A. Putri, S. Agustian, and R. Abdillah, “Penerapan Metode Logistic Regression untuk Klasifikasi Sentimen pada Dataset Twitter Terbatas,” Jurnal Sistem Informasi (ZONAsi), vol. 7, no. 1, pp. 95–107, Jan. 2025, doi: 10.31849/zn.v7i1.24804.
[31] O. Shobayo, S. Adeyemi-Longe, O. Popoola, and B. Ogunleye, “Innovative Sentiment Analysis and Prediction of Stock Price Using FinBERT, GPT-4 and Logistic Regression: A Data-Driven Approach,” Big Data and Cognitive Computing, vol. 8, no. 11, Nov. 2024, doi: 10.3390/bdcc8110143.
[32] C. M. Bishop, Pattern Recognition and Machine Learning. New York: Springer, 2006.
[33] L. Breiman, “Random Forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001.
Published
2026-07-27
How to Cite
NUGRAHA, Febriza Evan; ATHALLAH, Ibni Faiq; ISTONINGTYAS, Marrylinteri. ANALISIS KOMPARATIF KINERJA LOGISTIC REGRESSION DAN RANDOM FOREST BERBASIS TF-IDF UNTUK KLASIFIKASI SENTIMEN KOMENTAR TIKTOK TERHADAP PROGRAM MAKAN BERGIZI GRATIS. Journal of Information System, Applied, Management, Accounting and Research, [S.l.], v. 10, n. 3, p. 787-798, july 2026. ISSN 2598-8719. Available at: <https://journal.stmikjayakarta.ac.id/index.php/jisamar/article/view/2515>. Date accessed: 28 july 2026. doi: https://doi.org/10.52362/jisamar.v10i3.2515.