Classification of Sentiment in Reddit Forum Comments using Fine-Tuned IndoBERT (Case Study: Free Nutritious Meal Program)
Downloads
This study analyzes public opinion on the Free Nutritious Meals Program (MBG) policy on the Reddit platform using the IndoBERT model. The study uses a dataset of 6,295 Reddit comments collected from 2024 to 2025. The preprocessing stage implemented a comprehensive suite of textual refinement procedures, encompassing normalization of colloquial language, emoji conversion, and translation into Indonesian to establish linguistic consistency across the corpus. Manual annotation was performed with meticulous care on a balanced dataset of 3,000 comments, evenly allocated among positive, negative, and neutral classes. An 80:20 train–test split was applied to enhance the reliability of model training and evaluation. A fine‑tuned IndoBERT model exhibited outstanding performance, attaining 99.00% for accuracy, F1‑score, and precision. Applied to the full dataset, the model predicted neutral sentiment as predominant (59.44%), followed by positive (34.11%) and negative (6.45%) sentiments, a distribution that suggests a measured and reflective public discourse on the topic. Model reliability was further supported by a mean confidence score of 0.8502 and a processing throughput of 137.93 samples per second, indicating strong potential for deployment as an effective tool for near real‑time sentiment monitoring in public policy contexts.
B. Auxier and M. Anderson, “Social Media Use in 2021,” Pew Research Center, 2021. [Online]. Available: www.pewresearch.org.
M. A. Afandi and I. Suri, “Effectiveness of Government Communication in Delivering Public Policy on Social Media,” Jurnal Komunikasi Pemerintah, vol. 9, no. 1, pp. 23–38, 2025.
W. Wang et al., “IndoBERT: Pre-Trained Model for Bahasa Indonesia,” IndoNLP Research Group, 2020. [Online]. Available: https://huggingface.co/indobenchmark.
K. Pham, K. C. Rao Kathala, and S. Palakurthi, “Reddit Sentiment Analysis on the Impact of AI Using VADER, TextBlob, and BERT,” Procedia Computer Science, vol. 258, pp. 85–94, 2025.
B. Boe, “PRAW: The Python Reddit API Wrapper,” GitHub Documentation, 2023. [Online]. Available: https://praw.readthedocs.io/en/stable/index.html.
A. Kiftiyah et al., “Program Makan Bergizi Gratis (MBG) dalam Perspektif Keadilan Sosial dan Dinamika Sosial-Politik,” Pancasila: Jurnal Keindonesiaan, vol. 5, no. 1, pp. 45–62, 2025.
R. A. Munir, “Analisis Sentimen Cuitan di Media Sosial X tentang Program Makan Bergizi Gratis dengan Metode NLP,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 13, no. 1, pp. 123–135, 2025.
Nurjoko and A. Rahardi, “Model Indo-BERT untuk Identifikasi Sentimen Kekerasan Verbal di Twitter,” IJCCS (Indonesian Journal of Computing and Cybernetics Systems), vol. 18, no. 2, pp. 156–168, 2024.
M. R. Nur, Y. Wibisono, and R. Megasari, “Analisis Sentimen dan Pemodelan Topik pada Post tentang Merek Teknologi di X Menggunakan Fine-Tuning IndoBERT dan BERTopic,” Jurnal Komputer Teknologi Informasi Sistem Informasi (JUKTISI), vol. 4, no. 2, pp. 45–58, 2025.
F. A. Wagay and Jahiruddin, “Classification of Mental Illnesses from Reddit Posts Using Sentence-BERT Embeddings and Neural Networks,” Procedia Computer Science, vol. 258, pp. 234–243, 2025.
V. Agustina, A. Herliana, and E. P. Korespondensi, “Analisis Sentimen Publik atas Kebijakan Efisiensi Anggaran 2025 dengan Text Mining dan Natural Language Processing,” Jurnal Sistem Informasi dan Teknologi, vol. 7, no. 2, pp. 78–89, 2025.
N. Proferes et al., “Studying Reddit: A Systematic Overview of Disciplines, Approaches, Methods, and Ethics,” Social Media and Society, vol. 7, no. 2, pp. 1–14, 2021.
D. Prasetia et al., “Analisis Sentimen Pengguna Aplikasi MyBluebird dengan Algoritma Naïve Bayes di Playstore,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 13, no. 2, pp. 145–156, 2025.
Z. Liu et al., “Improving Sentiment Analysis Accuracy with Emoji Embedding,” Journal of Safety Science and Resilience, vol. 2, no. 4, pp. 289–298, 2021.
S. S. Berutu et al., “Data Preprocessing Approach for Machine Learning-Based Sentiment Classification,” Jurnal Infotel, vol. 15, no. 4, pp. 317–325, 2023.
N. Babanejad, A. Agrawal, A. An, and M. Papagelis, “A Comprehensive Analysis of Preprocessing for Word Representation Learning in Affective Tasks,” 2020. [Online]. Available: https://github.com/NastaranBa/.
H. T. Duong and T. A. Nguyen-Thi, “Preprocessing Techniques and Data Augmentation for Sentiment Analysis,” Computational Social Networks, vol. 8, no. 1, pp. 1–15, 2021.
R. Srinivasan and C. N. Subalalitha, “Sentimental Analysis from Imbalanced Code-Mixed Data Using Machine Learning Approaches,” Distributed and Parallel Databases, vol. 41, nos. 1–2, pp. 37–52, 2023.
A. Kunaefi, Z. Abidin, and R. Kusumawati, “Klasifikasi Berita Hoaks Bahasa Indonesia Menggunakan IndoBERT Fine-Tuning dengan Pendekatan Focal Loss pada Data Tidak Seimbang,” JIPI (Jurnal Ilmiah Penelitian dan Pembelajaran Informatika), vol. 10, no. 2, pp. 89–102, 2025.
M. Faza and M. Taufik, “Penerapan Model IndoBERT untuk Deteksi Potensi Sumber Stres dalam Teks Media,” Jurnal Teknologi Informasi dan Ilmu Komputer, vol. 12, no. 3, pp. 567–578, 2025.
N. Ratnaswari, N. C. Wibowo, and D. S. Y. Kartika, “Analisis Sentimen Menggunakan Metode Lexicon-Based dan Support Vector Machine pada Presiden dan Wakil Presiden Indonesia Periode 2024–2029,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 13, no. 1, pp. 67–78, 2025.
A. A. P. Wibowo and N. Hidayat, “Eksplorasi Linguistik Komputasional dalam Analisis Bahasa Alami untuk Mengungkap Evolusi Dialek Digital di Era Media Sosial Global,” Journal of New Trends in Sciences, vol. 1, no. 3, pp. 45–52, 2023.




3.png)