Klasifikasi Sentimen Teks Code-Mixed Indonesia–Inggris Non-Formal Pada X Menggunakan Model Fine-Tuned DistilBERT

  • Syifa Arifah Nurbayani * Mail Program Studi Teknik Informatika, Fakultas Sains dan Teknologi, Universitas Islam Negeri Sunan Gunung Djati Bandung, Indonesia
  • Dian Sa'adillah Maylawati Program Studi Teknik Informatika, Fakultas Sains dan Teknologi, Universitas Islam Negeri Sunan Gunung Djati Bandung, Indonesia, Indonesia
  • Aldy Rialdy Atmadja Program Studi Teknik Informatika, Fakultas Sains dan Teknologi, Universitas Islam Negeri Sunan Gunung Djati Bandung, Indonesia, Indonesia
Keywords: Sentiment Classification; Code-Mixed; DistilBERT; Social Media; Lightweight Transformer

Abstract

The rapid growth of social media has increased the use of Indonesian–English code-mixed language in digital communication, particularly on social media X. The non-formal characteristics of social media text, such as slang, abbreviations, emojis, and language switching within a single sentence, make sentiment analysis more challenging than monolingual text. This study aims to perform sentiment classification on code-mixed text by evaluating the performance of the lightweight Transformer model DistilBERT and comparing it with IndoBERTweet. The study adopts the CRISP-DM methodology, which consists of Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, and Deployment stages. The dataset comprises 1,108 primary data collected from social media X between 2022 and 2026 and 5,048 secondary data obtained from previous research. Three experimental scenarios were applied: translation into Indonesian, translation into English, and raw data without translation. The results show that DistilBERT achieved its best performance on the English translation scenario, with accuracies of 73.66% on the primary dataset and 72.00% on the secondary dataset. Meanwhile, IndoBERTweet obtained the highest performance on the Indonesian translation scenario, achieving an accuracy of 79.01%. These findings indicate that the alignment between the language of the input data and the pre-training characteristics of the model significantly affects sentiment classification performance on non-formal code-mixed text.

References

Nisrina Hanifa Setiono and Yunita Sari, “Exploring the Impact of Back-Translation on BERT’s Performance in Sentiment Analysis of Code-Mixed Language Data ,” vol. 19, Apr. 2025, doi: https://doi.org/10.22146/ijccs.104757.

Cuk Tho, Yaya Heryadi, Iman Herwidiana Kartowisastro, and Widodo Budiharto, “Code-Mixed Sentiment Analysis Indonesian-English Using Transformer Model,” ICIC Express Letters, vol. 17, no. 11, pp. 1295–1302, Feb. 2023, doi: 10.24507/icicel.17.11.1295.

A. F. Hidayatullah, “Code-Mixed Sentiment Analysis on Indonesian-Javanese-English Text Using Transformer Models,” in 2024 8th International Conference on Information Technology, Information Systems and Electrical Engineering (ICITISEE), IEEE, Aug. 2024, pp. 340–345. doi: 10.1109/ICITISEE63424.2024.10730138.

A. F. Hidayatullah, R. A. Apong, D. T. C. Lai, and A. Qazi, “Pre-trained language model for code-mixed text in Indonesian, Javanese, and English using transformer,” Soc. Netw. Anal. Min., vol. 15, no. 1, p. 30, Mar. 2025, doi: 10.1007/s13278-025-01444-9.

W. Mustikawati and Kundharu Saddhono, “Analisis Penggunaan Bahasa Jaksel Pada Video Berjudul ‘Language Barriers, Culture Shock & Tck Identity Crisis Ft. Mella Carli’ Di Youtube: Kajian Sosiolinguistik Campur Kode,” BASA Journal of Language & Literature, vol. 4, no. 2, pp. 89–94, Oct. 2024, doi: https://doi.org/10.33474/basa.v4i2.22031.

G. Singh, “Sentiment Analysis of Code-Mixed Social Media Text (Hinglish),” Feb. 2021, doi: https://doi.org/10.48550/arXiv.2102.12149.

A. K. Jena, A. K. Jena, K. M. Gopal, A. Tripathy, and N. Panda, “Optimized Feature Selection Approach for Semi-Supervised Sentiment Analysis of E-Commerce Feedback,” Journal of Computer Science, vol. 21, no. 2, pp. 363–379, Feb. 2025, doi: 10.3844/jcssp.2025.363.379.

A. Perera and A. Caldera, “Sentiment Analysis of Code-Mixed Text: A Comprehensive Review,” JUCS - Journal of Universal Computer Science, vol. 30, no. 2, pp. 242–261, Feb. 2024, doi: 10.3897/jucs.98708.

Q. Wu, P. Wang, and C. Huang, “MeisterMorxrc at SemEval-2020 Task 9: Fine-Tune Bert and Multitask Learning for Sentiment Analysis of Code-Mixed Tweets,” in Proceedings of the Fourteenth Workshop on Semantic Evaluation, Stroudsburg, PA, USA: International Committee for Computational Linguistics, 2020, pp. 1294–1297. doi: 10.18653/v1/2020.semeval-1.174.

S. Javdan, T. Shangipour ataei, and B. Minaei-Bidgoli, “IUST at SemEval-2020 Task 9: Sentiment Analysis for Code-Mixed Social Media Text Using Deep Neural Networks and Linear Baselines,” in Proceedings of the Fourteenth Workshop on Semantic Evaluation, Stroudsburg, PA, USA: International Committee for Computational Linguistics, 2020, pp. 1270–1275. doi: 10.18653/v1/2020.semeval-1.170.

A. Patil, V. Patwardhan, A. Phaltankar, G. Takawane, and R. Joshi, “Comparative Study of Pre-Trained BERT Models for Code-Mixed Hindi-English Data,” in 2023 IEEE 8th International Conference for Convergence in Technology (I2CT), IEEE, Apr. 2023, pp. 1–7. doi: 10.1109/I2CT57861.2023.10126273.

L. W. Astuti, Y. Sari, and S. -, “Code-Mixed Sentiment Analysis using Transformer for Twitter Social Media Data,” International Journal of Advanced Computer Science and Applications, vol. 14, no. 10, 2023, doi: 10.14569/IJACSA.2023.0141053.

F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization,” Sep. 2021, [Online]. Available: http://arxiv.org/abs/2109.04607

J. Devlin, M.-W. Chang, K. Lee, K. T. Google, and A. I. Language, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.” [Online]. Available: https://github.com/tensorflow/tensor2tensor

V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” Mar. 2020.

Y. Aliyu, A. Sarlan, K. Usman Danyaro, A. S. B. A. Rahman, and M. Abdullahi, “Sentiment Analysis in Low-Resource Settings: A Comprehensive Review of Approaches, Languages, and Data Sources,” IEEE Access, vol. 12, pp. 66883–66909, 2024, doi: 10.1109/ACCESS.2024.3398635.

M. A. Hasanah, S. Soim, and A. S. Handayani, “Implementasi CRISP-DM Model Menggunakan Metode Decision Tree dengan Algoritma CART untuk Prediksi Curah Hujan Berpotensi Banjir,” Journal of Applied Informatics and Computing, vol. 5, no. 2, pp. 103–108, Oct. 2021, doi: 10.30871/jaic.v5i2.3200.

A. D. Wijaya and B. Bram, “A Sociolinguistic Analysis Of Indoglish Phenomenon In South Jakarta,” PROJECT (Professional Journal of English Education), vol. 4, no. 4, p. 672, Jul. 2021, doi: 10.22460/project.v4i4.p672-684.

C. Tho, Y. Heryadi, L. Lukas, and A. Wibowo, “Code-mixed sentiment analysis of Indonesian language and Javanese language using Lexicon based approach,” J. Phys. Conf. Ser., vol. 1869, no. 1, p. 012084, Apr. 2021, doi: 10.1088/1742-6596/1869/1/012084.

D. I. Putri, A. N. Alfian, M. Y. Putra, and P. D. Mulyo, “IndoBERT Model Analysis: Twitter Sentiments on Indonesia’s 2024 Presidential Election,” Journal of Applied Informatics and Computing, vol. 8, no. 1, pp. 7–12, Jul. 2024, doi: 10.30871/jaic.v8i1.7440.

F. Barbieri, L. E. Anke, and J. Camacho-Collados, “XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond,” May 2022, [Online]. Available: http://arxiv.org/abs/2104.12250

H. Adel et al., “Improving Crisis Events Detection Using DistilBERT with Hunger Games Search Algorithm,” Mathematics, vol. 10, no. 3, Feb. 2022, doi: 10.3390/math10030447.

R. A. Casonatto, T. De Pádua Grillo Souza, and A. M. Mariano, “Quality and Risk Management in Data Mining: A CRISP-DM Perspective.,” Procedia Comput. Sci., vol. 242, pp. 161–168, 2024, doi: 10.1016/j.procs.2024.08.257.

J. Khan, K. Ahmad, S. K. Jagatheesaperumal, and K. A. Sohn, “Textual variations in social media text processing applications: challenges, solutions, and trends,” Artif. Intell. Rev., vol. 58, no. 3, Mar. 2025, doi: 10.1007/s10462-024-11071-z.

N. Calderon et al., “Measuring the Robustness of NLP Models to Domain Shifts,” Apr. 2024, [Online]. Available: http://arxiv.org/abs/2306.00168

S. Chen, L. Neves, and T. Solorio, “Mitigating Temporal-Drift: A Simple Approach to Keep NER Models Crisp,” Online Workshop, 2021. [Online]. Available: https://github.com/

A. Bansal and I. Gangwani, “Zero-Training Temporal Drift Detection for Transformer Sentiment Models: A Comprehensive Analysis on Authentic Social Media Streams,” Nov. 2025, [Online]. Available: http://arxiv.org/abs/2512.20631

Dimensions Badge
Published
2026-07-25
Section
Articles