Klasifikasi Sentimen Teks Code-Mixed Indonesia–Inggris Non-Formal Pada X Menggunakan Model Fine-Tuned DistilBERT
Abstract
The rapid growth of social media has increased the use of Indonesian–English code-mixed language in digital communication, particularly on social media X. The non-formal characteristics of social media text, such as slang, abbreviations, emojis, and language switching within a single sentence, make sentiment analysis more challenging than monolingual text. This study aims to perform sentiment classification on code-mixed text by evaluating the performance of the lightweight Transformer model DistilBERT and comparing it with IndoBERTweet. The study adopts the CRISP-DM methodology, which consists of Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, and Deployment stages. The dataset comprises 1,108 primary data collected from social media X between 2022 and 2026 and 5,048 secondary data obtained from previous research. Three experimental scenarios were applied: translation into Indonesian, translation into English, and raw data without translation. The results show that DistilBERT achieved its best performance on the English translation scenario, with accuracies of 73.66% on the primary dataset and 72.00% on the secondary dataset. Meanwhile, IndoBERTweet obtained the highest performance on the Indonesian translation scenario, achieving an accuracy of 79.01%. These findings indicate that the alignment between the language of the input data and the pre-training characteristics of the model significantly affects sentiment classification performance on non-formal code-mixed text.
References
Nisrina Hanifa Setiono and Yunita Sari, “Exploring the Impact of Back-Translation on BERT’s Performance in Sentiment Analysis of Code-Mixed Language Data ,” vol. 19, Apr. 2025, doi: https://doi.org/10.22146/ijccs.104757.
Cuk Tho, Yaya Heryadi, Iman Herwidiana Kartowisastro, and Widodo Budiharto, “Code-Mixed Sentiment Analysis Indonesian-English Using Transformer Model,” ICIC Express Letters, vol. 17, no. 11, pp. 1295–1302, Feb. 2023, doi: 10.24507/icicel.17.11.1295.
A. F. Hidayatullah, “Code-Mixed Sentiment Analysis on Indonesian-Javanese-English Text Using Transformer Models,” in 2024 8th International Conference on Information Technology, Information Systems and Electrical Engineering (ICITISEE), IEEE, Aug. 2024, pp. 340–345. doi: 10.1109/ICITISEE63424.2024.10730138.
A. F. Hidayatullah, R. A. Apong, D. T. C. Lai, and A. Qazi, “Pre-trained language model for code-mixed text in Indonesian, Javanese, and English using transformer,” Soc. Netw. Anal. Min., vol. 15, no. 1, p. 30, Mar. 2025, doi: 10.1007/s13278-025-01444-9.
W. Mustikawati and Kundharu Saddhono, “Analisis Penggunaan Bahasa Jaksel Pada Video Berjudul ‘Language Barriers, Culture Shock & Tck Identity Crisis Ft. Mella Carli’ Di Youtube: Kajian Sosiolinguistik Campur Kode,” BASA Journal of Language & Literature, vol. 4, no. 2, pp. 89–94, Oct. 2024, doi: https://doi.org/10.33474/basa.v4i2.22031.
G. Singh, “Sentiment Analysis of Code-Mixed Social Media Text (Hinglish),” Feb. 2021, doi: https://doi.org/10.48550/arXiv.2102.12149.
A. K. Jena, A. K. Jena, K. M. Gopal, A. Tripathy, and N. Panda, “Optimized Feature Selection Approach for Semi-Supervised Sentiment Analysis of E-Commerce Feedback,” Journal of Computer Science, vol. 21, no. 2, pp. 363–379, Feb. 2025, doi: 10.3844/jcssp.2025.363.379.
A. Perera and A. Caldera, “Sentiment Analysis of Code-Mixed Text: A Comprehensive Review,” JUCS - Journal of Universal Computer Science, vol. 30, no. 2, pp. 242–261, Feb. 2024, doi: 10.3897/jucs.98708.
Q. Wu, P. Wang, and C. Huang, “MeisterMorxrc at SemEval-2020 Task 9: Fine-Tune Bert and Multitask Learning for Sentiment Analysis of Code-Mixed Tweets,” in Proceedings of the Fourteenth Workshop on Semantic Evaluation, Stroudsburg, PA, USA: International Committee for Computational Linguistics, 2020, pp. 1294–1297. doi: 10.18653/v1/2020.semeval-1.174.
S. Javdan, T. Shangipour ataei, and B. Minaei-Bidgoli, “IUST at SemEval-2020 Task 9: Sentiment Analysis for Code-Mixed Social Media Text Using Deep Neural Networks and Linear Baselines,” in Proceedings of the Fourteenth Workshop on Semantic Evaluation, Stroudsburg, PA, USA: International Committee for Computational Linguistics, 2020, pp. 1270–1275. doi: 10.18653/v1/2020.semeval-1.170.
A. Patil, V. Patwardhan, A. Phaltankar, G. Takawane, and R. Joshi, “Comparative Study of Pre-Trained BERT Models for Code-Mixed Hindi-English Data,” in 2023 IEEE 8th International Conference for Convergence in Technology (I2CT), IEEE, Apr. 2023, pp. 1–7. doi: 10.1109/I2CT57861.2023.10126273.
L. W. Astuti, Y. Sari, and S. -, “Code-Mixed Sentiment Analysis using Transformer for Twitter Social Media Data,” International Journal of Advanced Computer Science and Applications, vol. 14, no. 10, 2023, doi: 10.14569/IJACSA.2023.0141053.
F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization,” Sep. 2021, [Online]. Available: http://arxiv.org/abs/2109.04607
J. Devlin, M.-W. Chang, K. Lee, K. T. Google, and A. I. Language, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.” [Online]. Available: https://github.com/tensorflow/tensor2tensor
V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” Mar. 2020.
Y. Aliyu, A. Sarlan, K. Usman Danyaro, A. S. B. A. Rahman, and M. Abdullahi, “Sentiment Analysis in Low-Resource Settings: A Comprehensive Review of Approaches, Languages, and Data Sources,” IEEE Access, vol. 12, pp. 66883–66909, 2024, doi: 10.1109/ACCESS.2024.3398635.
M. A. Hasanah, S. Soim, and A. S. Handayani, “Implementasi CRISP-DM Model Menggunakan Metode Decision Tree dengan Algoritma CART untuk Prediksi Curah Hujan Berpotensi Banjir,” Journal of Applied Informatics and Computing, vol. 5, no. 2, pp. 103–108, Oct. 2021, doi: 10.30871/jaic.v5i2.3200.
A. D. Wijaya and B. Bram, “A Sociolinguistic Analysis Of Indoglish Phenomenon In South Jakarta,” PROJECT (Professional Journal of English Education), vol. 4, no. 4, p. 672, Jul. 2021, doi: 10.22460/project.v4i4.p672-684.
C. Tho, Y. Heryadi, L. Lukas, and A. Wibowo, “Code-mixed sentiment analysis of Indonesian language and Javanese language using Lexicon based approach,” J. Phys. Conf. Ser., vol. 1869, no. 1, p. 012084, Apr. 2021, doi: 10.1088/1742-6596/1869/1/012084.
D. I. Putri, A. N. Alfian, M. Y. Putra, and P. D. Mulyo, “IndoBERT Model Analysis: Twitter Sentiments on Indonesia’s 2024 Presidential Election,” Journal of Applied Informatics and Computing, vol. 8, no. 1, pp. 7–12, Jul. 2024, doi: 10.30871/jaic.v8i1.7440.
F. Barbieri, L. E. Anke, and J. Camacho-Collados, “XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond,” May 2022, [Online]. Available: http://arxiv.org/abs/2104.12250
H. Adel et al., “Improving Crisis Events Detection Using DistilBERT with Hunger Games Search Algorithm,” Mathematics, vol. 10, no. 3, Feb. 2022, doi: 10.3390/math10030447.
R. A. Casonatto, T. De Pádua Grillo Souza, and A. M. Mariano, “Quality and Risk Management in Data Mining: A CRISP-DM Perspective.,” Procedia Comput. Sci., vol. 242, pp. 161–168, 2024, doi: 10.1016/j.procs.2024.08.257.
J. Khan, K. Ahmad, S. K. Jagatheesaperumal, and K. A. Sohn, “Textual variations in social media text processing applications: challenges, solutions, and trends,” Artif. Intell. Rev., vol. 58, no. 3, Mar. 2025, doi: 10.1007/s10462-024-11071-z.
N. Calderon et al., “Measuring the Robustness of NLP Models to Domain Shifts,” Apr. 2024, [Online]. Available: http://arxiv.org/abs/2306.00168
S. Chen, L. Neves, and T. Solorio, “Mitigating Temporal-Drift: A Simple Approach to Keep NER Models Crisp,” Online Workshop, 2021. [Online]. Available: https://github.com/
A. Bansal and I. Gangwani, “Zero-Training Temporal Drift Detection for Transformer Sentiment Models: A Comprehensive Analysis on Authentic Social Media Streams,” Nov. 2025, [Online]. Available: http://arxiv.org/abs/2512.20631
Copyright (c) 2026 Syifa Arifah Nurbayani, Dian Sa'adillah Maylawati, Aldy Rialdy Atmadja

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors retain copyright and grant the EXPLORER right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (CC BY-SA 4.0) that allows others to share (copy and redistribute the material in any medium or format) and adapt (remix, transform, and build upon the material) the work for any purpose, even commercially with an acknowledgement of the work's authorship and initial publication in EXPLORER.
Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in EXPLORER.
Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).





.png)















