TY - GEN
T1 - Sentiment analysis of noisy malay text using a large language model
AU - Khalip, Khairul Imran
AU - Khalif, Ku Muhammad Naim Ku
AU - Aziz, Mohd Khairul Bazli Mohd
AU - Gegov, Alexander
N1 - Publisher Copyright:
© 2025 IEICES/Kyushu University. All rights reserved.
PY - 2025/10/30
Y1 - 2025/10/30
N2 - Due to the informality of social media, Malay user-generated content sentiment analysis is dificult. Existing methods struggle to capture cultural and contextual details. This study proposes publishing an open-source annotated dataset, fine-tuning an open-source large language model (LLM), and using an open-source chatbot interface to createa robust sentiment analysis model for noisy Malay text. The research addresses three main issues: lack of labelled Malay social media data, in suficient generic Malay language models, and lack of practical sentiment analysis tools. Its three goals are to create a diverse dataset with accurate sentiment labels, parameter-efficiently fine-tune an LLM, and export the model for an interactive chatbot. The process involves collecting social media data using Contextual Lexical Adaptation, pre-processing and analysing it, fine-tuning the TinyLlama LLM using LoRA, and comparingit to traditional models. Real-world applications, such as sentiment analysis of Malaysian tweets, will be shown using a locally deployed chatbot interface for fine-tuned model inference. This study lays the groundwork for practical sentiment analysis, benefiting businesses, researchers, and politicians seeking data-driven insights. This research aims to revolutionise open-source Malay sentiment analysis by addressing current limitations through an integrated approach.
AB - Due to the informality of social media, Malay user-generated content sentiment analysis is dificult. Existing methods struggle to capture cultural and contextual details. This study proposes publishing an open-source annotated dataset, fine-tuning an open-source large language model (LLM), and using an open-source chatbot interface to createa robust sentiment analysis model for noisy Malay text. The research addresses three main issues: lack of labelled Malay social media data, in suficient generic Malay language models, and lack of practical sentiment analysis tools. Its three goals are to create a diverse dataset with accurate sentiment labels, parameter-efficiently fine-tune an LLM, and export the model for an interactive chatbot. The process involves collecting social media data using Contextual Lexical Adaptation, pre-processing and analysing it, fine-tuning the TinyLlama LLM using LoRA, and comparingit to traditional models. Real-world applications, such as sentiment analysis of Malaysian tweets, will be shown using a locally deployed chatbot interface for fine-tuned model inference. This study lays the groundwork for practical sentiment analysis, benefiting businesses, researchers, and politicians seeking data-driven insights. This research aims to revolutionise open-source Malay sentiment analysis by addressing current limitations through an integrated approach.
KW - Fine-Tuning
KW - Large Language Model
KW - Sentiment Analysis
UR - https://www.scopus.com/pages/publications/105027057660
U2 - 10.5109/7395763
DO - 10.5109/7395763
M3 - Conference contribution
AN - SCOPUS:105027057660
VL - 11
T3 - International Exchange and Innovation Conference on Engineering and Sciences
SP - 1904
EP - 1909
BT - Proceedings of International Exchange and Innovation Conference on Engineering & Sciences (IEICES)
PB - Kyushu University
T2 - 11th International Exchange and Innovation Conference on Engineering and Sciences, IEICES 2025
Y2 - 30 October 2025 through 31 October 2025
ER -