Tag: deep learning NLP Indonesia

  • Comparison of Bidirectional-LSTM and GRU Models for Sentiment Analysis in Bahasa Indonesia

    Authors: Muhammad Rajab Fachrizal, Annisa Paramitha Fadillah, Lusi Melian

    DOI: 10.1109/INCITEST64888.2024.11121460

    Abstract

    Bidirectional-LSTM and GRU are models that can be used to process sequential data, including text data. This study, compares the two deep learning models to classify text for sentiment analysis using the Bahasa. The dataset used is the JKN BPJS Kesehatan mobile application user review data obtained from the Google Play Store site. After text preprocessing, the amount of data to be processed is 93517 with three target labels, positive, negative, and neutral. By using several model parameters such as Number of Units, Activation, Batch Size, Dropout, and other parameters, the test results obtained are that the Bidirectional-LSTM model has a slight accuracy value of 96.70% and higher precision, recall, and F1Score values compared to the GRU model. © 2024 IEEE.

    Author keywords

    Bahasa; Bi-LSTM; deep learning; GRU; sentiment analysis; text classification

    This article can be accessed at https://www.scopus.com/pages/publications/105015853213

  • A Bert Model to Detect Provocative Hoax

    A Bert Model to Detect Provocative Hoax

    Authors: Rio Yunanto, Eri Prasetyo Wibowo, Rianto R

    Abstract

    The information flood makes social media users vulnerable to becoming victims of provocative hoaxes or even spreading hoaxes themselves. This research examines the capabilities of two variants of Bidirectional Encoder Representations from Transformers (BERT) models for the Indonesian language (IndoBERT Base Model and Indonesian BERT base model 522M) in developing the detection of provocative hoaxes in the Indonesian language. The proposed method used two variants of the monolingual BERT model for the Indonesian language from the Huggingface library. The proposed method’s architectural flow starts with data collection and labelling from community hoax collector websites, followed by pre-processing. The cleaned data is then divided into training and test data to proceed to the fine-tuning stage, where several layers and weights of the BERT model are adjusted to fit the desired classification task. The experimental results of the study show that the recommended Indonesian BERT variant for the detection of provocative hoaxes is the IndoBERT Base Model with a learning rate of 1e-5, a batch size of 32, and a maximum length limit of 128 tokens, achieving an average training accuracy of 99,22%, with a training time of 21min 52s. The research findings also indicate that a learning rate 1e-5 can produce better test accuracy than a learning rate of 2e-5 or 3e-5. The detection model of provocative hoaxes using Indonesian BERT variants needs to be improved, especially in terms of collecting a large amount of hoax data, to enhance the accuracy of the provocative hoax detection model. © School of Engineering, Taylor’s University.

    Author keywords

    Accuracy; Classification; Hoax; Provocative; Transformer

    This article can be accessed at https://www.scopus.com/pages/publications/85180943734