Authors: Rio Yunanto, Eri Prasetyo Wibowo, Rianto R
Abstract
The information flood makes social media users vulnerable to becoming victims of provocative hoaxes or even spreading hoaxes themselves. This research examines the capabilities of two variants of Bidirectional Encoder Representations from Transformers (BERT) models for the Indonesian language (IndoBERT Base Model and Indonesian BERT base model 522M) in developing the detection of provocative hoaxes in the Indonesian language. The proposed method used two variants of the monolingual BERT model for the Indonesian language from the Huggingface library. The proposed method’s architectural flow starts with data collection and labelling from community hoax collector websites, followed by pre-processing. The cleaned data is then divided into training and test data to proceed to the fine-tuning stage, where several layers and weights of the BERT model are adjusted to fit the desired classification task. The experimental results of the study show that the recommended Indonesian BERT variant for the detection of provocative hoaxes is the IndoBERT Base Model with a learning rate of 1e-5, a batch size of 32, and a maximum length limit of 128 tokens, achieving an average training accuracy of 99,22%, with a training time of 21min 52s. The research findings also indicate that a learning rate 1e-5 can produce better test accuracy than a learning rate of 2e-5 or 3e-5. The detection model of provocative hoaxes using Indonesian BERT variants needs to be improved, especially in terms of collecting a large amount of hoax data, to enhance the accuracy of the provocative hoax detection model. © School of Engineering, Taylor’s University.
Author keywords
Accuracy; Classification; Hoax; Provocative; Transformer
This article can be accessed at https://www.scopus.com/pages/publications/85180943734