habiburrahman.pro
← Writing

Fine-tuning IndoBERT for Indonesian sentiment

· 1 min read

Indonesian social media text is messy. It mixes slang, abbreviations, regional words, and English in the same sentence. A model pretrained on Indonesian text, like IndoBERT, is a much better starting point than a generic multilingual model.

The setup

  • Three labels: Positive, Neutral, and Negative.
  • Light text normalization before tokenization.
  • A classification head on top of the pretrained encoder, fine-tuned end to end.
from transformers import AutoTokenizer, AutoModelForSequenceClassification
 
name = "indobenchmark/indobert-base-p1"
tokenizer = AutoTokenizer.from_pretrained(name)
model = AutoModelForSequenceClassification.from_pretrained(name, num_labels=3)

Serving it

The model sits behind a small Flask REST API that returns the predicted label together with the full probability distribution, so the UI can show how confident each prediction is.

A prediction without a confidence score is only half an answer.