Fine-tuning IndoBERT for Indonesian sentiment
· 1 min read
Indonesian social media text is messy. It mixes slang, abbreviations, regional words, and English in the same sentence. A model pretrained on Indonesian text, like IndoBERT, is a much better starting point than a generic multilingual model.
The setup
- Three labels: Positive, Neutral, and Negative.
- Light text normalization before tokenization.
- A classification head on top of the pretrained encoder, fine-tuned end to end.
from transformers import AutoTokenizer, AutoModelForSequenceClassification
name = "indobenchmark/indobert-base-p1"
tokenizer = AutoTokenizer.from_pretrained(name)
model = AutoModelForSequenceClassification.from_pretrained(name, num_labels=3)Serving it
The model sits behind a small Flask REST API that returns the predicted label together with the full probability distribution, so the UI can show how confident each prediction is.
A prediction without a confidence score is only half an answer.