AI/ML · November 1, 2025
Singlish-Aware Hate Speech Guardrail
Built a hate-speech guardrail for Singlish and code-switched text, then tested whether RoBERTa embeddings with XGBoost could outperform an end-to-end classifier.

Overview
Developed as part of the 11th DAP (Data Associates Program) under SMUBIA (Business Intelligence & Analytics Club), this project began with one question: can rich transformer embeddings paired with traditional machine-learning classifiers match end-to-end deep learning for hate-speech detection?
We generated 768-dimensional representations with BERT and RoBERTa, compared vanilla and fine-tuned embeddings across four classifiers, and then adapted the encoder to Singlish and code-switched text through continued masked-language-model pre-training.
Dataset: HateXplain, with 15,383 training samples and 1,924 test samples labelled as hate, offensive, or normal, plus annotator rationales.
Embedding Space Analysis
Embedding Geometry
We first measured the geometric structure of embeddings from vanilla (pre-trained) vs. fine-tuned encoders:
| Model | Silhouette | kNN Purity |
|---|---|---|
| Vanilla RoBERTa [cls] | −0.003 | 44.8% |
| Vanilla BERT [cls] | −0.004 | 41.0% |
| Fine-Tuned BERT Binary [cls] | +0.181 | 82.7% |
| Fine-Tuned RoBERTa 3-Class [cls] | +0.177 | 72.1% |
Key finding: Fine-tuning yields a ~50× improvement in cluster quality (silhouette ≈ 0.00 → ≈ 0.18). Vanilla models show near-zero or negative silhouette scores — classes are geometrically indistinguishable. Fine-tuning reshapes all 768 dimensions to create real cluster structure, making shallow classifiers effective.
3-Class Classification Results
| Embeddings | Classifier | Accuracy | Macro F1 |
|---|---|---|---|
| Fine-Tuned RoBERTa [cls] | XGBoost | 70.3% | 0.691 |
| Fine-Tuned RoBERTa [cls] | Random Forest | 70.0% | 0.686 |
| Fine-Tuned RoBERTa [cls] | Logistic Regression | 68.8% | 0.681 |
| Fine-Tuned BERT [cls] | XGBoost | 69.5% | 0.684 |
| Vanilla RoBERTa [cls] | Logistic Regression | 62.2% | 0.610 |
| End-to-end RoBERTa + classification head | (reference) | 70.0% | 0.689 |
The embedding + XGBoost approach exceeded the end-to-end RoBERTa baseline (70.3% vs 70.0%).
Binary Classification Results
| Embeddings | Classifier | Accuracy | Macro F1 |
|---|---|---|---|
| Fine-Tuned BERT [rationale] | XGBoost | 86.5% | 0.840 |
| Fine-Tuned BERT [rationale] | MLP (256→128→64) | 85.7% | 0.835 |
| End-to-end BERT + classification head | (reference) | 87.7% | 0.874 |
Within 1.2% of end-to-end fine-tuning — at a fraction of the compute cost.
Key Insights
- Fine-tuned embeddings drove most of the performance gain. Moving from vanilla to fine-tuned embeddings gained about 8% F1. Changing classifiers within the same embedding space moved F1 by 2–4%.
- XGBoost used the fine-tuned feature space better than the linear head. End-to-end RoBERTa reached 70.0% accuracy, while frozen embeddings with XGBoost reached 70.3%.
- Frozen embeddings kept classifier training modular and GPU-free. We could compare several traditional classifiers without fine-tuning the encoder for each run.
Hack-Day Extension: A Singlish-Aware Guardrail
The original benchmark answered a narrow question on HateXplain. The hack-day extension asked a harder one: would the pipeline recognise hate and offensive language when people switch between English, Singlish, and local slang?
We continued masked-language-model pre-training on Singaporean text before extracting embeddings. This gave the encoder more exposure to local vocabulary and phrasing without changing the guardrail's three outputs: hate, offensive, or normal.
Reflection: Adversarial Data Was the Hard Part
Generating credible adversarial examples was harder than wiring up the classifier. General chat models often returned language that was sanitised, stereotypical, or recognisably unlike natural Singlish. Using the same generator for both training and evaluation would also risk measuring the generator's writing style rather than the guardrail's ability to generalise.
In this bounded experiment, coding agents were more useful for producing structured test fixtures inside a reviewable engineering workflow. Gemini also appeared more willing to return some Asian-localised examples. That is an observation from this experiment, not evidence that Gemini broadly has weaker safeguards for Asian content; model versions, prompt wording, and locale coverage may also explain the difference.
Every generated example must therefore remain untrusted until human review, and evaluation data needs a source independent from the training-data generator. Real regional data would be stronger, provided its licensing, privacy, and annotation quality were handled properly.
Datasets
- HateXplain: Labels for hate, offensive, and normal text, plus annotator rationales used for rationale-weighted pooling
- Singaporean text corpus: Unlabelled text used for continued masked-language-model pre-training and local vocabulary exposure
Technologies
Python, PyTorch, Hugging Face Transformers, Scikit-learn, XGBoost, Weights & Biases, NumPy, Pandas, Matplotlib