SR — Shariff Rashid, home
← Back to projects

AI/ML · November 1, 2025

Singlish-Aware Hate Speech Guardrail

Built a hate-speech guardrail for Singlish and code-switched text, then tested whether RoBERTa embeddings with XGBoost could outperform an end-to-end classifier.

Singlish-Aware Hate Speech Guardrail

Overview

Developed as part of the 11th DAP (Data Associates Program) under SMUBIA (Business Intelligence & Analytics Club), this project began with one question: can rich transformer embeddings paired with traditional machine-learning classifiers match end-to-end deep learning for hate-speech detection?

We generated 768-dimensional representations with BERT and RoBERTa, compared vanilla and fine-tuned embeddings across four classifiers, and then adapted the encoder to Singlish and code-switched text through continued masked-language-model pre-training.

Dataset: HateXplain, with 15,383 training samples and 1,924 test samples labelled as hate, offensive, or normal, plus annotator rationales.


Embedding Space Analysis

Embedding Geometry

We first measured the geometric structure of embeddings from vanilla (pre-trained) vs. fine-tuned encoders:

ModelSilhouettekNN Purity
Vanilla RoBERTa [cls]−0.00344.8%
Vanilla BERT [cls]−0.00441.0%
Fine-Tuned BERT Binary [cls]+0.18182.7%
Fine-Tuned RoBERTa 3-Class [cls]+0.17772.1%

Key finding: Fine-tuning yields a ~50× improvement in cluster quality (silhouette ≈ 0.00 → ≈ 0.18). Vanilla models show near-zero or negative silhouette scores — classes are geometrically indistinguishable. Fine-tuning reshapes all 768 dimensions to create real cluster structure, making shallow classifiers effective.

3-Class Classification Results

EmbeddingsClassifierAccuracyMacro F1
Fine-Tuned RoBERTa [cls]XGBoost70.3%0.691
Fine-Tuned RoBERTa [cls]Random Forest70.0%0.686
Fine-Tuned RoBERTa [cls]Logistic Regression68.8%0.681
Fine-Tuned BERT [cls]XGBoost69.5%0.684
Vanilla RoBERTa [cls]Logistic Regression62.2%0.610
End-to-end RoBERTa + classification head(reference)70.0%0.689

The embedding + XGBoost approach exceeded the end-to-end RoBERTa baseline (70.3% vs 70.0%).

Binary Classification Results

EmbeddingsClassifierAccuracyMacro F1
Fine-Tuned BERT [rationale]XGBoost86.5%0.840
Fine-Tuned BERT [rationale]MLP (256→128→64)85.7%0.835
End-to-end BERT + classification head(reference)87.7%0.874

Within 1.2% of end-to-end fine-tuning — at a fraction of the compute cost.

Key Insights

  1. Fine-tuned embeddings drove most of the performance gain. Moving from vanilla to fine-tuned embeddings gained about 8% F1. Changing classifiers within the same embedding space moved F1 by 2–4%.
  2. XGBoost used the fine-tuned feature space better than the linear head. End-to-end RoBERTa reached 70.0% accuracy, while frozen embeddings with XGBoost reached 70.3%.
  3. Frozen embeddings kept classifier training modular and GPU-free. We could compare several traditional classifiers without fine-tuning the encoder for each run.

Hack-Day Extension: A Singlish-Aware Guardrail

The original benchmark answered a narrow question on HateXplain. The hack-day extension asked a harder one: would the pipeline recognise hate and offensive language when people switch between English, Singlish, and local slang?

We continued masked-language-model pre-training on Singaporean text before extracting embeddings. This gave the encoder more exposure to local vocabulary and phrasing without changing the guardrail's three outputs: hate, offensive, or normal.

Reflection: Adversarial Data Was the Hard Part

Generating credible adversarial examples was harder than wiring up the classifier. General chat models often returned language that was sanitised, stereotypical, or recognisably unlike natural Singlish. Using the same generator for both training and evaluation would also risk measuring the generator's writing style rather than the guardrail's ability to generalise.

In this bounded experiment, coding agents were more useful for producing structured test fixtures inside a reviewable engineering workflow. Gemini also appeared more willing to return some Asian-localised examples. That is an observation from this experiment, not evidence that Gemini broadly has weaker safeguards for Asian content; model versions, prompt wording, and locale coverage may also explain the difference.

Every generated example must therefore remain untrusted until human review, and evaluation data needs a source independent from the training-data generator. Real regional data would be stronger, provided its licensing, privacy, and annotation quality were handled properly.


Datasets

  • HateXplain: Labels for hate, offensive, and normal text, plus annotator rationales used for rationale-weighted pooling
  • Singaporean text corpus: Unlabelled text used for continued masked-language-model pre-training and local vocabulary exposure

Technologies

Python, PyTorch, Hugging Face Transformers, Scikit-learn, XGBoost, Weights & Biases, NumPy, Pandas, Matplotlib

← Back to projects