Hello, I’m Shariff.
Data Scientist & Software Engineer
I build applied AI systems and dependable software. I’m completing a BSc in Information Systems at Singapore Management University, with a second major in Computer Science (Artificial Intelligence).
Recently, I’ve taken AI products from prototype to production by building structured AI workflows, tracing prompts, improving document-processing reliability, and designing evaluations with domain experts.
Work experience
GovTech (Government Technology Agency)Data Scientist InternMay 2026 – Present
SCG Digital Governance · Central Digital Assurance (CDA-IM8)
- Own full-stack development, AI engineering, and production operations for two AI products; built one from scratch and took over the second through launch.
- Instrumented both products with Langfuse tracing and prompt version management; reworked prompts to improve cache reuse and reduce latency.
- Built a queue system for document jobs to resolve pre-production memory failures, then deployed shared PDF and DOCX ingestion for reuse across the team’s AI products.
HTX (Home Team Science and Technology Agency)Software Engineering InternJan 2026 – Apr 2026
xDigital · AI Products Team
- Developed features in a full-stack TypeScript codebase for AI-assisted government report workflows.
- Built a custom MCP server that exposed report creation, file upload, and lookup as typed tools for agent workflows.
- Added Playwright end-to-end coverage to GitLab CI for critical report journeys ahead of launch.
Featured projects
iPiD Growth IntelligenceFour connected AI workflows link market monitoring, account research, content production, and sales outreach to cited evidence and review history.Singlish Hate-Speech GuardrailBuilt a hate-speech guardrail for Singlish and code-switched text, then tested whether RoBERTa embeddings with XGBoost could outperform an end-to-end classifier.PII Anonymisation: AgenticBuilt a fully local document anonymiser that combines NER, generative replacement, and adversarial re-identification, improving anonymisation from 87% to 97%.Hate Speech with BERTFull RoBERTa fine-tuning reached 0.689 macro-F1; LoRA trained only 0.71% of the parameters but fell to 0.630.
Older projects
Experiments
Some stuff I made for fun :)