About Me
I'm a data scientist and AI/ML engineer based in Jakarta, Indonesia — building systems at the intersection of language, data, and decisions that actually matter. My work spans predictive modeling and ML pipelines on one side, and production LLM systems on the other: RAG pipelines, AI agents, and text-to-SQL tools that let non-technical people query real financial data in plain English. I also contribute to NLP research on Southeast Asian languages.
I studied Mathematics at Universitas Indonesia, graduating in 2023 — where I first realized that the most interesting problems live at the intersection of math, language, and human behavior. That conviction has shaped everything I've worked on since.
You can take a look at my experience on my LinkedIn. · Open for full-time AI/ML engineering and data science roles — feel free to reach out.
Experience
-
2025 — now
Independent AI/ML Engineer — building a portfolio of production LLM systems while consulting for clients across fintech and consumer domains.
-
2024
Data Analytics at Bank Saqu / Bank Jasa Jakarta (Astra & WeLab Group)
-
2023 — now
NLP Research Contributor at SEACROWD — building open-source NLP resources for Southeast Asian languages
-
2023
Product Operations Intern at Deall Jobs (YC 2022)
-
2022 — 23
Data Scientist Intern at Home Credit Indonesia — project-based intern with Rakamin
-
2022
Digital Consulting Analyst at Advisia
-
2022
Big Data Analytics Intern at Kimia Farma — project-based intern with Rakamin
-
2022
Business Analyst Intern at DailySocial.id
-
2021 — 22
Product Management Intern at SejutaCita (YC 2022)
-
2019 — 23
🎓 B.Sc. Mathematics — Universitas Indonesia. Thesis on NLP & Twitter sentiment analysis of football fan bases.
Interests
I'm drawn to natural language processing — specifically how we can build models that understand the nuances of Southeast Asian languages, which are still deeply underrepresented in AI research. Working with SEACROWD made this feel less like an academic concern and more like something that actually matters for the hundreds of millions of people whose languages rarely appear in a training dataset.
I'm also genuinely interested in applied ML in finance: credit risk modeling, behavioral prediction, and how data can help build fairer systems for people historically excluded from financial services. That interest became a full credit risk decision engine — PD modeling, SHAP explainability, expected loss, portfolio analytics, and stress testing on top of a 300K-applicant consumer lending dataset.
Lately I've been going deep into large language models — not just using them, but understanding how to build real systems around them. I've now shipped five production LLM projects, each one building on the last: from conversation memory and RAG pipelines, through AI agents that reason across live tools, to a full production API with observability, all the way to a Fintech LLM Analyst that takes plain-English questions and translates them into SQL queries against real financial data. The through-line across all of it is the same question: what does it actually look like when AI works for the people using it, not just the people building it? More recently that's pushed into document intelligence — a resilient vision-language pipeline for extracting structured data from receipts, a citation-grounded RAG system for internal documents, and a QLoRA fine-tune of a sub-3B model for fintech text-to-SQL.
Tech Stack
Outside of Work
I play a lot of games — and I mean a lot. My range is embarrassingly wide: one moment I'm sweating in CS or grinding ranked in MLBB, the next I'm in Roblox without a hint of shame. A big part of what I play these days is honestly influenced by Windah Basudara — if he touches it, there's a good chance I end up trying it.
I watch anime. My favorite is One Piece — I've accepted that it will outlive most of my life decisions. My favorite character is Buggy, which I think says a lot about me and I'm fine with that. Outside the big three, the shows that genuinely stuck with me are Frieren, Saiki K, and Bocchi the Rock.
I also follow football. I support Manchester United, which at this point is less a hobby and more a test of character. We're not great. We've been not great for a while. But here we are.
Awards & Research
- SEA-VL: Multicultural Vision-Language Dataset for Southeast Asia. Co-author. Cahyawijaya et al., ACL 2025 (Long Papers). Contributed to dataset preprocessing and documentation. [ACL Anthology]
- Dato' Dr. Low Tuck Kwong Scholarship.
- Most Outstanding Student — Mathematics Dept., UI. Recognized out of 600+ students.
- 1st Runner Up — ASEAN Data Science Challenge 2022. Proposed GetHired, a platform for youth employment readiness. Competed against 200+ teams.
- 1st Place — Ganesha Student Innovation Summit Challenge 2022. Outperformed 100+ individual participants and 10 finalist teams.
- 1st Winner — International Youth Summit for Renewable Energy 2021. Waste to Energy Challenge. Awarded USD 1,800. Outperformed 200+ participants.
Featured In
Selected Projects
RAG Evaluation & Optimization System
An automated evaluation harness that scores RAG pipelines across multiple metrics and wires them into a GitHub Actions quality gate.
- Architecture: Scores based on faithfulness, answer relevancy, context recall, and context precision across 7 configs, with a second LLM-as-judge layer.
- Key Insight: Chasing an eval score gap led to root-causing a stale Chroma index (chunks silently duplicated on re-index) rather than a real architecture difference.
Credit Risk Decision Engine
An end-to-end credit risk system built on a 307K-applicant dataset, wrapped in a 4-page interactive dashboard.
- Core Models: PD modeling (WOE/Logistic Regression vs. LightGBM) and risk grading.
- Features: SHAP explainability, expected loss calculations, portfolio analytics, stress testing, and PSI-based drift monitoring.
- Key Insight: Discovered that setting
is_unbalance=Truebarely moved ranking metrics but spiked predicted default probability by 5x—which would have catastrophically inflated downstream Expected Loss figures.
Document Intelligence RAG System
RAG over mixed-format internal documents, where every answer cites its exact source down to the page, slide, or row range.
- Sources: PDFs, Word docs, slide decks, spreadsheets, and CSVs.
- Architecture: Structure-aware chunking per format, local embeddings, Postgres/pgvector retrieval, and full Langfuse tracing on every query.
Resilient Document Extraction Engine — VLM Pipeline
A production-minded VLM pipeline for extracting structured data from receipts.
- Pipeline: Groq Llama Vision as the primary extractor with a GPT-4o fallback.
- Reliability: A self-healing Pydantic schema flags arithmetic inconsistencies instead of hard-failing.
- Key Insight: Naive downscaling to cut inference cost tanked accuracy on right-aligned totals by 22%; an OpenCV multi-contour crop step recovered nearly all of it at the same resolution.
text2sql-finetune — QLoRA Fine-Tune for Fintech SQL
Fine-tuning a sub-3B model (Qwen2.5-Coder-1.5B) for fintech text-to-SQL with QLoRA.
- Evaluation: Scored strictly on live SQLite execution accuracy rather than textual similarity.
- Key Insight: Both base and fine-tuned models hit 0% execution accuracy — but for opposite reasons: the base model buried valid SQL in chatty prose, while the fine-tune over-complicated simple queries after inheriting the training set's bias toward multi-table joins.
P4 · LLM API — Production Backend
A production-grade LLM API with RAG, persistent chat history, and full observability.
- Observability: Every model call logged via Langfuse tracing — inspectable and debuggable, not a black box.
- Key Insight: The gap between "works in a notebook" and "works in production" is mostly observability — this project closed that gap.
View All Projects →
12 more projects spanning production LLM systems, applied ML, NLP research, and early work.