Skip to content

ML / AI Engineer · Forward-Deployed Engineer · Taipei

I’m Sakkarin.

I teach machines to make decisions — and then I measure whether they actually did.

Thai engineer in Taipei. By day I build Claude Code tooling and C# fund systems at SystemWeb; the rest of the time I build RL agents, world models, and LLM workflows — in public, with the null results left in.

Open to ML / AI Engineer and Forward-Deployed Engineer roles

fig. 1 — tabular Q-learning, live in your browser · 12×8 · α .5 · γ .95

§ 01

Now

What’s on the bench this month.

updated Sep 28, 2026

Researching
janus-chrysalis Phase 2. The RSSM training path cleared its loss-curve gate; next up is validating the non-stationarity instrument before Gate G2.
At work
Claude Code plugins and a graph-based agent memory for our engineering team, alongside the C# Transfer Agent System.
Learning
The inference layer: CUDA, NVIDIA NIM, NeMo, TensorRT-LLM. The next RAG project gets measured latency, not target latency.
Looking for
ML / AI Engineer or Forward-Deployed Engineer roles, somewhere models meet real users and slightly messy data.

§ 02

Selected work

Public repos only, so you can check the numbers. Every metric shows its n, and every case study has a box listing what the numbers don’t show.

all work →
activeevidence: null result

janus-chrysalis

Open research on world models in multi-agent RL, trained in TypeScript, co-authored with Claude, and published even when the effect disappears.

Can you measure how much a world model's prediction error comes from the other agent's learning, rather than from the model just being new? I built a freeze-intervention instrument to isolate that signal in a two-agent gridworld. The first 9 seeds looked significant (p ≈ 0.039). A pre-registered replication took it to p ≈ 0.388, so the write-up reports a null result. The work runs as a human + Claude collaboration: a daily autonomous agent opens PRs, and I review and merge them.

Pre-registered replication
p ≈ 0.388
n = 12 seeds, 8 negative · vs p ≈ 0.039 at n = 9 · did not replicate
G2 loss-curve gate
6 / 6 pass
n = 3 seeds × 2 agents · total loss 2.2–2.9 → 1.4–1.8 over 20 episodes
Test suite
165 passing
n = 177 tests, 12 todo, 0 fail
Autonomous loop
77 stand-ups
n = Jul – Sep 2026

TypeScript · TensorFlow.js (tfjs-node) · Node 22 · RSSM / Dreamer-style world model · node:test

pausedevidence: inconclusive

MaxEnt actor-critic for portfolio allocation

A reproducible Soft Actor-Critic research harness for long-only allocation. Built properly, measured against baselines, and not yet beating them.

Python · PyTorch · Gymnasium-style env · SAC (twin Q, auto-entropy)

fig. 1, again

The gridworld up top is a tiny version of the same idea: an agent, a reward, and a number you can check against the optimum. Draw a wall and watch it re-learn.

back to the demo ↑

§ 03

Lab bench

Smaller and local-only experiments. One line each, with the evidence level shown, not hidden.

  • planner-executor

    Python · MLX · Claude CLI

    Frontier model plans and rescues, a local Qwen 9B on MLX executes behind per-step JSON gates. On a 15-document extraction pilot it matched all-frontier accuracy at ~56% of the serving cost.

    evidence: preliminaryn = 15 docslocal only
  • Fortuna

    SwiftUI · FoundationModels · SwiftData

    My own iPhone finance app, and yes, I use it daily. An on-device Apple Foundation Model answers budget questions through 5 finance tools; card statements import straight from PDF. No servers, no dependencies.

    evidence: in daily use63 testslocal only
  • constitutional-guardrail-gateway

    FastAPI · pydantic · YAML policy

    OpenAI-compatible gateway that enforces a written 10-principle constitution as policy-as-code, with PII / injection detectors and a hash-chained audit log. Red-team eval designed, not yet run.

    evidence: self-check62 testslocal only
  • model-router-alpha

    scikit-learn · MiniLM · FastAPI

    Routes each query to the cheapest Claude tier that clears a quality bar. The eval design is the point: calibrated classifiers, bootstrap CIs, a κ-checked LLM judge.

    evidence: synthetic only830-prompt test sets
  • PASSAGE

    LangGraph · pydantic

    KYC/AML onboarding for a fictional bank: a 4-stage LangGraph pipeline with schema-checked handoffs, a rules-based risk engine, and a human review step.

    evidence: synthetic onlyoffline core34 tests
  • RL from scratch

    PyTorch · Gymnasium · React

    Single-file PPO (CartPole, Pendulum) and a Dreamer-v1-style world model with a live React training view. The PPO results are seed-sensitive, and I report them that way.

    evidence: measuredseed-sensitive

§ 04

Up next

Planned, not built. Listed so nobody mistakes a target for a result.

  • rag-fin-qa

    planned

    Q&A over SEC filings and earnings calls with GPU-served retrieval and faithfulness evals. Scaffolded; the latency and faithfulness targets are targets, not results.

    RAG · Ragas · Triton

  • cuda-quant-kernels

    planned

    Hand-written CUDA kernels for quant workloads, profiled with Nsight — my own kernels this time, not the samples.

    CUDA · Nsight

§ 05

How I work

Three habits I keep coming back to, at work and on nights-and-weekends projects.

Plan first, read widely, then go explore

Most of my projects start life as a plan doc and a reading list. janus-chrysalis keeps notes on every paper it leans on. Then I try the new approach anyway, because that’s the fun part, and the plan tells me whether it actually made things better or just made them different.

Small steps, real numbers

Ship something small, measure it, then decide. I’d rather show you a modest number I can defend than a big one I can’t. That’s why every number on this site comes with its n, and why one of my favourite results is a null result.

Start from the person, not the tech

Figure out what they actually need first. So far that’s been fund administrators, a Canadian bank’s data team, Thailand’s largest mobile operator, and, for Fortuna, me squinting at my own credit-card statement. The model comes second.

§ 06

Day job

Where agent tooling meets a real engineering team, and the C# fund systems it helps ship.

full experience →

Oct 2023 — now · Taipei, Taiwan

SystemWeb Technologies

Software Developer

active

I build the AI tooling our engineering team works with every day, on top of the C# fund-administration systems we ship.

Claude Code plugins published internally
3
per-feature development time
days → hours
engineers using them daily
3–4
documents in the agent memory graph
500
  • Architected and published 3 plugins to our internal Claude Code plugin marketplace, for task execution, ideation, and AI-native SDLC onboarding for engineers new to AI-assisted work. Per-feature development went from days to hours, and 3–4 engineers use them every day.
  • Designed a graph-based persistent memory layer for our agent system, replacing ad-hoc context handling. It indexes 500 documents so agents can recall what happened in earlier sessions.
  • Underneath it all: the C# Transfer Agent System for asset- and wealth-management clients (and the Portfolio Management System before it), built with a 5–10 person team split between Thailand and Taiwan.

Claude Code · Agent memory (graph) · LLM multi-agent systems · C# · JavaScript · SQL

§ 07

Writing

all posts →

§ 08

Off the keyboard

Brazilian jiu-jitsu

The most honest feedback loop I know: bad policy, immediate negative reward.

Cooking

Recipe testing is just hyperparameter search with better snacks.

Films

How I switch my brain off after a long training run. Somebody else’s story, somebody else’s decisions.

§ 09

Let’s talk

Open to ML / AI Engineer and Forward-Deployed Engineer roles. Based in Taipei (UTC+8).