Skip to content

work · evidence-labelled

Work

Three tiers. Case studies are public repos you can check. Lab bench items are real code that’s smaller, local-only, or not measured yet. Up next is planned work, marked as such.

§ A

Case studies

activeevidence: null result

janus-chrysalis

Open research on world models in multi-agent RL, trained in TypeScript, co-authored with Claude, and published even when the effect disappears.

Can you measure how much a world model's prediction error comes from the other agent's learning, rather than from the model just being new? I built a freeze-intervention instrument to isolate that signal in a two-agent gridworld. The first 9 seeds looked significant (p ≈ 0.039). A pre-registered replication took it to p ≈ 0.388, so the write-up reports a null result. The work runs as a human + Claude collaboration: a daily autonomous agent opens PRs, and I review and merge them.

Pre-registered replication
p ≈ 0.388
n = 12 seeds, 8 negative · vs p ≈ 0.039 at n = 9 · did not replicate
G2 loss-curve gate
6 / 6 pass
n = 3 seeds × 2 agents · total loss 2.2–2.9 → 1.4–1.8 over 20 episodes
Test suite
165 passing
n = 177 tests, 12 todo, 0 fail
Autonomous loop
77 stand-ups
n = Jul – Sep 2026

TypeScript · TensorFlow.js (tfjs-node) · Node 22 · RSSM / Dreamer-style world model · node:test

pausedevidence: inconclusive

MaxEnt actor-critic for portfolio allocation

A reproducible Soft Actor-Critic research harness for long-only allocation. Built properly, measured against baselines, and not yet beating them.

Python · PyTorch · Gymnasium-style env · SAC (twin Q, auto-entropy)

§ B

Lab bench

Local-only repos don’t get links, so their numbers stay one-liners with caveats rather than headline claims.

  • planner-executor

    Python · MLX · Claude CLI

    Frontier model plans and rescues, a local Qwen 9B on MLX executes behind per-step JSON gates. On a 15-document extraction pilot it matched all-frontier accuracy at ~56% of the serving cost.

    evidence: preliminaryn = 15 docslocal only
  • Fortuna

    SwiftUI · FoundationModels · SwiftData

    My own iPhone finance app, and yes, I use it daily. An on-device Apple Foundation Model answers budget questions through 5 finance tools; card statements import straight from PDF. No servers, no dependencies.

    evidence: in daily use63 testslocal only
  • constitutional-guardrail-gateway

    FastAPI · pydantic · YAML policy

    OpenAI-compatible gateway that enforces a written 10-principle constitution as policy-as-code, with PII / injection detectors and a hash-chained audit log. Red-team eval designed, not yet run.

    evidence: self-check62 testslocal only
  • model-router-alpha

    scikit-learn · MiniLM · FastAPI

    Routes each query to the cheapest Claude tier that clears a quality bar. The eval design is the point: calibrated classifiers, bootstrap CIs, a κ-checked LLM judge.

    evidence: synthetic only830-prompt test sets
  • PASSAGE

    LangGraph · pydantic

    KYC/AML onboarding for a fictional bank: a 4-stage LangGraph pipeline with schema-checked handoffs, a rules-based risk engine, and a human review step.

    evidence: synthetic onlyoffline core34 tests
  • RL from scratch

    PyTorch · Gymnasium · React

    Single-file PPO (CartPole, Pendulum) and a Dreamer-v1-style world model with a live React training view. The PPO results are seed-sensitive, and I report them that way.

    evidence: measuredseed-sensitive

§ C

Up next

  • rag-fin-qa

    planned

    Q&A over SEC filings and earnings calls with GPU-served retrieval and faithfulness evals. Scaffolded; the latency and faithfulness targets are targets, not results.

    RAG · Ragas · Triton

  • cuda-quant-kernels

    planned

    Hand-written CUDA kernels for quant workloads, profiled with Nsight — my own kernels this time, not the samples.

    CUDA · Nsight