janus-chrysalis
Open research on world models in multi-agent RL, trained in TypeScript, co-authored with Claude, and published even when the effect disappears.
Can you measure how much a world model's prediction error comes from the other agent's learning, rather than from the model just being new? I built a freeze-intervention instrument to isolate that signal in a two-agent gridworld. The first 9 seeds looked significant (p ≈ 0.039). A pre-registered replication took it to p ≈ 0.388, so the write-up reports a null result. The work runs as a human + Claude collaboration: a daily autonomous agent opens PRs, and I review and merge them.
- Pre-registered replication
- p ≈ 0.388
- n = 12 seeds, 8 negative · vs p ≈ 0.039 at n = 9 · did not replicate
- G2 loss-curve gate
- 6 / 6 pass
- n = 3 seeds × 2 agents · total loss 2.2–2.9 → 1.4–1.8 over 20 episodes
- Test suite
- 165 passing
- n = 177 tests, 12 todo, 0 fail
- Autonomous loop
- 77 stand-ups
- n = Jul – Sep 2026
TypeScript · TensorFlow.js (tfjs-node) · Node 22 · RSSM / Dreamer-style world model · node:test