Skip to content

case study

MaxEnt actor-critic for portfolio allocation

A reproducible Soft Actor-Critic research harness for long-only allocation. Built properly, measured against baselines, and not yet beating them.

pausedevidence: inconclusive
period
Apr 2026 — May 2026
status as of
Sep 23, 2026
size
~1.5k Python · 13 tests
stack
Python · PyTorch · Gymnasium-style env · SAC (twin Q, auto-entropy) · YAML configs · pytest
links
Repository ↗

tl;dr

A SAC-style max-entropy agent that allocates daily across SPY, QQQ, TLT, GLD and cash. The point is the harness: seeded runs, config snapshots, a six-way ablation sweep, and baselines evaluated the same way as the agent. On real data its validation Sharpe (2.00) edges past equal-weight (1.67) and random (1.75). That's a single seed with a worse drawdown, so I don't read it as an edge.

Problem

RL papers for portfolio allocation often report one striking backtest. What's usually missing is what's needed to believe it: fixed seeds, the exact config, ablations that show which ingredient matters, and naive baselines scored on the same split with the same metrics. I wanted a harness where an "RL beats the market" result would have to survive all of that first.

Approach

  • Environment: daily, long-only, no leverage, with an explicit cash bucket. The observation is a rolling feature window plus current weights. Actions are target weights, projected onto the simplex with a softmax. Reward is log growth after transaction costs, with optional volatility and drawdown penalties.
  • Agent: SAC with a tanh-squashed Gaussian actor, twin Q-networks with target networks, and automatic entropy tuning. Version 2 adds validation-driven checkpoint selection and benchmark-relative rewards.
  • Discipline: every run writes its config snapshot and seed. The ablation runner applies six overrides: entropy off, low and high transaction cost, risk penalty on, short and long window. Buy-and-hold, periodic rebalance and a random policy are scored on the same split.

Architecture

maxent-actorcritic-portfolio pipelineA YAML config and seed, optionally modified by the ablation runner, configure a data provider (synthetic or Yahoo) and a long-only portfolio environment with a cash bucket and transaction costs. A SAC agent with twin Q-networks and automatic entropy tuning trains against it. An evaluator scores the agent and three baselines (buy-and-hold, periodic rebalance, random) on the same split and writes results plus a config snapshot into a per-run directory.weightsrewardsame splitAblation runner6 overridesSAC agenttwin Q · auto-entropyYAML config + seedsnapshotted per runData providersynthetic | YahooPortfolio envlong-only · cash · costsEvaluatorSharpe · MDD · turnoverBaselinesB&H · rebalance · random
fig. 1 — Config-driven training, ablation, and evaluation pipeline

Results

Real market data, validation split, seed 7:

PolicySharpeMax drawdownAnn. returnTurnover
SAC agent2.00−8.0%18.2%0.007
Random policy1.75−4.6%18.0%0.865
Equal-weight (buy & hold)1.67−4.4%14.1%0

What’s next

  • Multi-seed runs with confidence intervals, then the held-out test split, once.
  • Re-run the ablation sweep on real data. The synthetic generator was there to prove the harness, not the strategy.
  • Walk-forward evaluation across regimes before any claim about skill.