All work04 / 2026

Reinforcement learning · simulation

AutoDoseRL

Learning to stay in range.

A PPO controller adjusting insulin from glucose history, with explicit safeguards. Built and evaluated in simulation.

Take a closer look
PythonPyTorchStable-Baselines3Gymnasiumsimglucose

Saved simulation results

The average isn’t everyone.

Time in range 70–180 mg/dL

PPO policy88.7%
PID baseline59.0%
Basal-bolus baseline95.2%
1.0%PPO time below range
Below 70 mg/dL

Ten virtual adults, unseen episodes from the training cohort. Basal-bolus has meal/carbohydrate information that PPO does not receive. This comparison uses saved results; it runs no medical controller.

01 / The question

What can a controller learn without meal announcements?

02 / What I built

I built a continuous-action PPO controller using rolling glucose readings, insulin-on-board context, and time features. An explicit constraint layer sits between the learned action and the simulator.

One generalized policy was trained across ten adult virtual patients and evaluated on unseen episodes from that same cohort. The learned policy receives no meal announcements.

03 / What happened

The saved evaluation gives 88.7% mean time in range and 1.0% time below range. PID reaches 59.0% time in range; basal-bolus reaches 95.2%, with meal/carbohydrate information unavailable to the RL policy.

04 / Where it stops

These are simulator results, not clinical validation. Unseen episodes are not unseen patients. The 1.0% average time below range includes about 7.1% for adult #007, so the mean does not describe every patient. Baseline information differences matter to the comparison.

05 / Keep exploring

One more? / 05

Visual robustness