Reinforcement learning · simulation
AutoDoseRL
Learning to stay in range.
A PPO controller adjusting insulin from glucose history, with explicit safeguards. Built and evaluated in simulation.
Take a closer lookTime in range 70–180 mg/dL
Below 70 mg/dL
Ten virtual adults, unseen episodes from the training cohort. Basal-bolus has meal/carbohydrate information that PPO does not receive. This comparison uses saved results; it runs no medical controller.
01 / The question
What can a controller learn without meal announcements?
02 / What I built
I built a continuous-action PPO controller using rolling glucose readings, insulin-on-board context, and time features. An explicit constraint layer sits between the learned action and the simulator.
One generalized policy was trained across ten adult virtual patients and evaluated on unseen episodes from that same cohort. The learned policy receives no meal announcements.
03 / What happened
The saved evaluation gives 88.7% mean time in range and 1.0% time below range. PID reaches 59.0% time in range; basal-bolus reaches 95.2%, with meal/carbohydrate information unavailable to the RL policy.
04 / Where it stops
These are simulator results, not clinical validation. Unseen episodes are not unseen patients. The 1.0% average time below range includes about 7.1% for adult #007, so the mean does not describe every patient. Baseline information differences matter to the comparison.
05 / Keep exploring
One more? / 05