This project formulated pandemic intervention planning as an infinite-horizon Markov Decision Process over Jakarta, Samarinda, Makassar, and Sorong. In a team of three, I helped implement a stochastic Gymnasium environment with city-level SEIRD dynamics, economic and budget variables, and independent lockdown, travel-restriction, and vaccination controls.
The main environment has a 22-dimensional continuous state and approximately 5.3 million joint actions. Reduced formulations with up to 4,096 states were also evaluated using Value Iteration, Policy Iteration, and Q-Learning. PPO, A2C, and a restricted uniform-national DQN were trained for 50,000 environment steps. PPO produced the best held-out mean return of 32.36 ± 12.98 across 30 seeds and eradicated the outbreak in all 50 descriptive simulations while retaining 44.65% of the intervention budget.