Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal
Modal proposes decoupling RL rollout workers from trainer clusters to use distributed GPU capacity across datacenters.
“IO wants all four of these at the same time. Enough GPU, same region, fast fabric, and available now. Any of these like is manageable, but all four of them that are pretty hard to get at the same time.”
Modal engineer Nan Jiang identifies a structural mismatch in RL post-training: the standard loop requires a tightly coupled single cluster with RDMA for fast weight sync, but available GPU capacity is distributed across providers and regions. He proposes decoupling the rollout/sampling fleet from the trainer so cheaper, scattered compute can handle trajectory generation while the tightly coupled cluster focuses only on training updates. This framing — 'cathedral' for trainers, 'bazaar' for rollouts — has practical implications for teams trying to scale RL post-training without booking monolithic reserved clusters.