Post-training infrastructure at Proximal
At Proximal, we believe that post-training is the best way to evaluate how our data affects model performance. It gives us the ability to experiment with various methods to train on more exotic forms of data and understand what models can and cannot learn from them. These experiments run in bursts, which requires us to have readily available infrastructure to support our on-demand experimental needs.
As such, we built a training stack around on-demand GPUs via Modal and internal sandbox infrastructure capable of supporting the scale we operate at. Training and inference run on GPUs from Modal, where we can add inference capacity during a run without restarting training, while rollouts run in sandboxes on a Kubernetes cluster. In this post, we highlight the design choices behind both halves, then show them at work training Qwen3.8-27B on agentic coding tasks, raising DeepSWE Pass@1 by 8.4 points in 44 steps.
System overview
LoRA weights over HTTPS, after each step
↺ LoRA weights go back to every replica over HTTPS after each step
Training on Modal
We built our training stack off an internal fork of Miles which is a popular open-source RL framework based on Slime. When first designing our stack, we started with 64 B300s of reserved capacity on Modal. We quickly realized we needed something that could scale beyond 64. To ensure high training throughput, we opted to use single-replica on-demand GPU containers rather than full 8-GPU nodes, as the latter added latency if we were unable to get a node scheduled. The on-demand replicas are subject to preemption. As such, we needed a design that was fault-tolerant and could scale to an arbitrarily large number of GPUs, while still preserving optimal inference routing between them and keeping training as efficient as possible. In this section, we detail the design choices we made.
Asynchronous training
We first decided to have our training replicas within the reserved-capacity cluster to optimize for data-parallel communication. Depending on the workload, we then allocate the remaining GPUs in the cluster to inference, supplementing them with single-replica on-demand GPUs.
Each rollout starts and ends on a single inference replica. Because we train LoRA adapters, a replica can serve several adapter versions at once. As soon as a new adapter is synced and ready, new rollouts start on it while older ones finish on the adapter they started with.
Asynchronous training
- 01LaunchNew groups start on the newest adapter and keep it to the end
- 02CollectFinished groups go to the rollout store
- 03TrainEach step takes 128 finished groups
- 04PublishReplicas load the new adapter next to the older ones still in use
Capturing what was sampled
To train on a rollout, you need the exact token IDs the model saw and produced on every call, plus the log-probability of each generated token. By default, Miles runs capture as a separate service that the training job starts alongside the trainer, in front of the inference engines. That assumes the engines are nearby since they are typically within the same cluster, and ours are not guaranteed to be.
In one of our experiments, Modal placed the trainer node in Sweden (eu-north-1) while the model replicas were in US West, so every call crossed the Atlantic and back. Over the course of a rollout, which makes on the order of 100 model calls, that is a few minutes of pure network wait per rollout. As such, we moved capture inside the inference replica, which removes this extra hop.
Growing inference during a run
During RL, a majority of the compute is in inference. The ideal ratio between inference and trainer replicas to maximize trainer utilization depends on various constantly changing factors. For example, if the model starts training on harder tasks later on in its curriculum, the number of zero-variance groups could potentially increase. This leads to fewer useable groups for RL which in turn requires more inference throughput to compensate. Hence, we designed our system with the ability to scale up inference replicas as necessary, smoothly going from a 3:1 to a 4:1 inference-to-trainer ratio, allowing us to keep a mean staleness of 1.7 even at that scale.
- 01Warm upLoad the current adapter
- 02JoinReceive every later weight push
- 03ServeTake new rollouts
Routing
Requests can have very different response lengths and compute requirements, so we cannot simply rely on Modal's request-level routing. Instead, we implement a custom router based on a best guess of how large each request will grow. This is a static form of the session admission that MISA-T proposes for RL rollouts. We randomly query K replicas' KV-cache headroom and extrapolate each replica's live rollouts before sending the new request to the GPU with the most room, accounting for whether the replica is an 8-GPU reserved-capacity replica or a single on-demand GPU.
This allowed us to hit an extremely high cache hit rate of 99.6%. We monitor our inference workers for thrashing during our runs, which we define as a constant cycle of cache evictions. Over the course of the run, the median episode for a single GPU was only 80 seconds.
Fault tolerance
Suddenly losing serving capacity does not stop a run. An incident took out 40% of our inference capacity mid-step. However, the run recovered on its own with replacement containers. All 128 GPUs were serving again 19 minutes after the first one dropped out. The step that spanned the outage took 37 minutes, against 35 minutes for the steps before and after it.
requests running on each of 128 serving GPUs · 45 minutes
To ensure that finished rollouts are durable, we replace Miles's default in-memory buffer with a rollout store, and after every step we save both the checkpoint and the store to Modal Volumes. In the event of an incident, we are able to resume from the last saved step with its unused rollouts intact.
Weight sync
Given that the trainer and inference workers are not in the same RDMA domain, we needed an efficient and reliable way to sync the updated weights to both the reserved and on-demand inference workers. We opted to sync the weights over HTTPS, as it was faster than writing them to a shared Modal Volume. We sharded the transfer over multiple ranks to use more of the available bandwidth, which further improved performance when the inference workers were busy handling inference requests. Lastly, we experimented with syncing only the weight deltas but found that the additional overhead incurred to reconstruct the weights was not worth it, given that LoRA adapters are already small.
Since replicas can join mid-run, we write the new weights to a Modal Volume for newly started replicas to read from. We start routing traffic to the new weight version as soon as one replica in each of our reserved and on-demand pools is serving it.
Training at long context
Much of our data involve rollouts at long context length, which requires us to use various strategies for efficient training. For example on Qwen 3.8 27B rollouts at max context length, keeping all activations, which native Miles supports, would require too much memory even with 4-way Tensor Parallelism on B300s. Instead, we implement a custom attention recomputation scheme on top of Miles in which we keep only FlashAttention's output and log-sum-exp from the first pass. This reduced our step time by 7.3%.
Inside a long-context training step
share of GPU time · 8 × B300 at TP 4 · samples averaging 231k tokens
Sandbox infrastructure
Since we are training on agentic coding tasks, we need to be able to run untrusted agent-written code in a secure, isolated fashion. Each rollout must have its own sandbox, and a single RL run can need thousands of them concurrently.
We operate a Kubernetes cluster with support for gVisor-based and microVM-based sandboxes with internet access disabled.
The agent harness runs independently and communicates via our lightweight command daemon (pxd) installed in the sandbox.
Bin packing
Sandboxes spend most of their time idle waiting for model inference. Our tasks declare between 1 and 8 vCPUs, with a mean of 2.3. CPU utilization is ~20% on average, and 100% at p99. Memory utilization is ~5% on average, with a p99 of 87%. We currently configure sandbox pods to request CPU at 25% of their hard limit and memory at 40% of their hard limit. Balancing this overcommitting for more full nodes and cache affinity, we pack an average of 26 sandboxes into 16-vCPU nodes without a single sandbox being evicted for node pressure.
Faster sandbox starts
We prioritize scheduling sandboxes on nodes that share the image over maximizing bin packing to optimize sandbox startup time. Image-aware scheduling raised our cache-hit rate from about 1% of sandbox launches to about 37%. An image cache hit reduces sandbox startup time from minutes to only 2 seconds.
Sandbox Scheduling
image affinity and cache reuse
Sandbox spin-up is mostly dictated by two things: how many cache hits we can get and how fast we can unpack an image in the case we don't. The former is solved by the scheduler affinity, but the latter is a combination of a few optimizations. One of the biggest optimizations we made was using zstd instead of gzip for image compression, which speeds up image unpacking by 15-25%. Cold starts scale with total image size as the compressed images need to be downloaded and decompressed. Cache hits are constant with respect to image size as nothing needs to be downloaded or moved.
Sandbox starts
ours vs SOTA provider · 20 images · both on gVisor
Latency
The difference in command latency is mainly determined by the time it takes for our agent's request to reach the sandbox. Our advantage in latency comes as an added benefit from owning our infrastructure.
Lower command latency reduces rollout duration, which means that rollouts can be more on-policy, increasing training efficiency and throughput.
Command latency
one command or no-op in a running sandbox · ours vs SOTA provider · same image · both on gVisor
Case study: Qwen3.8-27B on DeepSWE
To validate our infrastructure, we post-trained Qwen3.8-27B on an internal set of agentic coding tasks. We trained a rank-32 LoRA with 128 environments per step and a group size of 8 using vanilla Group Relative Policy Optimization (GRPO). We categorized our tasks based on difficulty and ordered them from easy to hard, with the ability to retry tasks with no passing attempts later in training.
We evaluated every checkpoint on DeepSWE with eight attempts on each of its 113 tasks, comparing each with the base model on the same tasks. We followed DeepSWE v1.1's methodology, which grades only what the agent explicitly commits.
DeepSWE
pass@1 · 113 tasks × 8 attempts
During our run, the model's Pass@1 on the full 113-task DeepSWE set climbed to 37.2%, 8.4 percentage points above the base model. Pass@8, the chance that at least one of eight attempts solves a task, increased from 67.7% to 86.6%.
Flawed tasks
To get a clearer understanding of the model's underlying capabilities, we wanted a fair test set for evaluation. DeepSWE has known flaws: some tasks produce false negatives, and some evaluate the model on factors outside the task specification. Epoch AI found issues in at least 23 of its 113 tasks, and Scrim found defects in 37, naming 13 of them. Our own internal review flagged 56.
Dropping any one list's flagged tasks leaves the pass-rate gain intact and even more significant. This highlights the importance of having a fair test set to understand the actual capabilities of a model.
Gains after dropping flawed tasks
DeepSWE · gain over the base model in points · all 8 attempts per task
The median DeepSWE task has 44 target tests (that must go from failing to passing) and 165 existing tests (that must remain passing). Thus, to have a clearer performance measure during training, we judge performance with partial credit as a binary pass fails an attempt over a single flawed test. The target-test score is the share of a task's target tests that an attempt passes. It rose from 56.1% for the base model to 78.5%, and from 52.8% to 74.6% when an attempt that breaks any existing test scores zero.
Partial credit on DeepSWE
| Checkpoint | Target-test score | Strict-partial |
|---|---|---|
| Base | 56.1% | 52.8% |
| Step 5 | 59.6% | 56.8% |
| Step 10 | 63.8% | 60.4% |
| Step 16 | 71.3% | 67.7% |
| Step 21 | 73.0% | 69.4% |
| Step 27 | 75.8% | 71.6% |
| Step 33 | 76.2% | 72.3% |
| Step 38 | 77.7% | 73.4% |
| Step 44 | 78.5% | 74.6% |
We build on open-source Miles and SGLang, and contribute back where we can: SGLang #43047 and Miles #3977.
At Proximal, we believe that it is critical to train models in order to push the frontier of data research. If any of the problems discussed in the blog are interesting to you, we are hiring.
Citation
Please cite this work as:
@article{proximal2026posttraininginfra,
author = {Evan Chu and Wei Hern Lim and Manasbir Bagri and Akira Yoshiyama and Brendan Graham},
title = {{Post-training infrastructure at Proximal}},
journal = {Proximal Blog},
year = {2026},
url = {https://www.proximal.ai/blog/posttraining-infra/},
}