Post-training infrastructure at Proximal

At Proximal, we believe that post-training is the best way to evaluate how our data affects model performance. It gives us the ability to experiment with various methods to train on more exotic forms of data and understand what models can and cannot learn from them. These experiments run in bursts, which requires us to have readily available infrastructure to support our on-demand experimental needs.

As such, we built a training stack around on-demand GPUs via Modal and internal sandbox infrastructure capable of supporting the scale we operate at. Training and inference run on GPUs from Modal, where we can add inference capacity during a run without restarting training, while rollouts run in sandboxes on a Kubernetes cluster. In this post, we highlight the design choices behind both halves, then show them at work training Qwen3.8-27B on agentic coding tasks, raising DeepSWE Pass@1 by 8.4 points in 44 steps.

System overview

Modal · SGLangInference
Kubernetes · gVisorSandboxes
Modal · MegatronTrainer

LoRA weights over HTTPS, after each step

↺ LoRA weights go back to every replica over HTTPS after each step

Figure 1. Each rollout runs its tool calls in a gVisor sandbox on Kubernetes and its model calls on SGLang replicas on Modal. A capture server on each replica records the exact tokens and logprobs that were sampled. Graded trajectories go to the trainer, which sends the updated LoRA adapter back to every replica after each step.

Training on Modal

We built our training stack off an internal fork of Miles which is a popular open-source RL framework based on Slime. When first designing our stack, we started with 64 B300s of reserved capacity on Modal. We quickly realized we needed something that could scale beyond 64. To ensure high training throughput, we opted to use single-replica on-demand GPU containers rather than full 8-GPU nodes, as the latter added latency if we were unable to get a node scheduled. The on-demand replicas are subject to preemption. As such, we needed a design that was fault-tolerant and could scale to an arbitrarily large number of GPUs, while still preserving optimal inference routing between them and keeping training as efficient as possible. In this section, we detail the design choices we made.

Asynchronous training

We first decided to have our training replicas within the reserved-capacity cluster to optimize for data-parallel communication. Depending on the workload, we then allocate the remaining GPUs in the cluster to inference, supplementing them with single-replica on-demand GPUs.

Each rollout starts and ends on a single inference replica. Because we train LoRA adapters, a replica can serve several adapter versions at once. As soon as a new adapter is synced and ready, new rollouts start on it while older ones finish on the adapter they started with.

Asynchronous training

Modal · MegatronTrainer
Each replicaAdapters
In-flight budgetRollouts
A rollout keeps the adapter it started on, older to newerMore than three versions behind: never trained on
  1. 01LaunchNew groups start on the newest adapter and keep it to the end
  2. 02CollectFinished groups go to the rollout store
  3. 03TrainEach step takes 128 finished groups
  4. 04PublishReplicas load the new adapter next to the older ones still in use
Figure 2. Generation and training run at the same time. As rollouts finish, new groups start on the newest published weights which keeps the budget busy. A rollout keeps the adapter it started on, so the same replica serves several adapters at once.

Capturing what was sampled

To train on a rollout, you need the exact token IDs the model saw and produced on every call, plus the log-probability of each generated token. By default, Miles runs capture as a separate service that the training job starts alongside the trainer, in front of the inference engines. That assumes the engines are nearby since they are typically within the same cluster, and ours are not guaranteed to be.

In one of our experiments, Modal placed the trainer node in Sweden (eu-north-1) while the model replicas were in US West, so every call crossed the Atlantic and back. Over the course of a rollout, which makes on the order of 100 model calls, that is a few minutes of pure network wait per rollout. As such, we moved capture inside the inference replica, which removes this extra hop.

Growing inference during a run

During RL, a majority of the compute is in inference. The ideal ratio between inference and trainer replicas to maximize trainer utilization depends on various constantly changing factors. For example, if the model starts training on harder tasks later on in its curriculum, the number of zero-variance groups could potentially increase. This leads to fewer useable groups for RL which in turn requires more inference throughput to compensate. Hence, we designed our system with the ability to scale up inference replicas as necessary, smoothly going from a 3:1 to a 4:1 inference-to-trainer ratio, allowing us to keep a mean staleness of 1.7 even at that scale.

Reserved nodes · always on
On-demand replicas · one GPU each
Joining mid-run
Adapter archive · Modal VolumeNew replicas start from the current adapter
  1. 01Warm upLoad the current adapter
  2. 02JoinReceive every later weight push
  3. 03ServeTake new rollouts
Figure 3. Reserved 8-GPU nodes stay up for the whole run. The single-GPU pool is able to grow during the run. A new replica starts from the current adapter, receives every later weight push and takes new rollouts.

Routing

Requests can have very different response lengths and compute requirements, so we cannot simply rely on Modal's request-level routing. Instead, we implement a custom router based on a best guess of how large each request will grow. This is a static form of the session admission that MISA-T proposes for RL rollouts. We randomly query K replicas' KV-cache headroom and extrapolate each replica's live rollouts before sending the new request to the GPU with the most room, accounting for whether the replica is an 8-GPU reserved-capacity replica or a single on-demand GPU.

This allowed us to hit an extremely high cache hit rate of 99.6%. We monitor our inference workers for thrashing during our runs, which we define as a constant cycle of cache evictions. Over the course of the run, the median episode for a single GPU was only 80 seconds.

Fault tolerance

Suddenly losing serving capacity does not stop a run. An incident took out 40% of our inference capacity mid-step. However, the run recovered on its own with replacement containers. All 128 GPUs were serving again 19 minutes after the first one dropped out. The step that spanned the outage took 37 minutes, against 35 minutes for the steps before and after it.

requests running on each of 128 serving GPUs · 45 minutes

Reserved nodes · 3 of 4 killed
On-demand replicas · 21 of 96 killed
1,570 rollout attempts relaunched
Figure 4. Each row is one serving GPU and each cell 20 seconds. A red cell is a GPU whose replica has stopped logging or has been killed and has not yet been replaced: 21 of the 96 single-GPU replicas and 3 of the 4 reserved nodes. The bars show the 1,570 rollout attempts launched in these 45 minutes.

To ensure that finished rollouts are durable, we replace Miles's default in-memory buffer with a rollout store, and after every step we save both the checkpoint and the store to Modal Volumes. In the event of an incident, we are able to resume from the last saved step with its unused rollouts intact.

Weight sync

Given that the trainer and inference workers are not in the same RDMA domain, we needed an efficient and reliable way to sync the updated weights to both the reserved and on-demand inference workers. We opted to sync the weights over HTTPS, as it was faster than writing them to a shared Modal Volume. We sharded the transfer over multiple ranks to use more of the available bandwidth, which further improved performance when the inference workers were busy handling inference requests. Lastly, we experimented with syncing only the weight deltas but found that the additional overhead incurred to reconstruct the weights was not worth it, given that LoRA adapters are already small.

Since replicas can join mid-run, we write the new weights to a Modal Volume for newly started replicas to read from. We start routing traffic to the new weight version as soon as one replica in each of our reserved and on-demand pools is serving it.

Training at long context

Much of our data involve rollouts at long context length, which requires us to use various strategies for efficient training. For example on Qwen 3.8 27B rollouts at max context length, keeping all activations, which native Miles supports, would require too much memory even with 4-way Tensor Parallelism on B300s. Instead, we implement a custom attention recomputation scheme on top of Miles in which we keep only FlashAttention's output and log-sum-exp from the first pass. This reduced our step time by 7.3%.

Inside a long-context training step

share of GPU time · 8 × B300 at TP 4 · samples averaging 231k tokens

One training stepshare of GPU time
One attention layerkernel time, to scale
Full recomputekeeps nothing
828 ms
Selective recomputekeeps Q, K, V per layer
out of memory
Recompute + output cacheours · keeps outputs
687 ms

Skipping the second forward removes 141 of 828 ms per layer. Over whole steps that measured 7.3% less time, for about 14 GiB more peak memory.

Figure 5. The top bar is a profile of one steady-state training step. The rows below are FA4's kernel time in one attention layer at 248k tokens. Full recompute runs the forward pass twice. Megatron's selective recompute avoids that by keeping every layer's inputs. We keep only the outputs.

Sandbox infrastructure

Since we are training on agentic coding tasks, we need to be able to run untrusted agent-written code in a secure, isolated fashion. Each rollout must have its own sandbox, and a single RL run can need thousands of them concurrently.

We operate a Kubernetes cluster with support for gVisor-based and microVM-based sandboxes with internet access disabled.

The agent harness runs independently and communicates via our lightweight command daemon (pxd) installed in the sandbox.

Bin packing

Sandboxes spend most of their time idle waiting for model inference. Our tasks declare between 1 and 8 vCPUs, with a mean of 2.3. CPU utilization is ~20% on average, and 100% at p99. Memory utilization is ~5% on average, with a p99 of 87%. We currently configure sandbox pods to request CPU at 25% of their hard limit and memory at 40% of their hard limit. Balancing this overcommitting for more full nodes and cache affinity, we pack an average of 26 sandboxes into 16-vCPU nodes without a single sandbox being evicted for node pressure.

Faster sandbox starts

We prioritize scheduling sandboxes on nodes that share the image over maximizing bin packing to optimize sandbox startup time. Image-aware scheduling raised our cache-hit rate from about 1% of sandbox launches to about 37%. An image cache hit reduces sandbox startup time from minutes to only 2 seconds.

Sandbox Scheduling

image affinity and cache reuse

Figure 6. Sandbox requests enter from the left where the letters identify their respective container images. Scheduling favors nodes with the image cached or already in use and then looks at the most full node for its reservation.

Sandbox spin-up is mostly dictated by two things: how many cache hits we can get and how fast we can unpack an image in the case we don't. The former is solved by the scheduler affinity, but the latter is a combination of a few optimizations. One of the biggest optimizations we made was using zstd instead of gzip for image compression, which speeds up image unpacking by 15-25%. Cold starts scale with total image size as the compressed images need to be downloaded and decompressed. Cache hits are constant with respect to image size as nothing needs to be downloaded or moved.

Sandbox starts

ours vs SOTA provider · 20 images · both on gVisor

New imagemedian start by image size, seconds
Cached imageall 20 images, seconds

Whiskers and shading are 95% intervals, bootstrapped over images. Each point is the median of a group of four images, placed at their median size.

Figure 7. We compare against a SOTA sandbox provider on the same 20 images on gVisor, timed from the create request to the first command (sh -c true) finishing. We did 5 trials for cached images and 1 for cold starts as we do not have control over the provider's cache. Our setup is faster in the new image and cached cases with lower variance in speed.

Latency

The difference in command latency is mainly determined by the time it takes for our agent's request to reach the sandbox. Our advantage in latency comes as an added benefit from owning our infrastructure.

Lower command latency reduces rollout duration, which means that rollouts can be more on-policy, increasing training efficiency and throughput.

Command latency

one command or no-op in a running sandbox · ours vs SOTA provider · same image · both on gVisor

Command, median
26 ms
130 ms
Command, 90th percentile
38 ms
287 ms
No-op, median
2 ms
64 ms
No-op, 90th percentile
3 ms
143 ms

Whiskers show 95% intervals, bootstrapped over each setup's 20 sandboxes.

Figure 8. The same SOTA sandbox provider as in Figure 7, in its default region. We measured the time to run one command (sh -c true) and a no-op in a sandbox that is already up. Each side made 600 calls of each kind to 20 sandboxes of the same gVisor image. Our setup is about 5x faster at the median and over 7x at the 90th percentile, with lower variance.

Case study: Qwen3.8-27B on DeepSWE

To validate our infrastructure, we post-trained Qwen3.8-27B on an internal set of agentic coding tasks. We trained a rank-32 LoRA with 128 environments per step and a group size of 8 using vanilla Group Relative Policy Optimization (GRPO). We categorized our tasks based on difficulty and ordered them from easy to hard, with the ability to retry tasks with no passing attempts later in training.

We evaluated every checkpoint on DeepSWE with eight attempts on each of its 113 tasks, comparing each with the base model on the same tasks. We followed DeepSWE v1.1's methodology, which grades only what the agent explicitly commits.

DeepSWE

pass@1 · 113 tasks × 8 attempts

Figure 9. Pass@1 by training step for every checkpoint we evaluated, compared with the base model on the same 113 tasks. The dashed line is the base model. Hover or tap a step to read its score.

During our run, the model's Pass@1 on the full 113-task DeepSWE set climbed to 37.2%, 8.4 percentage points above the base model. Pass@8, the chance that at least one of eight attempts solves a task, increased from 67.7% to 86.6%.

Flawed tasks

To get a clearer understanding of the model's underlying capabilities, we wanted a fair test set for evaluation. DeepSWE has known flaws: some tasks produce false negatives, and some evaluate the model on factors outside the task specification. Epoch AI found issues in at least 23 of its 113 tasks, and Scrim found defects in 37, naming 13 of them. Our own internal review flagged 56.

Dropping any one list's flagged tasks leaves the pass-rate gain intact and even more significant. This highlights the importance of having a fair test set to understand the actual capabilities of a model.

Gains after dropping flawed tasks

DeepSWE · gain over the base model in points · all 8 attempts per task

Figure 10. Each line drops the tasks one source flags as flawed: our internal review's, Epoch AI's, or the 13 tasks Scrim's write-up names out of the 37 it reports. All tasks is the full benchmark. Hover or tap a step to read every line; hover a line or tap its legend entry to pick it out.

The median DeepSWE task has 44 target tests (that must go from failing to passing) and 165 existing tests (that must remain passing). Thus, to have a clearer performance measure during training, we judge performance with partial credit as a binary pass fails an attempt over a single flawed test. The target-test score is the share of a task's target tests that an attempt passes. It rose from 56.1% for the base model to 78.5%, and from 52.8% to 74.6% when an attempt that breaks any existing test scores zero.

Partial credit on DeepSWE

CheckpointTarget-test scoreStrict-partial
Base56.1%52.8%
Step 559.6%56.8%
Step 1063.8%60.4%
Step 1671.3%67.7%
Step 2173.0%69.4%
Step 2775.8%71.6%
Step 3376.2%72.3%
Step 3877.7%73.4%
Step 4478.5%74.6%
Figure 11. Target-test score: target tests passed ÷ target tests. Strict-partial: the same, but 0 if any existing test broke.

We build on open-source Miles and SGLang, and contribute back where we can: SGLang #43047 and Miles #3977.

At Proximal, we believe that it is critical to train models in order to push the frontier of data research. If any of the problems discussed in the blog are interesting to you, we are hiring.


Citation

Please cite this work as:

@article{proximal2026posttraininginfra,
  author  = {Evan Chu and Wei Hern Lim and Manasbir Bagri and Akira Yoshiyama and Brendan Graham},
  title   = {{Post-training infrastructure at Proximal}},
  journal = {Proximal Blog},
  year    = {2026},
  url     = {https://www.proximal.ai/blog/posttraining-infra/},
}