Author | Bi Weihao
Editor | Li Shuiqing

According to a report by Zhidongxi on September 23, the latest paper authored by Liang Wenfeng, founder of DeepSeek, has just been made public. It is the first to systematically disclose the technical details of DeepSeek's agent‑training sandbox platform, DSec (DeepSeek Elastic Compute). The paper was submitted on September 19, and its author list includes more than 130 names, with Liang Wenfeng among them.

Paper link:
https://arxiv.org/pdf/2609.22978
The DSec platform first appeared in the DeepSeek V4 technical report, where its primary function is to provide a sandbox for agent training, enabling the stable operation of large-scale agent training. The paper explicitly states that, from RL training and evaluation between DeepSeek V3.2 and V4.1, all sandbox workloads were run on DSec; now, by making the technical details public, it can be said that the "Fenjue" system has been unveiled.
The paper indicates that the DSec platform is extremely large-scale, with one production unit comprising approximately 160 CPU nodes, 30,000 cores, and 250 TB of memory, hosting petabyte‑scale images.
In terms of capacity, the DSec platform processes approximately 3 million sandboxes per day, with peak concurrent sessions exceeding 380,000, a creation rate of over 5,000 per second, and the ability to launch up to 32,000 sandboxes simultaneously for a single training task.

▲Distribution of the number of sandboxes created for each task
So the question arises: why is such a massive number of sandboxes required in agent training? And when confronted with complex execution environments, how does DSec address challenges related to scale, scheduling, and resource management?
01.
Reinforcement learning agents naturally require massive amounts of simulation environments.
The training environment has become a bottleneck for development.
Reinforcement learning in traditional LLM training can often be conducted using static input–output mappings and reward signals. However, agent training differs fundamentally from LLM training: it requires actual interaction with a real-world environment, involving concrete tasks such as reviewing code, invoking tools, executing commands, and modifying files. With each step the model takes, the environmental state may change, and subsequent actions depend on the outcomes of previous steps.
This means that, during agent training, in addition to the model and the data, researchers must also maintain a large number of "work environments"—simulations that closely resemble real-world machines, allowing for the installation of dependencies, the execution of software, and the ability to revert to a clean state after each task, ready for the next rollout.
The problem is that these sandboxes are both numerous and heavy.
The paper shows that during training, a single task once simultaneously launched 32,000 sandboxes, yet these sandboxes were never fully utilized. As the agent executed its tasks, the sandboxes often remained in a waiting state for the next step, resulting in low CPU utilization. However, even when the CPU is idle, this does not mean that resources have been released; memory and writable state must still be continuously maintained.

▲CPU usage is intermittent, while memory and system state remain continuously occupied.
Consequently, the traditional approach of "launching a container and running a single task" can no longer sustain itself. The DSec platform was designed to address, in a unified manner, sandbox‑level capabilities such as batch creation, resource scheduling, environment replication, state persistence, suspension and resumption, and secure isolation.
02.
Four major environmental requirements: a single SDK manages them all.
Functions, Containers, and Virtual Machines
As the tasks that agents are called upon to perform grow increasingly complex, the underlying operating environments can no longer be handled with a one-size-fits-all approach.
The simplest task may only require a single function call—execute the code and return the result—while software engineering tasks demand a full Linux user-space environment, capable of installing dependencies, modifying code, and running tests. Security‑focused offensive and defensive work, as well as general computer use, impose even stricter isolation requirements. And if you need to run certain commercial applications, the required environment may approach that of an entire standalone machine.
DSec offers four backend options—FnCall, containers, Firecracker microVMs, and full virtual machines—each addressing distinct needs, from short-lived function invocations and software engineering to security‑sensitive workloads and full OS environments. Meanwhile, the training framework need not concern itself with whether the underlying execution environment is a container or a virtual machine; via the Python SDK (libdsec), it can create sandboxes, execute commands, and retrieve results directly, without requiring adaptation to different runtime types.

▲DSec offers four backend options.
DSec's four backend components are centrally orchestrated by a unified platform. After the training framework submits a request, the platform first performs identity and permission verification, then selects an appropriate node based on cluster load; the Edge service running on that node is responsible for creating the sandbox. Once the sandbox is launched, Aether and Chronus manage the communication between the platform and the sandbox's internal execution pipeline, while image data is provisioned on demand by 3FS.

▲DSec Architecture
However, as these environments expand from a few types to tens of thousands of instances, new challenges emerge: how to rapidly replicate a vast number of diverse environments.
03.
The more the environment, the harder replication becomes:
How does DeepSeek enable the rapid deployment of tens of thousands of sandboxes?
As mentioned earlier, the environments used for training agents are not only numerous but also highly complex in their combinations.
The paper analyzes data from a single production week: the container backend encompasses 11,266 base images, 102,171 workspaces, and 103 toolkits. During runtime, 67.8% of sandboxes further layer additional workspaces or toolkits on top of the base images. DeepSeek Harness is one such component that requires frequent updates.

If all these components are packaged into a single, monolithic image, any change at any layer could necessitate rebuilding and redistributing the entire image. As the number of environments grows, the costs of image maintenance and deployment will likewise increase.
DSec splits the base image, workspace, and toolkit into three independent read-only EROFS layers, which are then combined at sandbox startup using overlayfs. This way, when a particular component changes, only the corresponding layer is updated, eliminating the need to rebuild the entire image.
Image distribution employs on-demand loading. The paper finds that, during sandbox execution, the data actually accessed accounts for only 4.2% to 13.3% of the full image. Accordingly, DSec stores image data on 3FS and reads it on demand at runtime, while prefetching metadata locally and persisting writes to the node's local disk. Given that 3FS is better suited for large‑block, sequential reads, this approach also mitigates the performance overhead associated with small‑block, random I/O.

▲DSec decomposes the environment into composable layers
The results show that when 8,192 containers are deployed simultaneously in a burst, on-demand loading takes 35 minutes, whereas Docker's cold pull takes over 60 minutes. Additionally, the cumulative disk write volume per node drops from approximately 1,600 GB to about 700 GB.
The paper demonstrates that the environment in the DSec platform can also be set up by an agent. Using pack_diff, after the agent configures the environment, it generates an incremental snapshot, which can subsequently be used to restore a new sandbox.
04.
Rollout removes the GPU:
Training and execution are separate.
In the early design, the agent's inference and rollout processes shared the GPU pod with model training. Once a GPU task was preempted, the ongoing rollout would also be forcibly interrupted.
Starting with version 4.1, DeepSeek decouples the rollout process from the GPU‑based training environment and runs it independently on the DSec platform. The agent sandbox manages execution environments such as DeepSeek Harness, while the worker container handles the actual tasks; neither relies on GPU resources. As a result, when GPU‑intensive training is preempted, the rollout state can be preserved independently.

▲DSec's CPU scheduling, memory reclamation, layered images, and on-demand loading mechanisms.
If the cluster lacks sufficient capacity, DSec also supports scaling to the cloud. For example, once cluster utilization exceeds 80%, eligible sandboxes can be migrated to cloud-based virtual machines. To minimize the overhead of re-pulling images from the cloud, DeepSeek pre‑provides a deduplicated image repository of approximately 30 TB, with about 70% of its files actually accessed by container workloads. In production, 200 cloud VMs can handle roughly 30% of peak load.
05.
The more realistic the environment is
The greater the risk associated with the agent's actions.
DSec addresses scalability and efficiency challenges, but real-world environments still harbor other risks—for example, agents may not always complete their tasks along the expected path.
The paper documents a variety of anomalous behaviors: some agents search through logs, forge RPC requests, and even modify /bin/bash in an attempt to bypass normal task workflows. Additionally, other agents scan reachable services and fetch external code, seeking answers via paths not covered by the evaluation.

Even more troublesome, agents can sometimes corrupt the environment itself. The paper documents cases where recursive scanning of system files caused the kernel to crash, and even a simple `yes` command can rapidly swell log files to tens of gigabytes.

To address these risks, DSec primarily restricts the Agent's operational scope using AppArmor and eBPF: the former controls file and socket access, while the latter limits network access, with rules that can be dynamically adjusted based on the task's phase.
However, these measures currently only mitigate a portion of the risks, and kernel‑level vulnerabilities remain difficult to fully prevent.
06.
Conclusion: DSec Platform Capabilities
Becoming a key component of large-scale agent training.
DSec has unveiled an infrastructure solution designed for large-scale agent training. From sandbox creation and environment reuse to rollout scheduling, state persistence, and secure isolation, agent training is giving rise to a distinct set of infrastructure requirements.
As agent tasks grow longer and interaction cycles multiply, the scale of the execution environment will continue to expand. Ensuring the stable operation of tens of thousands—or even more—sandboxes while controlling resource costs and mitigating security risks will be a critical challenge that must be addressed as agent training scales further.
For DeepSeek's agent training, as model capabilities continue to improve, the execution platform that supports these tasks must keep pace. The engineering solution proposed by DSec may well represent one facet of the evolution of agent infrastructure at this stage.