GitHub Repo
PyTorch Precompilation Proposal Packages Kernels, Strengthens No-Compile Checks During Replay
The author reports that the full set of changes let 16 distributed processes replay a recommendation model without triggering compilation. The proposal is still a draft; validation across models and deployment benefits remain to be confirmed.

On September 27, PyTorch developers updated a precompilation proposal that aims to package the kernels needed for replay. The author reports that the full set of changes let 16 distributed processes complete 100 batches on a large recommendation model using FSDP2 without triggering compilation. At the time of review, PR #198789 was still a draft, not an officially released feature. [Proposal and tests](https://github.com/pytorch/pytorch/pull/198789)
The issue is that loading a computation graph does not guarantee that the underlying kernels are ready. Existing official documentation divides caches into layers, including computation graphs, Triton compilation results, and autotuning. Mega-Cache can export and load these artifacts, but checks that the PyTorch and Triton versions and CUDA GPU match. The documentation also explains that autotuning benchmarks candidate kernels and selects the fastest one; tuning again still has a cost, and when reusing a cache across processes, the runtime environment must be checked against its constraints. [Official caching documentation](https://docs.pytorch.org/tutorials/recipes/torch_compile_caching_tutorial.html)
The new approach proposes three steps: `capture_runtime()` records the Triton kernels and C++ binaries actually launched; `finalize_cache()` freezes the kernel launchers and cache contents; and `prepare_runtime()` validates and prepares the kernels before model artifacts are loaded. The path that directly uses Triton JIT also depends on whether Triton provides a runtime cache export interface. [API design](https://github.com/pytorch/pytorch/pull/198789)
The accompanying design for disabling compilation makes computation graph compilation, kernel cache misses, and autotuning raise errors directly, preventing the replay workflow from silently compiling as a fallback. For engineering teams, this makes “is precompilation complete?” a checkable runtime condition. However, the initial companion PR was closed and superseded by a later proposal, so the interface may still change. [Companion design](https://github.com/pytorch/pytorch/pull/198767)
From a deployment perspective, binding selected kernels to model artifacts could help check whether different processes use the same compute path and expose missing artifacts earlier. Such checks are especially relevant to deployments that need predictable startup behavior and launch multiple worker processes at once. However, the current results are the author’s test on a single model and cannot be generalized into performance or numerical guarantees for all models. Next, teams should track upstream review, full CI, reproduction on other workloads, and whether caches cover different input shapes. Before production deployment, they should also measure startup time and cache size to confirm that the compilation savings are not offset by artifact transfer and loading overhead. [Scope of validation](https://github.com/pytorch/pytorch/pull/198789)