It can't tell you if they work.
Hotpath is an AI agent that makes your code faster and proves every change is correct. It profiles, proposes, tests and benchmarks — and keeps only the changes that are both correct and measurably faster.
+ if not self._compiled and args[0].is_cuda: + # Compile the single-token decode path with CUDA graphs. + self.blocks = nn.ModuleList([torch.compile(b, mode='reduce-overhead') for b in self.blocks]) + self._compiled = True
tokens match the original exactly
no change is accepted without evidence
The patches crashed or changed the output. Hotpath never times a patch that fails the locked check, so a broken speedup can't be counted as a win.
attempts that broke correctness
Its 95% interval ran from 0.99 to 1.02. On this H100 the run-to-run noise was 1.1%, so the bar was 1.03× and an interval that excludes 1.0.
attempts not measurably faster
changes kept on a model's word alone
of AI-proposed attempts not shipped
Across 51 attempts on a tiny stand-in transformer, Hotpath shipped 3 changes worth 1.46× decode — every one correct and faster than the noise.
Like a scientist: form a hypothesis, run the experiment, keep the change only if the data supports it. Two steps are models. Four are plain, deterministic code the models cannot edit.
Where does the time actually go?
Runs the target under a profiler and summarises the widest bars. Amdahl's law: a function that takes 5% of runtime can only ever save 5%.
top: gemv kernel 12.3%
aten::native_layer_norm 7.4%torch.profiler · cProfileWhat should we try next?
The planner reads the profile, the hottest source, and every past result, and returns structured experiments. It stops proposing what already failed.
{"idea": "build the causal mask once",
"target_file": "model.py", "risk": "low"}planner · OpenAI APICan the idea be written as a patch?
Fast worker models write each idea as a patch in its own isolated worktree, so many experiments run in parallel without touching each other.
worktree .hotpath/worktrees/exp_5a3154d8e6 patch model.pyworkers · git worktrees
Is the output still the same?
Generated tokens must match the frozen reference exactly and logits must stay within tolerance. Tests are locked: a patch that touches them is rejected before a byte is written.
ok tokens seed=4 plen=200 n_new=16 ok logits seed=4 max abs diff 5.96e-07locked test suite
Is it faster than the noise?
Warmup runs, GPU sync, many trials, the median. The speedup must clear the machine-measured noise floor, and its 95% interval must exclude 1.0.
threshold max(1.03, 1 + 2·1.1%) = 1.03 ci95 [1.350, 1.372]warmup · sync · bootstrap CI
Does it still help on top of the rest?
Accepted patches stack, and each new one is measured against the stack it builds on. The bottleneck moves, so Hotpath re-profiles. An ablation pass re-measures what each shipped change contributes.
stack sdpa → reshape → mask once next re-profilebeam · re-profile · ablate
A strong planner reads profiles and history and chooses what to try. Cheap, fast workers write many candidate patches. The harness applies each one in its own worktree, runs the locked tests and the benchmark, and returns a structured verdict.
optimization moves · 9 seeds + free-form
verdict statuses
harness verdicts: 3 shipped · 48 not shipped
The planner proposed torch.compile with CUDA graphs to cut launch overhead, and the worker wrote the patch. The locked test crashed: the graph overwrote its own outputs. It was rejected before a single benchmark ran, and the reason was saved.
Correctness comes first, speed second. Tokens must match the original under greedy decoding, and logits must stay within tolerance. A fused kernel that reorders floating-point math passes; a change that alters what the model says does not, however fast it is.
tests failed (exit 1) · RuntimeError: accessing tensor output of CUDAGraphs that has been overwritten by a subsequent run
planner · reasoningFor small models with many single-token decode calls, much time is lost to kernel launch overhead.
+ if not self._compiled and args[0].is_cuda:+ # Compile the single-token decode path with CUDA graphs.+ self.blocks = nn.ModuleList([torch.compile(b, mode='reduce-overhead') for b in self.blocks])+ self._compiled = True
Out of 51 attempts, 3 shipped. A high rejection rate isn't the agent failing — it's the evidence that verification was necessary. If the model is usually wrong, correctness can't be something you trust. It has to be something you measure.
GPU time in these ops fell only from 9.42 to 9.14 ms, yet decode throughput rose 1.46×. Four attention ops became one fused kernel that took more GPU time, not less, and the mask stopped being rebuilt every call. The profile does not account for most of the gain; launch and CPU overhead, which a GPU-time profile does not capture, is the likely remainder. That is why Hotpath lets the benchmark, never the profile, decide. Re-checked afterwards outside Hotpath with the same locked test and benchmark: 904 → 1376 tokens/sec.
The run above is a stand-in model we wrote. This one is not: hotpath go against jaraco/inflect, a library nobody here wrote, with no human step between the URL and the draft PR — 5 min 3 s.
One commit per verified change, each the exact tree the harness tested. The benchmark was generated by Hotpath over plural_noun, compare, ud_match, _plnoun, _plequal, declared noisy at ±13% — which raised the acceptance bar to about 26%. The accepted head measured 1.409× with a 95% CI of [1.31, 1.45]. Planner gpt-4.1, worker Kimi-K2.7-Code, 283,112 tokens across 13 calls.
PR #1 failed the repository's CI — our committed benchmark imported Hotpath, which that CI does not have. That was fixed; on the follow-up PR the repository's own CI read 6 passed, 6 already failing on the base, 22 with no base result to compare, and 0 introduced. “0 introduced” means no failure was shown to be ours — not that every check is green.
Property tests carry a 200 ms deadline, so the suite fails under load. A timing-dependent definition of correct cannot gate a timing experiment.
The only benchmark was the whole test suite at 8.1% noise, so the bar rose to ~16%. The best candidate looked 4.6% faster with a CI of [0.99, 1.07] — not kept, no PR.
No benchmark? Hotpath writes one over your hottest functions and validates it by running it. Or bring one config file: the command that proves it's correct, the command that times it, what the agent may edit, and what it must never touch.
locked is enforced by the harness, not the prompt: an agent rewarded for "tests pass and it's faster" will sometimes try to edit the test. That patch is rejected before it runs.