Hotpath
Performance agent · Hack the North 2026

AI can write optimizations.

It can't tell you if they work.

Hotpath is an AI agent that makes your code faster and proves every change is correct. It profiles, proposes, tests and benchmarks — and keeps only the changes that are both correct and measurably faster.

$
exp_364b9d9ad5 · attempt 6 of 51model.py
+ if not self._compiled and args[0].is_cuda:
+     # Compile the single-token decode path with CUDA graphs.
+     self.blocks = nn.ModuleList([torch.compile(b, mode='reduce-overhead') for b in self.blocks])
+     self._compiled = True
apply
locked
tests
bench
verdicts · 51 attempts
shipped3
passed, not in final stack7
rejected · speed18
rejected · correctness22
patch failed1
verified · decode
0.00×

tokens match the original exactly

accepted chain · tokens/sec
baseline
fused attention
no .contiguous()
decode skips mask*
no .item() sync*
mask built once
§ 01 — the problem

Performance work is profile, guess, rewrite, re-test, re-benchmark — dozens of times. AI speeds up the guessing. Hotpath closes the loop.

Plausible, but broken

CUDA graphs, autocast, fused kernels. All reasonable. All broke the model.

The patches crashed or changed the output. Hotpath never times a patch that fails the locked check, so a broken speedup can't be counted as a win.

0/51

attempts that broke correctness

Faster, but it's noise

TF32 measured 1.004× faster. That's noise.

Its 95% interval ran from 0.99 to 1.02. On this H100 the run-to-run noise was 1.1%, so the bar was 1.03× and an interval that excludes 1.0.

0/51

attempts not measurably faster

The Hotpath rule

Correct and faster than the noise. Nothing else ships.

1058
989
1059
1020
1409
1390
0/51shipped · 0 unverified
0

changes kept on a model's word alone

0%

of AI-proposed attempts not shipped

Across 51 attempts on a tiny stand-in transformer, Hotpath shipped 3 changes worth 1.46× decode — every one correct and faster than the noise.

measured on an NVIDIA H100 80GB · gpt-4.1 planner, Kimi-K2.7-Code worker · full report
§ 02 — the loop

Six steps. The AI may be wrong. The harness can't be talked into it.

Like a scientist: form a hypothesis, run the experiment, keep the change only if the data supports it. Two steps are models. Four are plain, deterministic code the models cannot edit.

hotpath/harness.py · simplified
1harness

Profile

Where does the time actually go?

Runs the target under a profiler and summarises the widest bars. Amdahl's law: a function that takes 5% of runtime can only ever save 5%.

top: gemv kernel             12.3%
     aten::native_layer_norm   7.4%
torch.profiler · cProfile
2ai · may be wrong

Hypothesize

What should we try next?

The planner reads the profile, the hottest source, and every past result, and returns structured experiments. It stops proposing what already failed.

{"idea": "build the causal mask once",
 "target_file": "model.py", "risk": "low"}
planner · OpenAI API
3ai · may be wrong

Implement

Can the idea be written as a patch?

Fast worker models write each idea as a patch in its own isolated worktree, so many experiments run in parallel without touching each other.

worktree  .hotpath/worktrees/exp_5a3154d8e6
patch     model.py
workers · git worktrees
4the gate

Verify correctness

Is the output still the same?

Generated tokens must match the frozen reference exactly and logits must stay within tolerance. Tests are locked: a patch that touches them is rejected before a byte is written.

ok   tokens seed=4 plen=200 n_new=16
ok   logits seed=4 max abs diff 5.96e-07
locked test suite
5harness

Benchmark

Is it faster than the noise?

Warmup runs, GPU sync, many trials, the median. The speedup must clear the machine-measured noise floor, and its 95% interval must exclude 1.0.

threshold  max(1.03, 1 + 2·1.1%) = 1.03
ci95       [1.350, 1.372]
warmup · sync · bootstrap CI
6harness

Keep & compose

Does it still help on top of the rest?

Accepted patches stack, and each new one is measured against the stack it builds on. The bottleneck moves, so Hotpath re-profiles. An ablation pass re-measures what each shipped change contributes.

stack  sdpa → reshape → mask once
next   re-profile
beam · re-profile · ablate
The architecture

Big model plans. Fast model explores. The harness decides.

A strong planner reads profiles and history and chooses what to try. Cheap, fast workers write many candidate patches. The harness applies each one in its own worktree, runs the locked tests and the benchmark, and returns a structured verdict.

The loopaction / verdict
action · <patch>observation · verdict + reasonYour repoconfig.yamlHarnesssandboxed · models cannot edit itworktreelockedtestsbenchverdict + reason → historyPlanner + workersbig model plansfast model exploresproposes a patchVerified stackcorrect + faster

optimization moves · 9 seeds + free-form

fused SDPA attentionKV cache preallocationremove sync pointsbuild the causal mask oncetorch.compileCUDA graphsfused norm kernelsTF32speculative decoding

verdict statuses

acceptedrejected_speedrejected_correctnesspatch_failedlocked_filenot_selectedtimeouterror

Every rejection is recorded with its reason.

harness verdicts: 3 shipped · 48 not shipped

torch.compile with CUDA graphs
rejected · correctness · tests crashed: CUDA graph output overwritten
mixed-precision autocast
rejected · correctness · tests crashed: wrong autocast API
TF32 matmuls
rejected · speed · 1.004× vs parent · noise
drop the .item() sync
rejected · speed · 1.02× · below the 1.03× bar
preallocate the KV cache, again
patch failed · the patch changed nothing
build the causal mask once
kept · 1.36× vs parent · CI excludes 1.0
§ 03 — the trust moment

The agent proposed CUDA graphs. Hotpath threw it out.

The planner proposed torch.compile with CUDA graphs to cut launch overhead, and the worker wrote the patch. The locked test crashed: the graph overwrote its own outputs. It was rejected before a single benchmark ran, and the reason was saved.

See the harness
0attempts rejected on correctness
in this run, none of them timed
exit 1the locked test's verdict
5.96e-07max |Δlogit| of the shipped stack

Correctness comes first, speed second. Tokens must match the original under greedy decoding, and logits must stay within tolerance. A fused kernel that reorders floating-point math passes; a change that alters what the model says does not, however fast it is.

—torch.compile with CUDA graphs
rejected · correctness
harness · reason

tests failed (exit 1) · RuntimeError: accessing tensor output of CUDAGraphs that has been overwritten by a subsequent run

planner · reasoning

For small models with many single-token decode calls, much time is lost to kernel launch overhead.

diff
+ if not self._compiled and args[0].is_cuda:
+ # Compile the single-token decode path with CUDA graphs.
+ self.blocks = nn.ModuleList([torch.compile(b, mode='reduce-overhead') for b in self.blocks])
+ self._compiled = True
Why you need a verifier

Most AI-proposed optimizations fail. That's the point.

Out of 51 attempts, 3 shipped. A high rejection rate isn't the agent failing — it's the evidence that verification was necessary. If the model is usually wrong, correctness can't be something you trust. It has to be something you measure.

From idea to shipped changemeasured run · 51 attempts
Attempts (37 ideas + 14 retries)51Patches applied cleanly50 −1Passed correctness28 −22Faster than the noise10 −18Shipped in the final stack3 −7
0.0%shipped — correct and faster
0.0%broke correctness
0.0%not measurably faster
Results

965 → 1409 tokens/sec. Every step verified.

0.00×decode throughput vs baseline
0tokens/sec, best verified
3changes shipped
0changed outputs shipped
Throughput over the runmeasured · steps are accepted changes · * replaced by a faster lineage
90010001100120013001400150001020304050drop the .item() sync · within noisein-place MLP residual add · not selectedTF32 matmuls · within noisefused attentionno .contiguous()decode skips mask*no .item() sync*mask built once1409 tok/sexperimentstokens / sec
best verifiednoise band ±1.1%measured, not kept
Before · baseline profiletorch.profiler · 9.42 ms shown
aten ops, GPU self time · 9.42 msattention: 4 o…bmmmas…layer_normmmaddmmcatadd
After · head profiletorch.profiler · 9.14 ms shown
aten ops, GPU self time · 9.14 msattention: 1 fused ke…efficient_attentionlayer_normmmaddmmcatadd

The profile moved less than the benchmark.

GPU time in these ops fell only from 9.42 to 9.14 ms, yet decode throughput rose 1.46×. Four attention ops became one fused kernel that took more GPU time, not less, and the mask stopped being rebuilt every call. The profile does not account for most of the gain; launch and CPU overhead, which a GPU-time profile does not capture, is the likely remainder. That is why Hotpath lets the benchmark, never the profile, decide. Re-checked afterwards outside Hotpath with the same locked test and benchmark: 904 → 1376 tokens/sec.

mask build (tril, == 0) 0.80 ms → ≤ 0.10 ms · built oncecopy_ 0.33 ms → 0.06 ms · .contiguous() removedattention 4 ops · 2.19 ms → 1 kernel · 2.99 ms · fused SDPA
A real repository, a real pull request

A bare URL in. A verified pull request out.

The run above is a stand-in model we wrote. This one is not: hotpath go against jaraco/inflect, a library nobody here wrote, with no human step between the URL and the draft PR — 5 min 3 s.

1.472×vs baseline, 2 changes shipped
2 / 8candidates accepted
214existing tests, all green
0CI failures introduced (follow-up PR #2)
The pull requestmeasured · 2026-09-22
37767fddhotpath: set up benchmark, config, and verification check
c3e2d2f9perf: Reduce redundant isinstance checks in _plnoun by caching types
61f1edc7perf: Avoid repeated isinstance checks inside _plnoun for the common case

One commit per verified change, each the exact tree the harness tested. The benchmark was generated by Hotpath over plural_noun, compare, ud_match, _plnoun, _plequal, declared noisy at ±13% — which raised the acceptance bar to about 26%. The accepted head measured 1.409× with a 95% CI of [1.31, 1.45]. Planner gpt-4.1, worker Kimi-K2.7-Code, 283,112 tokens across 13 calls.

PR #1 failed the repository's CI — our committed benchmark imported Hotpath, which that CI does not have. That was fixed; on the follow-up PR the repository's own CI read 6 passed, 6 already failing on the base, 22 with no base result to compare, and 0 introduced. “0 introduced” means no failure was shown to be ours — not that every check is green.

And where it refusedno number is also a result
life4/textdistancestopped at the baseline

Property tests carry a 200 ms deadline, so the suite fails under load. A timing-dependent definition of correct cannot gate a timing experiment.

mahmoud/boltons7 candidates, 0 accepted

The only benchmark was the whole test suite at 8.1% noise, so the bar rose to ~16%. The best candidate looked 4.6% faster with a CI of [0.99, 1.07] — not kept, no PR.

§ 04 — quickstart

Any repo with a test suite.

No benchmark? Hotpath writes one over your hottest functions and validates it by running it. Or bring one config file: the command that proves it's correct, the command that times it, what the agent may edit, and what it must never touch.

$ pipx install hotpath-agent
$ hotpath go https://github.com/you/your-repo
configs/dryft_local.yaml

locked is enforced by the harness, not the prompt: an agent rewarded for "tests pass and it's faster" will sometimes try to edit the test. That patch is rejected before it runs.

HOTPATH