Documentation
Hotpath makes code faster and proves every change is correct. Install it, point it at a repository, and it comes back with a draft pull request of verified changes — or with the reason nothing was proved.
Install
Hotpath is a Python CLI. You need Python 3.11+ and Git; Docker and the GitHub gh CLI are optional.
pipx install hotpath-agent
hotpath doctor # Git, Docker, API keys, and the active acceleratorThe package is hotpath-agent; the command it installs is hotpath. The hotpath and hotpath-ai names on PyPI belong to unrelated projects. uv tool install hotpath-agent works too. For the latest unreleased code, pipx install git+https://github.com/Nijjea1/hotpath, or clone the repo and run hotpath.cmd (Windows) or ./hotpath.sh, which create a venv for you.
| You need | For | Without it |
|---|---|---|
| Python 3.11+, Git | everything | Hotpath will not start |
| OPENAI_API_KEY | real planner and worker models | use --provider mock (recorded patches, fully offline) |
| Docker | running untrusted repositories in a locked-down container | go asks before running code on your machine |
| gh or GITHUB_TOKEN | opening pull requests and reading CI | Hotpath pushes the branch and prints a pre-filled PR link |
Workflows
Any repository, one command
You have a GitHub URL (or a local checkout) and want a verified pull request.
hotpath doctor # git, Docker, keys, accelerator
hotpath check https://github.com/you/your-repo # seconds, read-only: blockers, cautions, cost
hotpath go https://github.com/you/your-repo # the nine stages, ending in a draft PR- Nine stages: setup, fetch, assess, baseline, benchmark, configure, optimize, publish, verify. Each one finishes or stops with the reason and a next step.
- The baseline must be green on 3 of 3 runs; a flaky or failing suite stops the run before anything is measured.
- No benchmark in the repo? Hotpath profiles the test suite, writes one over the hottest functions, and validates it by running it before trusting it.
- Nothing is pushed without your confirmation (or --yes), never to the default branch, and the PR opens as a draft.
- After the PR opens, Hotpath waits for the repository's own CI and only touches failures the PR introduced.
Inside your own repository
You maintain the code and want Hotpath as a repeatable tool with a config you control.
cd my-repo
hotpath init # writes .hotpath.yaml + .github/workflows/hotpath-verify.yml
git add .hotpath.yaml .gitignore .github && git commit -m "Set up Hotpath" && git push
hotpath run --pr # optimize, then push hotpath/<run id> and open a PR
hotpath serve # the dashboard for this repo's runs
hotpath pr --prune # later: ship only the changes that pull their weight- init detects your correctness check and benchmark, locks them, and asks what Hotpath may edit.
- Every command then finds the nearest .hotpath.yaml, so no config path is needed.
- The generated workflow re-runs the locked check on GitHub and fails any hotpath/* PR that touches a locked or non-editable file.
- Publishing refuses if your base branch moved since the run measured it; re-publishing the same run refreshes the same PR.
Offline demo, no API keys
You want to see the whole loop, including every rejection path, before spending anything.
git clone https://github.com/Nijjea1/hotpath && cd hotpath
pip install -e .
hotpath run configs/demo_repo.yaml # full loop with recorded patches
hotpath serve configs/demo_repo.yaml # http://127.0.0.1:8765
hotpath ablate configs/demo_repo.yaml # what each kept change contributed
hotpath export configs/demo_repo.yaml out/ # PR bundle: tree, changes.patch, REPORT.md
hotpath run configs/demo_repo_beam.yaml # beam search + a retry chain- The mock provider replaces the model, not the verification: every recorded patch still goes through a real worktree, the locked tests, and the benchmark.
- Several recorded patches are deliberately wrong (a test edit, a behaviour change, a no-op), so you see locked_file, rejected_correctness, rejected_speed and patch_failed for real.
Bring your own benchmark
You already know what 'faster' means for your code and want Hotpath to measure exactly that.
# bench.py — lock this file in .hotpath.yaml
from hotpath.benchlib import run
from mylib import workload
run(lambda: workload(), warmup=2, trials=15, seed=0)
# prints one JSON line: {"hotpath_benchmark": 1, "samples": [...], "metric": "seconds"}- Any command works as bench_cmd if it prints one JSON line with at least two positive samples.
- Add "higher_is_better": true for throughput metrics such as tokens/sec.
- GPU targets can declare a fixed workload matrix (benchmark.required_workloads); a candidate must keep ≥98% of its parent's throughput on every shape, not just win on aggregate.
The full stage-by-stage behaviour of go is in docs/GO.md; publishing details in docs/PULL_REQUESTS.md.
How it decides
Two steps are models, allowed to be creative and wrong. The rest is harness code the models cannot edit. A patch that fails a gate is rejected and recorded with the reason.
Where does the time actually go?
Runs the target under a profiler and summarises the widest bars. Amdahl's law: a function that takes 5% of runtime can only ever save 5%.
What should we try next?
The planner reads the profile, the hottest source, and every past result, and returns structured experiments. It stops proposing what already failed.
Can the idea be written as a patch?
Fast worker models write each idea as a patch in its own isolated worktree, so many experiments run in parallel without touching each other.
Is the output still the same?
Generated tokens must match the frozen reference exactly and logits must stay within tolerance. Tests are locked: a patch that touches them is rejected before a byte is written.
Is it faster than the noise?
Warmup runs, GPU sync, many trials, the median. The speedup must clear the machine-measured noise floor, and its 95% interval must exclude 1.0.
Does it still help on top of the rest?
Accepted patches stack, and each new one is measured against the stack it builds on. The bottleneck moves, so Hotpath re-profiles. An ablation pass re-measures what each shipped change contributes.
The accept rule
- Correct. The locked test_cmd exits 0. A failing candidate is never benchmarked.
- Noise floor. The untouched code is benchmarked baseline_repeats times; the variation between runs is the noise.
- Threshold. max(min_speedup, 1 + noise_multiplier × noise) — 3% on a quiet machine, higher on a noisy one.
- Confidence. 2,000 bootstrap resamples; the 95% interval of the median ratio must exclude 1.0.
- Quiet machine. Benchmarks take an exclusive lock, so nothing else runs while one is measured.
| Verdict | Meaning |
|---|---|
| accepted | correct, and faster than the threshold with a confidence interval clear of 1.0 |
| rejected_correctness | the locked tests failed; the benchmark never ran |
| rejected_speed | correct, but not measurably faster (or slower) |
| patch_failed | the edit could not be applied (search text not found, ambiguous, no-op) |
| locked_file | the patch touched a locked path; blocked before anything ran |
| not_selected | correct and faster, but a sibling was faster still |
| timeout | a test, benchmark or model call hit its hard limit |
| error | the benchmark crashed or printed no usable samples, or the harness hit an unexpected error |
Implementation: benchmark.py. Recorded runs, including refusals: the evidence ledger.
CLI reference
Every command and option below is generated from Hotpath's own argument parser — the same text as hotpath <command> --help. Global flag: -v for debug logging and full tracebacks.
Use it when: First thing on a new machine, and whenever a run fails for an environmental reason.
hotpath doctor
hotpath doctor --verify-keys # ask each endpoint whether the key still works
hotpath doctor --require-gpu # exit non-zero unless PyTorch sees an acceleratorShow all 3 options ▸Hide options ▾
- --json
- print the report as JSON
- --require-gpu
- fail unless PyTorch can use CUDA, ROCm, Intel XPU, or Apple MPS
- --verify-keys
- ask each model endpoint whether its key still works (a present key can be revoked)
Use it when: Setting up a repository you own so that run, pr and serve need no config path.
hotpath init
hotpath init --yes --test-cmd "python -m pytest -q" --bench-cmd "python bench.py" --editable "src/*.py"Show all 14 options ▸Hide options ▾
- [path]
- repository to set up (default: the current directory)
- --test-cmd TEST_CMD
- command that checks correctness (exit 0 = correct)
- --bench-cmd BENCH_CMD
- command that prints Hotpath benchmark JSON
- --profile-cmd PROFILE_CMD
- command that prints a Hotpath profile (optional)
- --editable EDITABLE
- comma-separated globs Hotpath may edit (default: *.py, or src/*.py)
- --lock LOCK
- an extra glob Hotpath must never edit (repeatable)
- --execution {local,docker}
- where candidate code runs
- --planner {openai,mock}
- model that proposes hypotheses
- --worker {openai,baseten,mock}
- model that writes each patch
- --mock-patches MOCK_PATCHES
- directory of recorded patches for offline `--provider mock` runs
- --name NAME
- name shown in reports (default: the directory name)
- -y, --yes
- accept detected defaults without prompting
- --force
- overwrite an existing .hotpath.yaml and workflow
- --no-workflow
- do not write the GitHub Actions check
Use it when: One command from a URL to a verified draft pull request.
hotpath go https://github.com/you/your-repo
hotpath go owner/repo --iterations 4 --beam 2 --max-minutes 30
hotpath go . --provider mock --mock-patches DIR --sandbox local --no-pr # offlineShow all 29 options ▸Hide options ▾
- [target]
- GitHub URL, owner/repo, any git URL, or a local checkout (default: the current directory)
- -y, --yes
- accept every default: local execution consent, the benchmark, and pushing the PR branch
- --provider {openai,mock}
- model provider (default: openai; mock is offline)
- --mock-patches MOCK_PATCHES
- recorded patches for --provider mock
- --iterations ITERATIONS
- search iterations (default: 3)
- --candidates CANDIDATES
- candidates per iteration
- --beam BEAM
- beam width: accepted heads kept per iteration (1 = greedy)
- --max-tokens MAX_TOKENS
- stop the search once model calls have used this many tokens
- --max-minutes MAX_MINUTES
- stop the search after this many minutes
- --sandbox {auto,local,docker}
- where the target's code runs (auto: Docker if it is running, else local with consent)
- --test-cmd TEST_CMD
- override the detected correctness check
- --bench-cmd BENCH_CMD
- use this benchmark (prints Hotpath JSON) instead of finding or generating one
- --no-generate
- never ask a model to write a benchmark
- --reuse-benchmark, --no-reuse-benchmark
- measure against the same generated benchmark as the last run on this repository, re-validated before use (default: on; without it two runs are not comparable)
- --regenerate-benchmark
- write a new benchmark even if the remembered one still validates
- --test-runs TEST_RUNS
- baseline test runs used to detect flaky tests
- --test-timeout TEST_TIMEOUT
- seconds per test-suite run
- --no-pr
- build the PR branch locally but do not push
- --pr-method {auto,gh,token,link}
- how to open the PR: gh CLI, GITHUB_TOKEN, or a pre-filled link (auto tries them in order)
- --ready
- open the PR ready for review instead of as a draft
- --no-open
- do not open the PR or dashboard in a browser
- --dashboard, --no-dashboard
- serve the live dashboard during the search and keep it up afterwards (default: on)
- --verify-ci, --no-verify-ci
- after opening the PR, wait for the repository's CI and fix what the PR broke (default: on; failures already present on the base are never touched)
- --ci-attempts CI_ATTEMPTS
- how many times to try repairing CI
- --ci-timeout CI_TIMEOUT
- seconds to wait for CI checks to settle
- --port PORT
- dashboard port (default: 8765)
- --workspaces WORKSPACES
- where clones and per-repo environments live (default: ./workspaces)
- --remote REMOTE
- git remote to push the PR branch to (default: origin)
- --resume
- continue the last `go` run on this repository
Use it when: Before spending anything: what would stop a run, and what it would cost. Exit 0 go, 1 caution, 2 stop.
hotpath check https://github.com/you/your-repo
hotpath check . --jsonShow all 4 options ▸Hide options ▾
- [path]
- local path or GitHub URL (default: the current directory)
- --json
- print the findings as JSON
- --iterations ITERATIONS
- iterations to price the estimate for
- --candidates CANDIDATES
- candidates per iteration to price for
Use it when: A read-only look at what Hotpath detects: ecosystem, tier, test and benchmark commands, editable and locked files.
hotpath assess path/to/repo
hotpath assess . --jsonShow all 2 options ▸Hide options ▾
- [path]
- repository to inspect (default: the current directory)
- --json
- print the report as JSON
Use it when: One optimization loop from a config you control.
hotpath run # nearest .hotpath.yaml
hotpath run configs/demo_repo.yaml --iterations 5 --beam 2
hotpath run --pr # publish if something was proved
hotpath run --resume run_ab12cd34 # continue a stopped runShow all 13 options ▸Hide options ▾
- [config]
- config file (default: the nearest .hotpath.yaml)
- --iterations ITERATIONS
- override search.iterations from the config
- --beam BEAM
- beam width: how many accepted heads to keep and expand each iteration (1 = greedy)
- --provider {mock,openai}
- override both planner and worker providers
- --export EXPORT
- directory to write the best accepted source tree into
- --autocommit
- snapshot uncommitted changes in the target into a commit before measuring (default: refuse)
- --resume RUN_ID
- continue a stopped or failed run from its stored beam instead of re-measuring a baseline; runs up to the larger of its original budget and --iterations
- --pr
- when the run finishes with a win, push a branch and open a pull request
- --base BASE
- branch the PR targets (default: the branch the run measured)
- --remote REMOTE
- git remote to push the PR branch to (default: origin)
- --draft
- open the PR as a draft
- --pr-method {auto,gh,token,link}
- how to open the PR: gh CLI, GITHUB_TOKEN, or a pre-filled link (auto tries them in order)
- --allow-moved-base
- publish even though the base branch moved since the run measured it
Use it when: Publishing a finished run as a pull request, one verified commit per change.
hotpath pr
hotpath pr --run-id run_ab12cd34 --prune --draft
hotpath pr --no-push # build the branch locally onlyShow all 9 options ▸Hide options ▾
- [config]
- config file (default: the nearest .hotpath.yaml)
- --run-id RUN_ID
- which run to publish (defaults to the latest)
- --prune
- ablate first and ship the verified pruned stack if there is one
- --no-push
- only build the local branch hotpath/<run id>
- --base BASE
- branch the PR targets (default: the branch the run measured)
- --remote REMOTE
- git remote to push the PR branch to (default: origin)
- --draft
- open the PR as a draft
- --pr-method {auto,gh,token,link}
- how to open the PR: gh CLI, GITHUB_TOKEN, or a pre-filled link (auto tries them in order)
- --allow-moved-base
- publish even though the base branch moved since the run measured it
Use it when: Reading a run: the experiment tree, the metric chart with its noise band, and every rejection reason.
hotpath serve # http://127.0.0.1:8765
hotpath serve --db path/to/hotpath.db # read-onlyShow all 4 options ▸Hide options ▾
- [config]
- config file; without one (and with --db) the dashboard is read-only
- --db DB
- database path (defaults to the config's)
- --host HOST
- loopback address to bind (non-loopback is refused)
- --port PORT
- port (default: 8765)
Use it when: Finding out what each accepted change actually contributed.
hotpath ablate
hotpath ablate --prune --json ablation.jsonShow all 4 options ▸Hide options ▾
- [config]
- config file (default: the nearest .hotpath.yaml)
- --run-id RUN_ID
- which run to ablate (defaults to the latest)
- --json JSON
- write the report here
- --prune
- also try dropping every change that did not pull its weight, together; keep the pruned stack only if it passes correctness and the full stack is not measurably faster
Use it when: A reviewable bundle without touching git remotes: optimized tree, changes.patch, REPORT.md.
hotpath export .hotpath.yaml out/
hotpath export .hotpath.yaml out/ --ablate --pruneShow all 5 options ▸Hide options ▾
- config
- config file (use .hotpath.yaml for the current repository)
- dest
- directory to write the bundle into
- --run-id RUN_ID
- which run to export (defaults to the latest)
- --ablate
- also run leave-one-out ablation and include the table (re-runs benchmarks)
- --prune
- ablate, then export the pruned stack instead of the head if pruning is verified (implies --ablate)
Configuration
hotpath init writes .hotpath.yaml for you, and hotpath go writes it into its setup commit. Every command then finds the nearest one. A config looks like this:
# configs/dryft_local.yaml (excerpt) — the config behind the run above
target: ../targets/torch_transformer
test_cmd: python tests/check.py # greedy tokens must match the frozen reference
bench_cmd: python bench.py # tokens_per_s across four prompt/batch shapes
profile_cmd: python hotprofile.py # torch.profiler summary
editable: ["model.py", "kernels/*.py", "*.py"]
locked: ["tests/*", "bench.py", "hotprofile.py", "reference*.py"]
benchmark: {min_speedup: 1.03, noise_multiplier: 2.0, baseline_repeats: 5, rebenchmark_parent: true}
search: {iterations: 6, candidates_per_iteration: 4, beam_width: 2, max_patch_retries: 1}Target
| Key | Default | Meaning |
|---|---|---|
| name | required | Name shown in reports and the dashboard. |
| target | required | Path to the repository, relative to the config file. |
| test_cmd | required | Correctness check. Exit 0 means correct; it is the only definition Hotpath uses. |
| bench_cmd | required | Prints one JSON line with samples (see Bring your own benchmark). |
| profile_cmd | none | Prints hotspots for the planner (cProfile or torch.profiler via hotpath.profilelib). |
| editable | required | Globs the agent may change. |
| locked | [] | Globs it must never touch, enforced in code before a byte is written. |
| correctness_contract | preserve outputs, ordering, exceptions… | Plain-language contract shown to the models. |
| workdir | .hotpath | Where worktrees and the SQLite store live, relative to the target. |
benchmark
| Key | Default | Meaning |
|---|---|---|
| min_speedup | 1.03 | Absolute floor: below this a change is never accepted. |
| noise_multiplier | 2.0 | Threshold = max(min_speedup, 1 + multiplier × baseline noise). |
| baseline_repeats | 3 | Baseline re-runs used to measure noise. |
| bootstrap_samples | 2000 | Resamples for the confidence interval. |
| confidence | 0.95 | The interval of the median ratio must exclude 1.0 at this level. |
| exclusive | true | Benchmarks run alone. Set false only when tests and benchmarks use different resources. |
| rebenchmark_parent | false | Re-measure the parent before each candidate to cancel machine drift (doubles benchmark cost). |
| required_workloads | [] | Workload IDs every benchmark must report. |
| min_workload_retention | 0.98 | Per-workload floor versus the parent. |
search
| Key | Default | Meaning |
|---|---|---|
| iterations | 3 | Plan → generate → verify rounds. |
| candidates_per_iteration | 3 | Hypotheses per round. |
| beam_width | 1 | Accepted heads kept and expanded each round; 1 is greedy. |
| max_patch_retries | 1 | Re-ask the worker once with the failure fed back; 0 disables. |
| max_parallel_workers | 4 | Concurrent worker model calls. |
| max_parallel_tests | 2 | Concurrent correctness runs. |
| planner_retries | 2 | Extra planner attempts after a transient API failure. |
provider
| Key | Default | Meaning |
|---|---|---|
| planner / worker | mock | openai (any OpenAI-compatible endpoint) or mock (recorded patches). |
| planner_model | gpt-4.1 | One call per iteration; keep it strong. |
| worker_model | gpt-4.1-mini | Many parallel calls; fast and cheap is the point. |
| worker_base_url | none | OpenAI-compatible endpoint for workers, e.g. https://inference.baseten.co/v1. |
| planner_api_key_env / worker_api_key_env | OPENAI_API_KEY | Which environment variable holds each key. |
| mock_patches_dir | none | Recorded patches for offline runs. |
execution
| Key | Default | Meaning |
|---|---|---|
| backend | docker | docker (untrusted code, fails closed) or local (trusted code only). |
| image | hotpath-runner:local | Reviewed runner image; pin a digest for real use. |
| cpus / memory_mb / pids_limit | 2 / 4096 / 128 | Container quotas. |
| gpu | none | e.g. device=0 for NVIDIA. |
| devices / group_add | [] | Device passthrough, e.g. /dev/kfd and /dev/dri for ROCm. |
timeouts, profile, context
| Key | Default | Meaning |
|---|---|---|
| timeouts.test / bench / profile / model | 300 / 900 / 300 / 180 s | Hard limits; a timeout becomes a structured verdict, never a crash. |
| profile.retain | 40 | Hotspot rows stored for the before/after diff. |
| context.max_hotspots | 12 | Hotspot rows shown to the planner. |
| context.max_source_chars | 14000 | Source budget per worker (go sizes it to the repository). |
Models & keys
The pattern is “big model plans, fast model explores”: one planner call per iteration, many worker calls in parallel. Both use the OpenAI API by default; workers can point at any OpenAI-compatible endpoint.
# once per shell
export OPENAI_API_KEY=sk-... # PowerShell: $env:OPENAI_API_KEY="sk-..."
# or once per machine: add OPENAI_API_KEY=... to ~/.hotpath/.env
# workers on Baseten (planner stays on OpenAI)
export BASETEN_API_KEY=...
hotpath init --worker baseten # writes worker_base_url + worker_api_key_env for you| Variable | What it does |
|---|---|
| OPENAI_API_KEY | Planner key (and worker key unless workers use Baseten). |
| BASETEN_API_KEY | When set, hotpath go and init put workers on Baseten (moonshotai/Kimi-K2.7-Code); the planner stays on OpenAI. |
| GITHUB_TOKEN / GH_TOKEN | Opens PRs when the gh CLI is not logged in. Without either, Hotpath prints a pre-filled PR link. |
| HOTPATH_HOME | Where user-level keys live (default ~/.hotpath; keys go in its .env). |
| HOTPATH_TORCH_DEVICE | Force cuda, xpu, mps or cpu for the bundled PyTorch helpers. |
| HOTPATH_AUTOCOMMIT | Same as run --autocommit: snapshot uncommitted target changes instead of refusing. |
| SENTRY_DSN | Optional tracing, logs and metrics. Empty means telemetry is off. |
| HOTPATH_DISABLE_SENTRY | Force telemetry off even if a DSN is present. |
Keys are read from the environment, then ./.env, then ~/.hotpath/.env. They never enter containers, commits, or model prompts. hotpath doctor --verify-keys checks each one against its endpoint.
Isolation & security
| Where code runs | Use it for | What it protects |
|---|---|---|
| execution.backend: docker | any repository you did not write (the default) | a fresh container per command: no network, read-only source, non-root, quotas, no host env, home, or Docker socket; fails closed |
| execution.backend: local | trusted code only (the bundled demos) | nothing — it runs as you |
| go --sandbox auto | the default for go | Docker when it is running; otherwise asks before running on the host in a per-target venv |
- Locked paths are enforced in code before a byte is written; the models are told about them, but the code is the guarantee.
- Correctness is relative to your locked tests: a passing suite is evidence against that contract, not a proof of equivalence.
- The dashboard only answers the local machine; put an authenticated reverse proxy in front for remote access.
Threat model: docs/ISOLATION.md. Reporting a vulnerability: SECURITY.md.
Platforms & GPUs
| Host / accelerator | Local execution | Docker execution |
|---|---|---|
| Windows / Linux / macOS CPU | Supported | Supported (Linux containers) |
| NVIDIA (Windows, Linux) | PyTorch CUDA | execution.gpu: device=0 with the NVIDIA container toolkit |
| AMD (Linux) | PyTorch ROCm | devices: [/dev/kfd, /dev/dri], group_add: [video, render] |
| Intel GPU (Linux) | PyTorch XPU | devices: [/dev/dri] |
| Apple GPU (macOS) | PyTorch MPS | Not available (Docker cannot expose Metal) |
The bundled PyTorch helpers pick CUDA/ROCm, Intel XPU, Apple MPS or CPU automatically; HOTPATH_TORCH_DEVICE forces one. Kernel edits are allowed when the kernel and its call site are editable, and must keep a correct PyTorch fallback. hotpath doctor --require-gpu fails unless an accelerator is usable. More in docs/PLATFORMS.md.
The dashboard
hotpath serve # http://127.0.0.1:8765 (go opens it for you)- Experiment tree — every candidate, retries and beam branches included; click a node for its hypothesis, diff, test output and exact verdict.
- Progress chart — the raw metric or speedup, with the keep-threshold drawn as a band around the baseline.
- Hotspot comparison — baseline versus head profile, with bounds instead of a fabricated “−100%” when a function drops out of the top N.
- Funnel — proposed → applied → correct → accepted → shipped.
- Create PR — appears only when a finished run has something verified to ship.
Every number on it is read from the run's SQLite store, written by the harness.
Troubleshooting
Every stop names its stage and prints a next: line. The common ones:
Set the key in your shell, or add it to ~/.hotpath/.env. To try Hotpath with no key at all: --provider mock.
The key is revoked or mistyped. hotpath doctor --verify-keys checks every configured key.
Start Docker Desktop (or the daemon) until docker info succeeds. For a repository you trust, --sandbox local runs it on the host.
Docker was not available and nobody could answer the prompt. Start Docker, or pass --sandbox local / --yes for trusted code.
Hotpath refuses to measure against a suite that is not green 3/3. Narrow it with --test-cmd (e.g. deselect network or timing-dependent tests).
Property tests with a deadline fail under load. Deselect them with --test-cmd, or add a conftest profile with deadline=None.
The worker could not write exact search/replace text. Try a stronger worker_model; hotpath go already sizes the source budget to the repository.
That is a result, not an error. A noisy benchmark raises the bar; a narrower benchmark over the hot function (or rebenchmark_parent: true) measures smaller wins.
Add a remote, or use --no-pr / hotpath export for a local bundle.
Rerun with hotpath -v go … for the traceback, add --resume to continue, and please open an issue.
Contributors: run python -m pytest from the interpreter Hotpath is installed into (activate the venv first).
Still stuck? Open an issue with hotpath doctor --json and the command output.
FAQ
Can a model's claim ever get a change accepted?
No. Correctness is the exit code of your locked test command; speed is a bootstrap confidence interval over repeated benchmark trials against a measured noise floor. The models only propose.
What code is sent to the model providers?
The profile summary, run history, and the source of editable files. Locked files and symlinks are never sent. Keys never enter containers, commits or prompts.
What does a run cost?
hotpath check prices it before you start (tokens and an approximate dollar figure). --max-tokens and --max-minutes stop the search cleanly and keep what has been proved.
What if nothing gets faster?
Then nothing is shipped and no PR is opened. The dashboard keeps every candidate with the reason it was rejected, because that is the honest result.
Will it push to my main branch?
Never. It pushes a hotpath/* branch after you confirm and opens a draft PR. Your checkout is never modified; commits are rebuilt from the exact trees the harness tested.
Does it work on GPU code?
Yes, through the bundled PyTorch helpers (CUDA events, torch.profiler) on CUDA, ROCm, XPU and MPS. A GPU result is evidence only for the device, driver and workload it was measured on.