Documentation

Hotpath makes code faster and proves every change is correct. Install it, point it at a repository, and it comes back with a draft pull request of verified changes — or with the reason nothing was proved.

Get started

Install

Hotpath is a Python CLI. You need Python 3.11+ and Git; Docker and the GitHub gh CLI are optional.

pipx install hotpath-agent
hotpath doctor          # Git, Docker, API keys, and the active accelerator

The package is hotpath-agent; the command it installs is hotpath. The hotpath and hotpath-ai names on PyPI belong to unrelated projects. uv tool install hotpath-agent works too. For the latest unreleased code, pipx install git+https://github.com/Nijjea1/hotpath, or clone the repo and run hotpath.cmd (Windows) or ./hotpath.sh, which create a venv for you.

You needForWithout it
Python 3.11+, GiteverythingHotpath will not start
OPENAI_API_KEYreal planner and worker modelsuse --provider mock (recorded patches, fully offline)
Dockerrunning untrusted repositories in a locked-down containergo asks before running code on your machine
gh or GITHUB_TOKENopening pull requests and reading CIHotpath pushes the branch and prints a pre-filled PR link
Recipes

Workflows

Any repository, one command

You have a GitHub URL (or a local checkout) and want a verified pull request.

hotpath doctor                                   # git, Docker, keys, accelerator
hotpath check https://github.com/you/your-repo   # seconds, read-only: blockers, cautions, cost
hotpath go    https://github.com/you/your-repo   # the nine stages, ending in a draft PR
  • Nine stages: setup, fetch, assess, baseline, benchmark, configure, optimize, publish, verify. Each one finishes or stops with the reason and a next step.
  • The baseline must be green on 3 of 3 runs; a flaky or failing suite stops the run before anything is measured.
  • No benchmark in the repo? Hotpath profiles the test suite, writes one over the hottest functions, and validates it by running it before trusting it.
  • Nothing is pushed without your confirmation (or --yes), never to the default branch, and the PR opens as a draft.
  • After the PR opens, Hotpath waits for the repository's own CI and only touches failures the PR introduced.

Inside your own repository

You maintain the code and want Hotpath as a repeatable tool with a config you control.

cd my-repo
hotpath init                 # writes .hotpath.yaml + .github/workflows/hotpath-verify.yml
git add .hotpath.yaml .gitignore .github && git commit -m "Set up Hotpath" && git push
hotpath run --pr             # optimize, then push hotpath/<run id> and open a PR
hotpath serve                # the dashboard for this repo's runs
hotpath pr --prune           # later: ship only the changes that pull their weight
  • init detects your correctness check and benchmark, locks them, and asks what Hotpath may edit.
  • Every command then finds the nearest .hotpath.yaml, so no config path is needed.
  • The generated workflow re-runs the locked check on GitHub and fails any hotpath/* PR that touches a locked or non-editable file.
  • Publishing refuses if your base branch moved since the run measured it; re-publishing the same run refreshes the same PR.

Offline demo, no API keys

You want to see the whole loop, including every rejection path, before spending anything.

git clone https://github.com/Nijjea1/hotpath && cd hotpath
pip install -e .
hotpath run configs/demo_repo.yaml            # full loop with recorded patches
hotpath serve configs/demo_repo.yaml          # http://127.0.0.1:8765
hotpath ablate configs/demo_repo.yaml         # what each kept change contributed
hotpath export configs/demo_repo.yaml out/    # PR bundle: tree, changes.patch, REPORT.md
hotpath run configs/demo_repo_beam.yaml       # beam search + a retry chain
  • The mock provider replaces the model, not the verification: every recorded patch still goes through a real worktree, the locked tests, and the benchmark.
  • Several recorded patches are deliberately wrong (a test edit, a behaviour change, a no-op), so you see locked_file, rejected_correctness, rejected_speed and patch_failed for real.

Bring your own benchmark

You already know what 'faster' means for your code and want Hotpath to measure exactly that.

# bench.py — lock this file in .hotpath.yaml
from hotpath.benchlib import run
from mylib import workload

run(lambda: workload(), warmup=2, trials=15, seed=0)
# prints one JSON line: {"hotpath_benchmark": 1, "samples": [...], "metric": "seconds"}
  • Any command works as bench_cmd if it prints one JSON line with at least two positive samples.
  • Add "higher_is_better": true for throughput metrics such as tokens/sec.
  • GPU targets can declare a fixed workload matrix (benchmark.required_workloads); a candidate must keep ≥98% of its parent's throughput on every shape, not just win on aggregate.

The full stage-by-stage behaviour of go is in docs/GO.md; publishing details in docs/PULL_REQUESTS.md.

The mechanism

How it decides

Two steps are models, allowed to be creative and wrong. The rest is harness code the models cannot edit. A patch that fails a gate is rejected and recorded with the reason.

1Profileharness

Where does the time actually go?

Runs the target under a profiler and summarises the widest bars. Amdahl's law: a function that takes 5% of runtime can only ever save 5%.

2Hypothesizeai

What should we try next?

The planner reads the profile, the hottest source, and every past result, and returns structured experiments. It stops proposing what already failed.

3Implementai

Can the idea be written as a patch?

Fast worker models write each idea as a patch in its own isolated worktree, so many experiments run in parallel without touching each other.

4Verify correctnessharness

Is the output still the same?

Generated tokens must match the frozen reference exactly and logits must stay within tolerance. Tests are locked: a patch that touches them is rejected before a byte is written.

5Benchmarkharness

Is it faster than the noise?

Warmup runs, GPU sync, many trials, the median. The speedup must clear the machine-measured noise floor, and its 95% interval must exclude 1.0.

6Keep & composeharness

Does it still help on top of the rest?

Accepted patches stack, and each new one is measured against the stack it builds on. The bottleneck moves, so Hotpath re-profiles. An ablation pass re-measures what each shipped change contributes.

The accept rule

  1. Correct. The locked test_cmd exits 0. A failing candidate is never benchmarked.
  2. Noise floor. The untouched code is benchmarked baseline_repeats times; the variation between runs is the noise.
  3. Threshold. max(min_speedup, 1 + noise_multiplier × noise) — 3% on a quiet machine, higher on a noisy one.
  4. Confidence. 2,000 bootstrap resamples; the 95% interval of the median ratio must exclude 1.0.
  5. Quiet machine. Benchmarks take an exclusive lock, so nothing else runs while one is measured.
VerdictMeaning
acceptedcorrect, and faster than the threshold with a confidence interval clear of 1.0
rejected_correctnessthe locked tests failed; the benchmark never ran
rejected_speedcorrect, but not measurably faster (or slower)
patch_failedthe edit could not be applied (search text not found, ambiguous, no-op)
locked_filethe patch touched a locked path; blocked before anything ran
not_selectedcorrect and faster, but a sibling was faster still
timeouta test, benchmark or model call hit its hard limit
errorthe benchmark crashed or printed no usable samples, or the harness hit an unexpected error

Implementation: benchmark.py. Recorded runs, including refusals: the evidence ledger.

Reference

CLI reference

Every command and option below is generated from Hotpath's own argument parser — the same text as hotpath <command> --help. Global flag: -v for debug logging and full tracebacks.

hotpath doctor [options]
check Git, Docker, credentials, and CUDA/ROCm/XPU/MPS availability

Use it when: First thing on a new machine, and whenever a run fails for an environmental reason.

hotpath doctor
hotpath doctor --verify-keys      # ask each endpoint whether the key still works
hotpath doctor --require-gpu      # exit non-zero unless PyTorch sees an accelerator
Show all 3 options ▸
--json
print the report as JSON
--require-gpu
fail unless PyTorch can use CUDA, ROCm, Intel XPU, or Apple MPS
--verify-keys
ask each model endpoint whether its key still works (a present key can be revoked)
hotpath init [path] [options]
set a repository up for Hotpath (.hotpath.yaml + CI check)

Use it when: Setting up a repository you own so that run, pr and serve need no config path.

hotpath init
hotpath init --yes --test-cmd "python -m pytest -q" --bench-cmd "python bench.py" --editable "src/*.py"
Show all 14 options ▸
[path]
repository to set up (default: the current directory)
--test-cmd TEST_CMD
command that checks correctness (exit 0 = correct)
--bench-cmd BENCH_CMD
command that prints Hotpath benchmark JSON
--profile-cmd PROFILE_CMD
command that prints a Hotpath profile (optional)
--editable EDITABLE
comma-separated globs Hotpath may edit (default: *.py, or src/*.py)
--lock LOCK
an extra glob Hotpath must never edit (repeatable)
--execution {local,docker}
where candidate code runs
--planner {openai,mock}
model that proposes hypotheses
--worker {openai,baseten,mock}
model that writes each patch
--mock-patches MOCK_PATCHES
directory of recorded patches for offline `--provider mock` runs
--name NAME
name shown in reports (default: the directory name)
-y, --yes
accept detected defaults without prompting
--force
overwrite an existing .hotpath.yaml and workflow
--no-workflow
do not write the GitHub Actions check
hotpath go [target] [options]
one command: assess a repo, build and test it, find a benchmark, optimize, and open a draft PR

Use it when: One command from a URL to a verified draft pull request.

hotpath go https://github.com/you/your-repo
hotpath go owner/repo --iterations 4 --beam 2 --max-minutes 30
hotpath go . --provider mock --mock-patches DIR --sandbox local --no-pr   # offline
Show all 29 options ▸
[target]
GitHub URL, owner/repo, any git URL, or a local checkout (default: the current directory)
-y, --yes
accept every default: local execution consent, the benchmark, and pushing the PR branch
--provider {openai,mock}
model provider (default: openai; mock is offline)
--mock-patches MOCK_PATCHES
recorded patches for --provider mock
--iterations ITERATIONS
search iterations (default: 3)
--candidates CANDIDATES
candidates per iteration
--beam BEAM
beam width: accepted heads kept per iteration (1 = greedy)
--max-tokens MAX_TOKENS
stop the search once model calls have used this many tokens
--max-minutes MAX_MINUTES
stop the search after this many minutes
--sandbox {auto,local,docker}
where the target's code runs (auto: Docker if it is running, else local with consent)
--test-cmd TEST_CMD
override the detected correctness check
--bench-cmd BENCH_CMD
use this benchmark (prints Hotpath JSON) instead of finding or generating one
--no-generate
never ask a model to write a benchmark
--reuse-benchmark, --no-reuse-benchmark
measure against the same generated benchmark as the last run on this repository, re-validated before use (default: on; without it two runs are not comparable)
--regenerate-benchmark
write a new benchmark even if the remembered one still validates
--test-runs TEST_RUNS
baseline test runs used to detect flaky tests
--test-timeout TEST_TIMEOUT
seconds per test-suite run
--no-pr
build the PR branch locally but do not push
--pr-method {auto,gh,token,link}
how to open the PR: gh CLI, GITHUB_TOKEN, or a pre-filled link (auto tries them in order)
--ready
open the PR ready for review instead of as a draft
--no-open
do not open the PR or dashboard in a browser
--dashboard, --no-dashboard
serve the live dashboard during the search and keep it up afterwards (default: on)
--verify-ci, --no-verify-ci
after opening the PR, wait for the repository's CI and fix what the PR broke (default: on; failures already present on the base are never touched)
--ci-attempts CI_ATTEMPTS
how many times to try repairing CI
--ci-timeout CI_TIMEOUT
seconds to wait for CI checks to settle
--port PORT
dashboard port (default: 8765)
--workspaces WORKSPACES
where clones and per-repo environments live (default: ./workspaces)
--remote REMOTE
git remote to push the PR branch to (default: origin)
--resume
continue the last `go` run on this repository
hotpath check [path] [options]
preflight a repository in seconds: what would stop a run, and what it would cost (exit 0 go, 1 caution, 2 stop)

Use it when: Before spending anything: what would stop a run, and what it would cost. Exit 0 go, 1 caution, 2 stop.

hotpath check https://github.com/you/your-repo
hotpath check . --json
Show all 4 options ▸
[path]
local path or GitHub URL (default: the current directory)
--json
print the findings as JSON
--iterations ITERATIONS
iterations to price the estimate for
--candidates CANDIDATES
candidates per iteration to price for
hotpath assess [path] [options]
read-only report: ecosystem, tests, benchmark, editable and locked files

Use it when: A read-only look at what Hotpath detects: ecosystem, tier, test and benchmark commands, editable and locked files.

hotpath assess path/to/repo
hotpath assess . --json
Show all 2 options ▸
[path]
repository to inspect (default: the current directory)
--json
print the report as JSON
hotpath run [config] [options]
run the optimization loop once

Use it when: One optimization loop from a config you control.

hotpath run                                   # nearest .hotpath.yaml
hotpath run configs/demo_repo.yaml --iterations 5 --beam 2
hotpath run --pr                              # publish if something was proved
hotpath run --resume run_ab12cd34             # continue a stopped run
Show all 13 options ▸
[config]
config file (default: the nearest .hotpath.yaml)
--iterations ITERATIONS
override search.iterations from the config
--beam BEAM
beam width: how many accepted heads to keep and expand each iteration (1 = greedy)
--provider {mock,openai}
override both planner and worker providers
--export EXPORT
directory to write the best accepted source tree into
--autocommit
snapshot uncommitted changes in the target into a commit before measuring (default: refuse)
--resume RUN_ID
continue a stopped or failed run from its stored beam instead of re-measuring a baseline; runs up to the larger of its original budget and --iterations
--pr
when the run finishes with a win, push a branch and open a pull request
--base BASE
branch the PR targets (default: the branch the run measured)
--remote REMOTE
git remote to push the PR branch to (default: origin)
--draft
open the PR as a draft
--pr-method {auto,gh,token,link}
how to open the PR: gh CLI, GITHUB_TOKEN, or a pre-filled link (auto tries them in order)
--allow-moved-base
publish even though the base branch moved since the run measured it
hotpath pr [config] [options]
publish a finished run as a pull request (one verified commit per change)

Use it when: Publishing a finished run as a pull request, one verified commit per change.

hotpath pr
hotpath pr --run-id run_ab12cd34 --prune --draft
hotpath pr --no-push                          # build the branch locally only
Show all 9 options ▸
[config]
config file (default: the nearest .hotpath.yaml)
--run-id RUN_ID
which run to publish (defaults to the latest)
--prune
ablate first and ship the verified pruned stack if there is one
--no-push
only build the local branch hotpath/<run id>
--base BASE
branch the PR targets (default: the branch the run measured)
--remote REMOTE
git remote to push the PR branch to (default: origin)
--draft
open the PR as a draft
--pr-method {auto,gh,token,link}
how to open the PR: gh CLI, GITHUB_TOKEN, or a pre-filled link (auto tries them in order)
--allow-moved-base
publish even though the base branch moved since the run measured it
hotpath serve [config] [options]
serve the dashboard (and allow starting runs from it)

Use it when: Reading a run: the experiment tree, the metric chart with its noise band, and every rejection reason.

hotpath serve                                  # http://127.0.0.1:8765
hotpath serve --db path/to/hotpath.db          # read-only
Show all 4 options ▸
[config]
config file; without one (and with --db) the dashboard is read-only
--db DB
database path (defaults to the config's)
--host HOST
loopback address to bind (non-loopback is refused)
--port PORT
port (default: 8765)
hotpath ablate [config] [options]
re-measure the accepted chain with each change removed

Use it when: Finding out what each accepted change actually contributed.

hotpath ablate
hotpath ablate --prune --json ablation.json
Show all 4 options ▸
[config]
config file (default: the nearest .hotpath.yaml)
--run-id RUN_ID
which run to ablate (defaults to the latest)
--json JSON
write the report here
--prune
also try dropping every change that did not pull its weight, together; keep the pruned stack only if it passes correctness and the full stack is not measurably faster
hotpath export config dest [options]
write a PR-ready bundle (optimized tree, diff, benchmark table, ablation)

Use it when: A reviewable bundle without touching git remotes: optimized tree, changes.patch, REPORT.md.

hotpath export .hotpath.yaml out/
hotpath export .hotpath.yaml out/ --ablate --prune
Show all 5 options ▸
config
config file (use .hotpath.yaml for the current repository)
dest
directory to write the bundle into
--run-id RUN_ID
which run to export (defaults to the latest)
--ablate
also run leave-one-out ablation and include the table (re-runs benchmarks)
--prune
ablate, then export the pruned stack instead of the head if pruning is verified (implies --ablate)
Your own code

Configuration

hotpath init writes .hotpath.yaml for you, and hotpath go writes it into its setup commit. Every command then finds the nearest one. A config looks like this:

# configs/dryft_local.yaml (excerpt) — the config behind the run above
target: ../targets/torch_transformer
test_cmd: python tests/check.py      # greedy tokens must match the frozen reference
bench_cmd: python bench.py           # tokens_per_s across four prompt/batch shapes
profile_cmd: python hotprofile.py    # torch.profiler summary
editable: ["model.py", "kernels/*.py", "*.py"]
locked: ["tests/*", "bench.py", "hotprofile.py", "reference*.py"]
benchmark: {min_speedup: 1.03, noise_multiplier: 2.0, baseline_repeats: 5, rebenchmark_parent: true}
search: {iterations: 6, candidates_per_iteration: 4, beam_width: 2, max_patch_retries: 1}

Target

KeyDefaultMeaning
namerequiredName shown in reports and the dashboard.
targetrequiredPath to the repository, relative to the config file.
test_cmdrequiredCorrectness check. Exit 0 means correct; it is the only definition Hotpath uses.
bench_cmdrequiredPrints one JSON line with samples (see Bring your own benchmark).
profile_cmdnonePrints hotspots for the planner (cProfile or torch.profiler via hotpath.profilelib).
editablerequiredGlobs the agent may change.
locked[]Globs it must never touch, enforced in code before a byte is written.
correctness_contractpreserve outputs, ordering, exceptions…Plain-language contract shown to the models.
workdir.hotpathWhere worktrees and the SQLite store live, relative to the target.

benchmark

KeyDefaultMeaning
min_speedup1.03Absolute floor: below this a change is never accepted.
noise_multiplier2.0Threshold = max(min_speedup, 1 + multiplier × baseline noise).
baseline_repeats3Baseline re-runs used to measure noise.
bootstrap_samples2000Resamples for the confidence interval.
confidence0.95The interval of the median ratio must exclude 1.0 at this level.
exclusivetrueBenchmarks run alone. Set false only when tests and benchmarks use different resources.
rebenchmark_parentfalseRe-measure the parent before each candidate to cancel machine drift (doubles benchmark cost).
required_workloads[]Workload IDs every benchmark must report.
min_workload_retention0.98Per-workload floor versus the parent.

search

KeyDefaultMeaning
iterations3Plan → generate → verify rounds.
candidates_per_iteration3Hypotheses per round.
beam_width1Accepted heads kept and expanded each round; 1 is greedy.
max_patch_retries1Re-ask the worker once with the failure fed back; 0 disables.
max_parallel_workers4Concurrent worker model calls.
max_parallel_tests2Concurrent correctness runs.
planner_retries2Extra planner attempts after a transient API failure.

provider

KeyDefaultMeaning
planner / workermockopenai (any OpenAI-compatible endpoint) or mock (recorded patches).
planner_modelgpt-4.1One call per iteration; keep it strong.
worker_modelgpt-4.1-miniMany parallel calls; fast and cheap is the point.
worker_base_urlnoneOpenAI-compatible endpoint for workers, e.g. https://inference.baseten.co/v1.
planner_api_key_env / worker_api_key_envOPENAI_API_KEYWhich environment variable holds each key.
mock_patches_dirnoneRecorded patches for offline runs.

execution

KeyDefaultMeaning
backenddockerdocker (untrusted code, fails closed) or local (trusted code only).
imagehotpath-runner:localReviewed runner image; pin a digest for real use.
cpus / memory_mb / pids_limit2 / 4096 / 128Container quotas.
gpunonee.g. device=0 for NVIDIA.
devices / group_add[]Device passthrough, e.g. /dev/kfd and /dev/dri for ROCm.

timeouts, profile, context

KeyDefaultMeaning
timeouts.test / bench / profile / model300 / 900 / 300 / 180 sHard limits; a timeout becomes a structured verdict, never a crash.
profile.retain40Hotspot rows stored for the before/after diff.
context.max_hotspots12Hotspot rows shown to the planner.
context.max_source_chars14000Source budget per worker (go sizes it to the repository).
Providers

Models & keys

The pattern is “big model plans, fast model explores”: one planner call per iteration, many worker calls in parallel. Both use the OpenAI API by default; workers can point at any OpenAI-compatible endpoint.

# once per shell
export OPENAI_API_KEY=sk-...              # PowerShell: $env:OPENAI_API_KEY="sk-..."
# or once per machine: add OPENAI_API_KEY=... to ~/.hotpath/.env

# workers on Baseten (planner stays on OpenAI)
export BASETEN_API_KEY=...
hotpath init --worker baseten             # writes worker_base_url + worker_api_key_env for you
VariableWhat it does
OPENAI_API_KEYPlanner key (and worker key unless workers use Baseten).
BASETEN_API_KEYWhen set, hotpath go and init put workers on Baseten (moonshotai/Kimi-K2.7-Code); the planner stays on OpenAI.
GITHUB_TOKEN / GH_TOKENOpens PRs when the gh CLI is not logged in. Without either, Hotpath prints a pre-filled PR link.
HOTPATH_HOMEWhere user-level keys live (default ~/.hotpath; keys go in its .env).
HOTPATH_TORCH_DEVICEForce cuda, xpu, mps or cpu for the bundled PyTorch helpers.
HOTPATH_AUTOCOMMITSame as run --autocommit: snapshot uncommitted target changes instead of refusing.
SENTRY_DSNOptional tracing, logs and metrics. Empty means telemetry is off.
HOTPATH_DISABLE_SENTRYForce telemetry off even if a DSN is present.

Keys are read from the environment, then ./.env, then ~/.hotpath/.env. They never enter containers, commits, or model prompts. hotpath doctor --verify-keys checks each one against its endpoint.

Trust

Isolation & security

Where code runsUse it forWhat it protects
execution.backend: dockerany repository you did not write (the default)a fresh container per command: no network, read-only source, non-root, quotas, no host env, home, or Docker socket; fails closed
execution.backend: localtrusted code only (the bundled demos)nothing — it runs as you
go --sandbox autothe default for goDocker when it is running; otherwise asks before running on the host in a per-target venv
  • Locked paths are enforced in code before a byte is written; the models are told about them, but the code is the guarantee.
  • Correctness is relative to your locked tests: a passing suite is evidence against that contract, not a proof of equivalence.
  • The dashboard only answers the local machine; put an authenticated reverse proxy in front for remote access.

Threat model: docs/ISOLATION.md. Reporting a vulnerability: SECURITY.md.

Hardware

Platforms & GPUs

Host / acceleratorLocal executionDocker execution
Windows / Linux / macOS CPUSupportedSupported (Linux containers)
NVIDIA (Windows, Linux)PyTorch CUDAexecution.gpu: device=0 with the NVIDIA container toolkit
AMD (Linux)PyTorch ROCmdevices: [/dev/kfd, /dev/dri], group_add: [video, render]
Intel GPU (Linux)PyTorch XPUdevices: [/dev/dri]
Apple GPU (macOS)PyTorch MPSNot available (Docker cannot expose Metal)

The bundled PyTorch helpers pick CUDA/ROCm, Intel XPU, Apple MPS or CPU automatically; HOTPATH_TORCH_DEVICE forces one. Kernel edits are allowed when the kernel and its call site are editable, and must keep a correct PyTorch fallback. hotpath doctor --require-gpu fails unless an accelerator is usable. More in docs/PLATFORMS.md.

Reading a run

The dashboard

hotpath serve                    # http://127.0.0.1:8765 (go opens it for you)
  • Experiment tree — every candidate, retries and beam branches included; click a node for its hypothesis, diff, test output and exact verdict.
  • Progress chart — the raw metric or speedup, with the keep-threshold drawn as a band around the baseline.
  • Hotspot comparison — baseline versus head profile, with bounds instead of a fabricated “−100%” when a function drops out of the top N.
  • Funnel — proposed → applied → correct → accepted → shipped.
  • Create PR — appears only when a finished run has something verified to ship.

Every number on it is read from the run's SQLite store, written by the harness.

When it stops

Troubleshooting

Every stop names its stage and prints a next: line. The common ones:

stopped at [1/9] Setup: OPENAI_API_KEY is not set

Set the key in your shell, or add it to ~/.hotpath/.env. To try Hotpath with no key at all: --provider mock.

OPENAI_API_KEY was rejected by the API

The key is revoked or mistyped. hotpath doctor --verify-keys checks every configured key.

stopped at [4/9] Baseline: … Docker is not reachable

Start Docker Desktop (or the daemon) until docker info succeeds. For a repository you trust, --sandbox local runs it on the host.

no consent to run the repository's code on this machine

Docker was not available and nobody could answer the prompt. Start Docker, or pass --sandbox local / --yes for trusted code.

The baseline tests fail or are flaky

Hotpath refuses to measure against a suite that is not green 3/3. Narrow it with --test-cmd (e.g. deselect network or timing-dependent tests).

Hypothesis deadline errors in the baseline

Property tests with a deadline fail under load. Deselect them with --test-cmd, or add a conftest profile with deadline=None.

Every candidate is patch_failed

The worker could not write exact search/replace text. Try a stronger worker_model; hotpath go already sizes the source budget to the repository.

Nothing accepted: all rejected_speed

That is a result, not an error. A noisy benchmark raises the bar; a narrower benchmark over the hot function (or rebenchmark_parent: true) measures smaller wins.

no pull request will be opened: no 'origin' remote

Add a remote, or use --no-pr / hotpath export for a local bundle.

Hotpath hit an unexpected error at [n/9] …

Rerun with hotpath -v go … for the traceback, add --resume to continue, and please open an issue.

~20 tests fail with No module named 'hotpath'

Contributors: run python -m pytest from the interpreter Hotpath is installed into (activate the venv first).

Still stuck? Open an issue with hotpath doctor --json and the command output.

Questions

FAQ

Can a model's claim ever get a change accepted?

No. Correctness is the exit code of your locked test command; speed is a bootstrap confidence interval over repeated benchmark trials against a measured noise floor. The models only propose.

What code is sent to the model providers?

The profile summary, run history, and the source of editable files. Locked files and symlinks are never sent. Keys never enter containers, commits or prompts.

What does a run cost?

hotpath check prices it before you start (tokens and an approximate dollar figure). --max-tokens and --max-minutes stop the search cleanly and keep what has been proved.

What if nothing gets faster?

Then nothing is shipped and no PR is opened. The dashboard keeps every candidate with the reason it was rejected, because that is the honest result.

Will it push to my main branch?

Never. It pushes a hotpath/* branch after you confirm and opens a draft PR. Your checkout is never modified; commits are rebuilt from the exact trees the harness tested.

Does it work on GPU code?

Yes, through the bundled PyTorch helpers (CUDA events, torch.profiler) on CUDA, ROCm, XPU and MPS. A GPU result is evidence only for the device, driver and workload it was measured on.