Operations
Performance Budgets
Every operation this repo owns falls into one of four performance tiers. Each tier has a hard budget enforced by benches/test_perf_budgets.sh. If a change pushes an operation past its budget's headroom, CI fails.
The tiers
| Tier | Budget | What belongs here |
|---|---|---|
| INSTANT | ≤ 500ms | Anything a human sees at a shell prompt |
| FAST | ≤ 2000ms | Heavier gates + sandboxed CLI reads |
| MEDIUM | ≤ 5000ms | Full diagnostic runs |
| ACCEPTED-SLOW | documented | Multi-second ops that are legitimately slow (see below) |
| OUT-OF-SCOPE | not gated | Multi-minute ops we don't gate per-run |
Reference baselines (2026-08-30, rousseau-cachyos-geekom-a9, Ryzen AI 9 HX 370)
All medians in milliseconds. Budget = 2× median (or the tier ceiling, whichever is larger).
INSTANT tier
| Operation | Baseline | Budget | Headroom |
|---|---|---|---|
dot version | 10 | 500 | 50× |
dot help | 13 | 500 | 38× |
dot help <cmd> | 15 | 500 | 33× |
dot search <keyword> | 15 | 500 | 33× |
| iCloud script single-run | 15 | 500 | 33× |
FAST tier
Currently empty.
dot status and dot diff were placed here at 2000ms against reference medians of 1510 ms and 1420 ms — 1.3× and 1.4× headroom, short of the ≥ 2× this document requires of every budget. They have since moved to MEDIUM; see that section for the measurements that prompted it.
The baselines of 28 ms and 32 ms recorded before those were not real. The perf sandbox never created chezmoi's source directory, so both commands aborted immediately with "no such file or directory" — and
_measurediscarded exit codes, so the gate timed the failure path and called it excellent. The sandbox now links the repo intoXDG_DATA_HOME.
The QA gates and test suites that used to sit in this tier moved to GATES below: they are not interactive operations, so a ceiling defined by human-perceived latency never described them.
MEDIUM tier
| Operation | Baseline | Budget | Headroom |
|---|---|---|---|
dot status (sandbox) | 2329 | 5000 | 2.1× |
dot diff (sandbox) | 2157 | 5000 | 2.3× |
dot doctor | 3892 | 5000 | 1.3× |
bench.sh --quick | 826 | 5000 | 6× |
The dot status and dot diff baselines are the worst of the two hosted runners, not the reference machine: macOS measured 2290 ms / 2157 ms and Ubuntu 2329 ms / 2136 ms, against 1139 ms / 1174 ms locally. Two unrelated platforms agreeing within 8% is a property of the commands rather than of one slow runner — both shell out to chezmoi, which walks the whole source tree, and that is disk-bound. A budget taken from the faster machine would have been a gate that only ever fired on other people's hardware.
dot doctor and bench.sh --quick are diagnostics that report findings through their exit status — dot doctor exits 1 whenever it finds issues, and bench.sh exits 1 when a shell breaches its own startup threshold. They are gated with _gate_diag, which permits exit 1 but still fails on 2+ (not-found, permission, signal, syntax error). That is far narrower than the blanket || true it replaced.
GATES tier (CI quality gates and test suites)
Budgets are 2× the median measured on the slowest supported platform, not the fastest. Medians below: rousseau-mbp-m1, macOS 26 (Darwin 25.6), 2026-08-30.
| Operation | macOS median | Linux median | Budget | Headroom |
|---|---|---|---|---|
check-version-consistency.sh | 68 | 14 | 500 | 7.4× |
docs-coverage.sh | 969 | 148 | 2000 | 2.1× |
| iCloud regression test (12 assertions) | 494 | 94 | 1000 | 2.0× |
| iCloud unit test (29 assertions) | 1743 | 498 | 3500 | 2.0× |
| iCloud manifest test (11 assertions) | 1575 | — † | 5000 | 3.2× |
traceability-coverage.sh | 2389 | 623 | 5000 | 2.1× |
test_dot_subcommand_smoke.sh | 3785 | 1308 | 7500 | 2.0× |
test_dot_help_registry_symmetry.sh | 4449 | 1381 | 9000 | 2.0× |
† The manifest test hashes an entire sandbox tree before and after every scenario, so it is disk-bound in a way the other gates are not. Its budget is set at 3.2× the macOS median rather than the 2.0× used above, deliberately: dot status and dot diff were first budgeted at 1.3× and 1.4×, and both breached on the first CI run that measured them. The Linux median is left unfilled until CI reports one — guessing it would defeat the point of a table of measurements.
These gates run 3–6× slower on macOS than on Linux — fork/exec is markedly more expensive there and every one of them is fork-heavy shell. CI covers both platforms, so the original Linux-only calibration could not hold, and five of these sat red as a result. Re-capture the Linux column when convenient; it is carried over from the 2026-08-30 Ryzen baseline and is not the binding constraint.
ACCEPTED-SLOW (documented, not gated per-run)
| Operation | Baseline | Reason |
|---|---|---|
test_dot_help_flag_universal.sh | ~11s | Invokes dot help --help on ~100 commands via subshell each. The coverage it provides justifies the cost; regression is caught by benches/test_help_gates_wall_clock.sh at the suite level. |
OUT-OF-SCOPE (not gated per-run)
| Operation | Why we don't gate |
|---|---|
chezmoi apply | Fresh macOS: minutes. Depends on iCloud sync + package installs. Gated at suite level only. |
install.sh full | Downloads + installs packages. Network-bound. |
dot upgrade | Runs mise upgrade, chezmoi apply, package manager upgrades. |
| Full test suite | 15+ minutes on CI. Gated by workflow timeout, not per-run assertion. |
Ratchet vs aspiration
The budgets above are regression gates, not aspirations. If a real optimisation lowers a baseline, edit the doc + the perf test to lower the budget too. If a change pushes something over the budget, the test fails and CI blocks the merge.
The aspirational shell-startup target (<30ms) is tracked separately in benches/bench.sh — that's a bench, not a budget.
Adding a new operation
When you add a new script that runs at a shell prompt:
- Time it 5 runs on a warm system:
for _ in {1..5}; do time bash your-script; done - Take the median.
- Place it in the tier where
budget ≥ 2 × median. If a median is 400ms it goes in FAST (500ms is uncomfortably tight); if it's 300ms, INSTANT is fine. - Add it to
benches/test_perf_budgets.shin the correct tier section. - Add its baseline to this doc.
Where the enforcement lives
- Per-op budget test:
benches/test_perf_budgets.sh - Suite-level wall-clock ratchet:
benches/test_help_gates_wall_clock.sh - CI wiring:
.github/workflows/ci.yml, jobquality-performance— runs on ubuntu-latest and macos-latest for every PR, with no|| trueand no budget scaling. On the gates the hosted macOS runner is comparable to or faster than the reference machine — docs-coverage 679 ms vs 969, traceability 1775 vs 2389, help-registry 3318 vs 4449,dot doctor1746 vs 4423. The exception is anything that drives chezmoi over the whole source tree:dot statusanddot diffmeasured ~2× the local median on both hosted platforms, which is why they sit in MEDIUM rather than FAST. Budget a new operation against the slowest platform it will run on, not against this machine.
How the time is measured
_measure uses the time keyword with TIMEFORMAT='%3R', not date +%s%N.
%N is a GNU extension. BSD date — macOS 14 and earlier — copies the literal N through, so every arithmetic conversion failed, _measure returned nothing, and an empty median compares as 0 against any budget. The gate reported every budget met on those machines, for any command, including one that could not parse. time is a shell builtin, is millisecond-accurate in bash 3.2, and needs no external clock at all.
Two consequences worth keeping:
_gate_max_rcrejects a non-numeric median outright rather than comparing it. Both fields go through-gt/-le, where bash reads a non-numeric operand as0— so a gate that cannot measure would otherwise report success.tests/regression/test_gate_integrity.shputs adatewithout%Nfirst onPATHand requires the same verdicts. A future timing rewrite that reintroduces the dependency fails there rather than going quiet.
Environment knobs
| Variable | Default | Effect |
|---|---|---|
PERF_BUDGET_PERCENT | 100 | Scales every budget. 0 makes them all impossible — that is how test_gate_integrity.sh proves the gate actually fires. |
PERF_GATE_FILTER | (unset) | Runs only gates whose label contains this substring. Skipped gates never call test_start, so TESTS_RUN == PASSED + FAILED still holds. |
DOT_CLI | bin/dot | Points the dot gates at another binary, so breakage detection can be exercised against a deliberately corrupted CLI. |
When a budget fires
The error looks like:
✗ instant_dot_help: median=612ms EXCEEDS budget=500ms
Steps:
- Bisect the change that pushed it over.
- Fix the regression, or
- If the increase is legitimate (real new work), move the operation to the next tier + update this doc +
test_perf_budgets.shin the same PR.
Never silently bump the budget. The tier a thing lives in is a promise to users.