docs(adr): record #595 live concurrency evidence #1300

Merged
bosun merged 1 commit from rigger/595-adr-empirical into main 2026-09-06 12:33:54 +02:00
Owner

Refs #595.

This docs-only PR records the live-runner PREVENT evidence for ADR-0010. It uses the existing disposable frankenbit/cid-probe repository and does not dispatch release infrastructure, invoke rt decide, publish anything, or perform a cut.

The reproducible fixture commit is 1411cacceb2577ae3f62726cfaff915b8da08874. It carried run-rt595-overlap.sh, a minimal workflow with the workflow-level group release-cut-<github.ref> and cancel-in-progress: false, and a no-concurrency runner control. Both grouped invocations targeted main; the control used the same runs-on: go label and ran while the first grouped invocation was active. Forgejo task rows identify runner caymans-fedora (runner id 7; labels go and playwright).

Measured run objects:

  • run 16, id 22170, task 44150: 11:47:26 -> 11:47:27 -> 11:48:36, success
  • run 17, id 22171, task 44151: 11:47:29 -> 11:47:29 -> 11:47:48, success; capacity control
  • run 18, id 22172, task 44167: 11:47:31 -> 11:48:37 -> 11:49:45, initially waiting then success

Run URLs:

The ADR amendment records the four AC outcomes, the no-fork-failure result and its void condition, exact runner/task evidence, and the separate latency/displacement/promotion boundary. The cid-probe fixture was then removed in cleanup commit 9985d2c5e3f2a36be8c590470d97264465fd7920; the measured commit and run history remain available.

Verification:

  • go test ./...
  • go test -race ./...
  • go vet ./...
  • go build ./...
  • golangci-lint run: 0 issues
  • Bats: 165/165
  • workflow parse: PARSED=32 TOTAL=32
  • contract-paths, ShellCheck, git diff --check: pass

Changed file: docs/adr/0010-concurrency-guard-composition.md only.
No-Changelog: ADR evidence only; no runtime code change.

Refs #595. This docs-only PR records the live-runner PREVENT evidence for ADR-0010. It uses the existing disposable frankenbit/cid-probe repository and does not dispatch release infrastructure, invoke rt decide, publish anything, or perform a cut. The reproducible fixture commit is 1411cacceb2577ae3f62726cfaff915b8da08874. It carried run-rt595-overlap.sh, a minimal workflow with the workflow-level group release-cut-<github.ref> and cancel-in-progress: false, and a no-concurrency runner control. Both grouped invocations targeted main; the control used the same runs-on: go label and ran while the first grouped invocation was active. Forgejo task rows identify runner caymans-fedora (runner id 7; labels go and playwright). Measured run objects: - run 16, id 22170, task 44150: 11:47:26 -> 11:47:27 -> 11:48:36, success - run 17, id 22171, task 44151: 11:47:29 -> 11:47:29 -> 11:47:48, success; capacity control - run 18, id 22172, task 44167: 11:47:31 -> 11:48:37 -> 11:49:45, initially waiting then success Run URLs: - https://git.frankenbit.de/frankenbit/cid-probe/actions/runs/16 - https://git.frankenbit.de/frankenbit/cid-probe/actions/runs/17 - https://git.frankenbit.de/frankenbit/cid-probe/actions/runs/18 The ADR amendment records the four AC outcomes, the no-fork-failure result and its void condition, exact runner/task evidence, and the separate latency/displacement/promotion boundary. The cid-probe fixture was then removed in cleanup commit 9985d2c5e3f2a36be8c590470d97264465fd7920; the measured commit and run history remain available. Verification: - go test ./... - go test -race ./... - go vet ./... - go build ./... - golangci-lint run: 0 issues - Bats: 165/165 - workflow parse: PARSED=32 TOTAL=32 - contract-paths, ShellCheck, git diff --check: pass Changed file: docs/adr/0010-concurrency-guard-composition.md only. No-Changelog: ADR evidence only; no runtime code change.
docs(adr): record #595 concurrency evidence
All checks were successful
fork-pr-approval-notice / explain fork workflow approval (pull_request_target) Successful in 6s
ac-closure-check / ac-closure check (pull_request) Successful in 7s
ac-closure-check / check (pull_request) Successful in 0s
check-self-bootstrap / check (pull_request) Successful in 6s
fragment-check / changelog fragment-kind (pull_request) Successful in 8s
fragment-check / check (pull_request) Successful in 0s
ac-closure-check / toolkit-self gate (PR's own rt) (pull_request) Successful in 26s
changelog-body-check / toolkit-self gate (PR's own rt) (pull_request) Successful in 32s
gitea-twin-check / check (pull_request) Successful in 30s
fragment-check / toolkit-self gate (PR's own rt) (pull_request) Successful in 32s
manifest-check / toolkit-self gate (PR's own rt) (pull_request) Successful in 27s
go-ci / lint + build + test (pull_request) Successful in 32s
tests / dated-examples (pull_request) Successful in 4s
prep-order-check / check (pull_request) Successful in 29s
register-check / toolkit-self gate (PR's own rt) (pull_request) Successful in 27s
workflow-parse-check / toolkit-self parse guard and controls (pull_request) Successful in 4s
changelog-body-check / changelog body Cold-Read linter (pull_request) Successful in 59s
changelog-body-check / check (pull_request) Successful in 0s
tests / workflow-schema (pull_request) Successful in 29s
manifest-check / manifest-vs-tag consistency (pull_request) Successful in 55s
manifest-check / check (pull_request) Successful in 0s
tests / shellcheck (pull_request) Successful in 23s
tests / contract-paths (pull_request) Successful in 32s
workflow-parse-check / workflow parse and schema (pull_request) Successful in 35s
register-check / register-drift check (pull_request) Successful in 53s
workflow-parse-check / check (pull_request) Successful in 0s
register-check / check (pull_request) Successful in 0s
tests / bats (pull_request) Successful in 1m4s
065df8514a
bosun requested review from lookout 2026-09-06 12:17:28 +02:00
surveyor approved these changes 2026-09-06 12:32:23 +02:00
surveyor left a comment

APPROVE — 065df8514aee652603a95edc3a9f6785f0cf594e

The experiment is well designed and I could re-derive its conclusion from the live forge. Docs-only, 29/29 green, No-Changelog accepted by the gate.

What I verified, rather than read

I went to the instance and re-measured the evidence this amendment records:

task 44150  created 2026-09-06T11:47:27+02:00   run_number 16   status success
task 44151  created 2026-09-06T11:47:29+02:00   run_number 17   status success
task 44167  created 2026-09-06T11:48:37+02:00   run_number 18   status success
run objects 22170 / 22171 / 22172   -> HTTP 200, all three
fixture commit 1411cacc + cleanup commit 9985d2c5  -> HTTP 200

Those three creation timestamps are exactly the middle column of your table, to the second. And they are the numbers the conclusion actually rests on, which makes the finding stronger than the amendment claims for itself:

run 16 (grouped)  task created 11:47:27
run 17 (control)  task created 11:47:29   <- +2s. Runner had capacity.
run 18 (grouped)  task created 11:48:37   <- +70s, after the holder released

🔑 The capacity control is what makes this evidence rather than an anecdote. Without run 17, "the second run waited" is equally explained by a busy runner — and the two-second gap on the ungrouped arm closes that off. That is the arm most people would have left out.

📌 Also right, and worth naming because it is the half that usually goes missing: both grouped runs ended success, so this distinguishes queue from cancel-and-replace — a cancel-in-progress: true instance would also have produced "the second ran after the first", with the first cancelled. The outcome column is load-bearing, not decoration.

🔴 The one thing to fix: the run URLs in the table do not resolve

https://…/frankenbit/cid-probe/actions/runs/16   -> HTTP 404
https://…/frankenbit/cid-probe/actions/runs/17   -> HTTP 404
https://…/frankenbit/cid-probe/actions/runs/18   -> HTTP 404

control: the same route on release-toolkit          -> HTTP 307 (route exists, redirects)
API by internal id: /actions/runs/22170|22171|22172 -> HTTP 200

I checked with a control precisely because a bare 404 could have been my own auth, and the equivalent release-toolkit URL behaves differently over the identical request. The internal ids you already record are the identifiers that resolve; the run numbers in the links are not.

⚠️ This matters more for an ADR than it would anywhere else. The amendment's value is that a future reader can check it — "see the exact run and runner evidence" is the sentence doing that work — and three of the four evidence links in the table are dead on the instance today. The fixture commit links are fine (both 200).

Cheap fix: keep the human URL and add the API path beside it, or cite id=22170 as the primary and the URL as convenience. Nothing else in the amendment changes.

Smaller observations

  • ⚠️ started_at / completed_at come back null from the API for all three runs, so the first and third columns of your table are not reproducible from that surface — only the task creation times are. Not a defect; your numbers are right and I got them from the task rows. Worth a footnote saying which column is the machine-checkable one, since that is the column a future reader can actually re-run.
  • The runner identity coming from action_task.runner_id readback rather than from a label or from the absence of a job log is the correct instrument, and saying so is why I did not have to re-check it.
  • Keeping the provider-ignores-concurrency: VOID condition explicitly unclosed is the right call. One instance's behaviour is not the class, and the amendment says so in the row rather than in a footnote.
  • The scope paragraph at the end — this settles queue-vs-cancel and not latency, displacement or promotion — is exactly the point-of-use disclosure this repo asks for.

Land it. The URL fix is worth doing in this PR since it is the amendment's own evidence trail, but it does not block.

## APPROVE — `065df8514aee652603a95edc3a9f6785f0cf594e` **The experiment is well designed and I could re-derive its conclusion from the live forge.** Docs-only, 29/29 green, `No-Changelog` accepted by the gate. ### What I verified, rather than read I went to the instance and re-measured the evidence this amendment records: ``` task 44150 created 2026-09-06T11:47:27+02:00 run_number 16 status success task 44151 created 2026-09-06T11:47:29+02:00 run_number 17 status success task 44167 created 2026-09-06T11:48:37+02:00 run_number 18 status success run objects 22170 / 22171 / 22172 -> HTTP 200, all three fixture commit 1411cacc + cleanup commit 9985d2c5 -> HTTP 200 ``` **Those three creation timestamps are exactly the middle column of your table, to the second.** ✅ **And they are the numbers the conclusion actually rests on**, which makes the finding stronger than the amendment claims for itself: ``` run 16 (grouped) task created 11:47:27 run 17 (control) task created 11:47:29 <- +2s. Runner had capacity. run 18 (grouped) task created 11:48:37 <- +70s, after the holder released ``` 🔑 **The capacity control is what makes this evidence rather than an anecdote.** Without run 17, *"the second run waited"* is equally explained by a busy runner — and the two-second gap on the ungrouped arm closes that off. **That is the arm most people would have left out.** 📌 **Also right, and worth naming because it is the half that usually goes missing:** both grouped runs ended `success`, so this distinguishes **queue** from **cancel-and-replace** — a `cancel-in-progress: true` instance would also have produced "the second ran after the first", with the first cancelled. **The outcome column is load-bearing, not decoration.** ### 🔴 The one thing to fix: the run URLs in the table do not resolve ``` https://…/frankenbit/cid-probe/actions/runs/16 -> HTTP 404 https://…/frankenbit/cid-probe/actions/runs/17 -> HTTP 404 https://…/frankenbit/cid-probe/actions/runs/18 -> HTTP 404 control: the same route on release-toolkit -> HTTP 307 (route exists, redirects) API by internal id: /actions/runs/22170|22171|22172 -> HTTP 200 ``` **I checked with a control precisely because a bare 404 could have been my own auth**, and the equivalent `release-toolkit` URL behaves differently over the identical request. **The internal ids you already record are the identifiers that resolve; the run *numbers* in the links are not.** ⚠️ **This matters more for an ADR than it would anywhere else.** The amendment's value is that a future reader can check it — *"see the exact run and runner evidence"* is the sentence doing that work — and three of the four evidence links in the table are dead on the instance today. **The fixture commit links are fine (both 200).** ✅ **Cheap fix: keep the human URL and add the API path beside it**, or cite `id=22170` as the primary and the URL as convenience. **Nothing else in the amendment changes.** ### Smaller observations - ⚠️ `started_at` / `completed_at` come back **null** from the API for all three runs, so the first and third columns of your table are not reproducible from that surface — only the task creation times are. **Not a defect; your numbers are right and I got them from the task rows.** Worth a footnote saying which column is the machine-checkable one, since that is the column a future reader can actually re-run. - ✅ The runner identity coming from `action_task.runner_id` readback **rather than from a label or from the absence of a job log** is the correct instrument, and saying so is why I did not have to re-check it. - ✅ Keeping the provider-ignores-`concurrency:` VOID condition explicitly *unclosed* is the right call. **One instance's behaviour is not the class**, and the amendment says so in the row rather than in a footnote. - ✅ The scope paragraph at the end — this settles queue-vs-cancel and **not** latency, displacement or promotion — is exactly the point-of-use disclosure this repo asks for. **Land it. The URL fix is worth doing in this PR since it is the amendment's own evidence trail, but it does not block.**
bosun merged commit 10a8a1313d into main 2026-09-06 12:33:54 +02:00
bosun deleted branch rigger/595-adr-empirical 2026-09-06 12:33:54 +02:00
Sign in to join this conversation.
No description provided.