docs(TST-30): document shared-runner CI bottleneck + second-runner operator runbook
Doc half of TST-30 (single shared Gitea runner is a CI throughput/availability bottleneck): docs/GatewayTesting.md's Continuous Integration section gains a "Runner capacity is shared and finite" subsection covering the maxParallel=1 instance-level runner shared with dohertj2/lmxopcua, the ~20-30 min queue latency observed under cross-repo contention, and Gitea 1.26's missing run cancel/delete API. The existing "windev tier down" degraded-mode paragraph now also covers "runner contended" as a reason to bypass the queue via CI_SHA=<sha> scripts/ci/run-windev-ci.sh <mode> or the manual windev worktree flow, generalizing it per the finding's design note. New operator runbook docs/runbooks/TST-30-second-ci-runner.md carries the actual runner registration (option a: second act_runner instance on 10.100.0.35 with the same container.network: traefik config, recommended; option b: dedicated labelled runner, escalation only; option c: runner on windev, rejected) plus verification steps and the no-cancel caveat. The optional workflow-level concurrency group is documented as unverified -- framed as "verify before relying on it" -- and left unimplemented in ci.yml, since registering the runner and any runs-on gating is operator/infra work outside this repo's tree. Tracking: TST-30 -> Done (doc half; runner registration operator-pending) in both registers + change-log row.
This commit is contained in:
@@ -0,0 +1,113 @@
|
||||
# TST-30 — Register A Second CI Runner (Operator Runbook)
|
||||
|
||||
Operator steps to relieve the single shared Gitea Actions runner that CI depends on. The
|
||||
repo-side half of TST-30 (documenting the shared-runner/no-cancel reality and the
|
||||
`run-windev-ci.sh` bypass) is already landed in `docs/GatewayTesting.md`; registering the
|
||||
second runner below is infrastructure work outside this repo's tree and is yours to execute.
|
||||
|
||||
## Why
|
||||
|
||||
All CI for this repo runs on one co-located `gitea-runner` container on docker host
|
||||
`10.100.0.35` with `maxParallel=1`. That runner is registered at the **instance** level, not
|
||||
scoped to this repo (`GET /repos/dohertj2/mxaccessgw/actions/runners` returns
|
||||
`total_count: 0`), so it is shared with `dohertj2/lmxopcua` and every job in every run across
|
||||
both repos executes serially on the single slot. A `mxaccessgw` push fans out to `portable`,
|
||||
`java`, `windows-x86`, and an active `lmxopcua` run blocks all of them — queue depth of
|
||||
~20–30 minutes was observed during TST-25 acceptance under cross-repo contention. Gitea 1.26
|
||||
also exposes **no run cancel or delete via the API**
|
||||
(`POST .../actions/runs/{id}/cancel` → 404, `DELETE .../actions/runs/{id}` → 400), so a
|
||||
superseded or hung run cannot be cleared and holds the slot until it finishes or times out.
|
||||
This is not a correctness problem — every job still reports accurately — but it undercuts the
|
||||
fast-feedback purpose of the TST-25 Windows tier and makes CI fragile to a single host: if
|
||||
`10.100.0.35` wedges or goes down, CI for both repos stops with no failover.
|
||||
|
||||
## Options (cheapest first)
|
||||
|
||||
- **(a) Register a second `act_runner` instance on `10.100.0.35` — recommended.** The host
|
||||
already runs `gitea-runner`; add a second `act_runner` container (or raise the existing
|
||||
runner's `maxParallel` where the docker-in-docker/resource budget allows) so at least two
|
||||
jobs run concurrently. Cheapest change, and it keeps the runner co-located on the
|
||||
`container.network: traefik` network that resolves `gitea:3000` — the property TST-03
|
||||
depended on. **Use the same `container.network: traefik` config as the existing runner.**
|
||||
- **(b) Dedicate a labelled runner to `mxaccessgw`.** Cleaner isolation — `lmxopcua` load
|
||||
never blocks this repo — but needs label wiring: register the new runner with a distinct
|
||||
label (e.g. `mxgw`) and change `.gitea/workflows/ci.yml`'s `runs-on:` for this repo's jobs
|
||||
to gate on that label (e.g. `runs-on: [ubuntu-latest, mxgw]`). Only do this if (a) proves
|
||||
insufficient — it is more moving parts for the same throughput gain, and it means `ci.yml`
|
||||
changes, which is out of scope for the doc-only half of TST-30.
|
||||
- **(c) Put the runner on windev / a second host — rejected as the primary fix.** windev is
|
||||
the Windows build target (`10.100.0.48`), not a CI host, and co-locating a Linux runner
|
||||
there loses the `gitea:3000` name resolution TST-03 relies on. Only consider if
|
||||
`10.100.0.35` genuinely runs out of capacity for a second instance.
|
||||
|
||||
Default to **(a)**. Escalate to (b) only if `lmxopcua` contention persists after a second
|
||||
instance is online (i.e., (a) is not sufficient because the two repos' combined load exceeds
|
||||
two slots).
|
||||
|
||||
## Preconditions
|
||||
|
||||
- SSH/docker access to `10.100.0.35`.
|
||||
- The existing `gitea-runner` container's compose/run config, to copy its
|
||||
`container.network: traefik` setting and registration token flow (repo memory
|
||||
`project_gitea_ci` records this configuration).
|
||||
- Admin access to Gitea (`gitea.dohertylan.com`) to mint a new runner registration token.
|
||||
|
||||
## Steps — option (a): second runner instance
|
||||
|
||||
1. On `10.100.0.35`, locate the existing `gitea-runner` container/compose definition and copy
|
||||
its configuration for a new instance (same `container.network: traefik`, same Docker
|
||||
socket mount if it uses docker-in-docker, a distinct container name/data volume).
|
||||
2. In Gitea, generate a new runner registration token (instance-level, since the existing
|
||||
runner is registered at the instance level too — Admin → Actions → Runners, or
|
||||
`POST /admin/actions/runners/registration-token`).
|
||||
3. Register and start the second `act_runner` instance with that token, pointed at the same
|
||||
Gitea origin.
|
||||
4. Confirm both runners show online: instance runner list in the Gitea admin UI, or the
|
||||
equivalent API listing.
|
||||
|
||||
## Verification
|
||||
|
||||
- Push two branches to `mxaccessgw` back-to-back (or trigger one `mxaccessgw` push while an
|
||||
`lmxopcua` run is in flight) and confirm both runs execute **concurrently**, not serially —
|
||||
the second run's jobs should start before the first finishes, not queue behind it.
|
||||
- `GET /repos/dohertj2/mxaccessgw/actions/runners` (or the instance runner listing) shows
|
||||
**≥2** runners online.
|
||||
- Re-run the TST-25 acceptance push (a plain push to a scratch branch) and confirm queue depth
|
||||
is materially lower than the ~20–30 minute baseline observed under a concurrent `lmxopcua`
|
||||
run.
|
||||
- Confirm `windows-x86` still resolves `gitea:3000` correctly from a job scheduled on the new
|
||||
runner instance (the `traefik` network property must hold for both instances).
|
||||
|
||||
## The no-cancel reality does not go away
|
||||
|
||||
A second runner relieves contention; it does not add a cancel/delete API — Gitea 1.26 still
|
||||
returns 404/400 for both. A stale or hung run on either runner still holds its slot until it
|
||||
finishes or times out. Two runners just means one stale run blocks at most half the capacity
|
||||
instead of all of it. Do not treat the second runner as a substitute for the escape hatch: a
|
||||
specific commit can still be verified out of band without waiting on either runner via
|
||||
`CI_SHA=<sha> scripts/ci/run-windev-ci.sh <build|test|live>` (Linux, needs SSH access to
|
||||
windev) or the manual windev worktree flow — see the "Runner capacity is shared and finite"
|
||||
section in `docs/GatewayTesting.md`.
|
||||
|
||||
## Optional: workflow-level `concurrency` group
|
||||
|
||||
As belt-and-suspenders against the missing cancel API, `.gitea/workflows/ci.yml` could add a
|
||||
top-level `concurrency` group (e.g. keyed on `${{ github.ref }}`) so a newer push to the same
|
||||
branch automatically supersedes an in-flight run instead of both running to completion.
|
||||
**Verify this Gitea deployment actually honors `concurrency` and cancels the superseded run
|
||||
before relying on it** — Gitea Actions' YAML surface does not track GitHub Actions feature
|
||||
parity release-for-release, and a `concurrency` block that is silently ignored would look like
|
||||
a working safeguard while doing nothing. If verified working, this is a `ci.yml` change (not
|
||||
covered by this runbook) and should land as its own small change with its own verification
|
||||
(push twice to the same branch quickly, confirm the first run's jobs cancel).
|
||||
|
||||
## Done criteria
|
||||
|
||||
- A second `act_runner` instance (or raised `maxParallel`) is online on `10.100.0.35` with the
|
||||
same `container.network: traefik` configuration as the existing runner.
|
||||
- `GET /repos/dohertj2/mxaccessgw/actions/runners` (or the instance listing) shows ≥2 runners.
|
||||
- Two concurrent runs (one `mxaccessgw`, one `lmxopcua`, or two `mxaccessgw` pushes) execute
|
||||
in parallel rather than serially.
|
||||
- `docs/GatewayTesting.md`'s shared-runner/no-cancel prose and the `run-windev-ci.sh` bypass
|
||||
remain accurate (they describe the bypass as still valid, which it is regardless of runner
|
||||
count).
|
||||
Reference in New Issue
Block a user