b604fed72b
Doc half of TST-30 (single shared Gitea runner is a CI throughput/availability bottleneck): docs/GatewayTesting.md's Continuous Integration section gains a "Runner capacity is shared and finite" subsection covering the maxParallel=1 instance-level runner shared with dohertj2/lmxopcua, the ~20-30 min queue latency observed under cross-repo contention, and Gitea 1.26's missing run cancel/delete API. The existing "windev tier down" degraded-mode paragraph now also covers "runner contended" as a reason to bypass the queue via CI_SHA=<sha> scripts/ci/run-windev-ci.sh <mode> or the manual windev worktree flow, generalizing it per the finding's design note. New operator runbook docs/runbooks/TST-30-second-ci-runner.md carries the actual runner registration (option a: second act_runner instance on 10.100.0.35 with the same container.network: traefik config, recommended; option b: dedicated labelled runner, escalation only; option c: runner on windev, rejected) plus verification steps and the no-cancel caveat. The optional workflow-level concurrency group is documented as unverified -- framed as "verify before relying on it" -- and left unimplemented in ci.yml, since registering the runner and any runs-on gating is operator/infra work outside this repo's tree. Tracking: TST-30 -> Done (doc half; runner registration operator-pending) in both registers + change-log row.
114 lines
7.1 KiB
Markdown
114 lines
7.1 KiB
Markdown
# TST-30 — Register A Second CI Runner (Operator Runbook)
|
||
|
||
Operator steps to relieve the single shared Gitea Actions runner that CI depends on. The
|
||
repo-side half of TST-30 (documenting the shared-runner/no-cancel reality and the
|
||
`run-windev-ci.sh` bypass) is already landed in `docs/GatewayTesting.md`; registering the
|
||
second runner below is infrastructure work outside this repo's tree and is yours to execute.
|
||
|
||
## Why
|
||
|
||
All CI for this repo runs on one co-located `gitea-runner` container on docker host
|
||
`10.100.0.35` with `maxParallel=1`. That runner is registered at the **instance** level, not
|
||
scoped to this repo (`GET /repos/dohertj2/mxaccessgw/actions/runners` returns
|
||
`total_count: 0`), so it is shared with `dohertj2/lmxopcua` and every job in every run across
|
||
both repos executes serially on the single slot. A `mxaccessgw` push fans out to `portable`,
|
||
`java`, `windows-x86`, and an active `lmxopcua` run blocks all of them — queue depth of
|
||
~20–30 minutes was observed during TST-25 acceptance under cross-repo contention. Gitea 1.26
|
||
also exposes **no run cancel or delete via the API**
|
||
(`POST .../actions/runs/{id}/cancel` → 404, `DELETE .../actions/runs/{id}` → 400), so a
|
||
superseded or hung run cannot be cleared and holds the slot until it finishes or times out.
|
||
This is not a correctness problem — every job still reports accurately — but it undercuts the
|
||
fast-feedback purpose of the TST-25 Windows tier and makes CI fragile to a single host: if
|
||
`10.100.0.35` wedges or goes down, CI for both repos stops with no failover.
|
||
|
||
## Options (cheapest first)
|
||
|
||
- **(a) Register a second `act_runner` instance on `10.100.0.35` — recommended.** The host
|
||
already runs `gitea-runner`; add a second `act_runner` container (or raise the existing
|
||
runner's `maxParallel` where the docker-in-docker/resource budget allows) so at least two
|
||
jobs run concurrently. Cheapest change, and it keeps the runner co-located on the
|
||
`container.network: traefik` network that resolves `gitea:3000` — the property TST-03
|
||
depended on. **Use the same `container.network: traefik` config as the existing runner.**
|
||
- **(b) Dedicate a labelled runner to `mxaccessgw`.** Cleaner isolation — `lmxopcua` load
|
||
never blocks this repo — but needs label wiring: register the new runner with a distinct
|
||
label (e.g. `mxgw`) and change `.gitea/workflows/ci.yml`'s `runs-on:` for this repo's jobs
|
||
to gate on that label (e.g. `runs-on: [ubuntu-latest, mxgw]`). Only do this if (a) proves
|
||
insufficient — it is more moving parts for the same throughput gain, and it means `ci.yml`
|
||
changes, which is out of scope for the doc-only half of TST-30.
|
||
- **(c) Put the runner on windev / a second host — rejected as the primary fix.** windev is
|
||
the Windows build target (`10.100.0.48`), not a CI host, and co-locating a Linux runner
|
||
there loses the `gitea:3000` name resolution TST-03 relies on. Only consider if
|
||
`10.100.0.35` genuinely runs out of capacity for a second instance.
|
||
|
||
Default to **(a)**. Escalate to (b) only if `lmxopcua` contention persists after a second
|
||
instance is online (i.e., (a) is not sufficient because the two repos' combined load exceeds
|
||
two slots).
|
||
|
||
## Preconditions
|
||
|
||
- SSH/docker access to `10.100.0.35`.
|
||
- The existing `gitea-runner` container's compose/run config, to copy its
|
||
`container.network: traefik` setting and registration token flow (repo memory
|
||
`project_gitea_ci` records this configuration).
|
||
- Admin access to Gitea (`gitea.dohertylan.com`) to mint a new runner registration token.
|
||
|
||
## Steps — option (a): second runner instance
|
||
|
||
1. On `10.100.0.35`, locate the existing `gitea-runner` container/compose definition and copy
|
||
its configuration for a new instance (same `container.network: traefik`, same Docker
|
||
socket mount if it uses docker-in-docker, a distinct container name/data volume).
|
||
2. In Gitea, generate a new runner registration token (instance-level, since the existing
|
||
runner is registered at the instance level too — Admin → Actions → Runners, or
|
||
`POST /admin/actions/runners/registration-token`).
|
||
3. Register and start the second `act_runner` instance with that token, pointed at the same
|
||
Gitea origin.
|
||
4. Confirm both runners show online: instance runner list in the Gitea admin UI, or the
|
||
equivalent API listing.
|
||
|
||
## Verification
|
||
|
||
- Push two branches to `mxaccessgw` back-to-back (or trigger one `mxaccessgw` push while an
|
||
`lmxopcua` run is in flight) and confirm both runs execute **concurrently**, not serially —
|
||
the second run's jobs should start before the first finishes, not queue behind it.
|
||
- `GET /repos/dohertj2/mxaccessgw/actions/runners` (or the instance runner listing) shows
|
||
**≥2** runners online.
|
||
- Re-run the TST-25 acceptance push (a plain push to a scratch branch) and confirm queue depth
|
||
is materially lower than the ~20–30 minute baseline observed under a concurrent `lmxopcua`
|
||
run.
|
||
- Confirm `windows-x86` still resolves `gitea:3000` correctly from a job scheduled on the new
|
||
runner instance (the `traefik` network property must hold for both instances).
|
||
|
||
## The no-cancel reality does not go away
|
||
|
||
A second runner relieves contention; it does not add a cancel/delete API — Gitea 1.26 still
|
||
returns 404/400 for both. A stale or hung run on either runner still holds its slot until it
|
||
finishes or times out. Two runners just means one stale run blocks at most half the capacity
|
||
instead of all of it. Do not treat the second runner as a substitute for the escape hatch: a
|
||
specific commit can still be verified out of band without waiting on either runner via
|
||
`CI_SHA=<sha> scripts/ci/run-windev-ci.sh <build|test|live>` (Linux, needs SSH access to
|
||
windev) or the manual windev worktree flow — see the "Runner capacity is shared and finite"
|
||
section in `docs/GatewayTesting.md`.
|
||
|
||
## Optional: workflow-level `concurrency` group
|
||
|
||
As belt-and-suspenders against the missing cancel API, `.gitea/workflows/ci.yml` could add a
|
||
top-level `concurrency` group (e.g. keyed on `${{ github.ref }}`) so a newer push to the same
|
||
branch automatically supersedes an in-flight run instead of both running to completion.
|
||
**Verify this Gitea deployment actually honors `concurrency` and cancels the superseded run
|
||
before relying on it** — Gitea Actions' YAML surface does not track GitHub Actions feature
|
||
parity release-for-release, and a `concurrency` block that is silently ignored would look like
|
||
a working safeguard while doing nothing. If verified working, this is a `ci.yml` change (not
|
||
covered by this runbook) and should land as its own small change with its own verification
|
||
(push twice to the same branch quickly, confirm the first run's jobs cancel).
|
||
|
||
## Done criteria
|
||
|
||
- A second `act_runner` instance (or raised `maxParallel`) is online on `10.100.0.35` with the
|
||
same `container.network: traefik` configuration as the existing runner.
|
||
- `GET /repos/dohertj2/mxaccessgw/actions/runners` (or the instance listing) shows ≥2 runners.
|
||
- Two concurrent runs (one `mxaccessgw`, one `lmxopcua`, or two `mxaccessgw` pushes) execute
|
||
in parallel rather than serially.
|
||
- `docs/GatewayTesting.md`'s shared-runner/no-cancel prose and the `run-windev-ci.sh` bypass
|
||
remain accurate (they describe the bypass as still valid, which it is regardless of runner
|
||
count).
|