Files
mxaccessgw/docs/runbooks/TST-30-second-ci-runner.md
T
Joseph Doherty b604fed72b docs(TST-30): document shared-runner CI bottleneck + second-runner operator runbook
Doc half of TST-30 (single shared Gitea runner is a CI throughput/availability
bottleneck): docs/GatewayTesting.md's Continuous Integration section gains a
"Runner capacity is shared and finite" subsection covering the maxParallel=1
instance-level runner shared with dohertj2/lmxopcua, the ~20-30 min queue
latency observed under cross-repo contention, and Gitea 1.26's missing run
cancel/delete API. The existing "windev tier down" degraded-mode paragraph now
also covers "runner contended" as a reason to bypass the queue via
CI_SHA=<sha> scripts/ci/run-windev-ci.sh <mode> or the manual windev worktree
flow, generalizing it per the finding's design note.

New operator runbook docs/runbooks/TST-30-second-ci-runner.md carries the
actual runner registration (option a: second act_runner instance on
10.100.0.35 with the same container.network: traefik config, recommended;
option b: dedicated labelled runner, escalation only; option c: runner on
windev, rejected) plus verification steps and the no-cancel caveat. The
optional workflow-level concurrency group is documented as unverified --
framed as "verify before relying on it" -- and left unimplemented in ci.yml,
since registering the runner and any runs-on gating is operator/infra work
outside this repo's tree.

Tracking: TST-30 -> Done (doc half; runner registration operator-pending) in
both registers + change-log row.
2026-08-07 07:47:37 -04:00

7.1 KiB
Raw Blame History

TST-30 — Register A Second CI Runner (Operator Runbook)

Operator steps to relieve the single shared Gitea Actions runner that CI depends on. The repo-side half of TST-30 (documenting the shared-runner/no-cancel reality and the run-windev-ci.sh bypass) is already landed in docs/GatewayTesting.md; registering the second runner below is infrastructure work outside this repo's tree and is yours to execute.

Why

All CI for this repo runs on one co-located gitea-runner container on docker host 10.100.0.35 with maxParallel=1. That runner is registered at the instance level, not scoped to this repo (GET /repos/dohertj2/mxaccessgw/actions/runners returns total_count: 0), so it is shared with dohertj2/lmxopcua and every job in every run across both repos executes serially on the single slot. A mxaccessgw push fans out to portable, java, windows-x86, and an active lmxopcua run blocks all of them — queue depth of ~2030 minutes was observed during TST-25 acceptance under cross-repo contention. Gitea 1.26 also exposes no run cancel or delete via the API (POST .../actions/runs/{id}/cancel → 404, DELETE .../actions/runs/{id} → 400), so a superseded or hung run cannot be cleared and holds the slot until it finishes or times out. This is not a correctness problem — every job still reports accurately — but it undercuts the fast-feedback purpose of the TST-25 Windows tier and makes CI fragile to a single host: if 10.100.0.35 wedges or goes down, CI for both repos stops with no failover.

Options (cheapest first)

  • (a) Register a second act_runner instance on 10.100.0.35 — recommended. The host already runs gitea-runner; add a second act_runner container (or raise the existing runner's maxParallel where the docker-in-docker/resource budget allows) so at least two jobs run concurrently. Cheapest change, and it keeps the runner co-located on the container.network: traefik network that resolves gitea:3000 — the property TST-03 depended on. Use the same container.network: traefik config as the existing runner.
  • (b) Dedicate a labelled runner to mxaccessgw. Cleaner isolation — lmxopcua load never blocks this repo — but needs label wiring: register the new runner with a distinct label (e.g. mxgw) and change .gitea/workflows/ci.yml's runs-on: for this repo's jobs to gate on that label (e.g. runs-on: [ubuntu-latest, mxgw]). Only do this if (a) proves insufficient — it is more moving parts for the same throughput gain, and it means ci.yml changes, which is out of scope for the doc-only half of TST-30.
  • (c) Put the runner on windev / a second host — rejected as the primary fix. windev is the Windows build target (10.100.0.48), not a CI host, and co-locating a Linux runner there loses the gitea:3000 name resolution TST-03 relies on. Only consider if 10.100.0.35 genuinely runs out of capacity for a second instance.

Default to (a). Escalate to (b) only if lmxopcua contention persists after a second instance is online (i.e., (a) is not sufficient because the two repos' combined load exceeds two slots).

Preconditions

  • SSH/docker access to 10.100.0.35.
  • The existing gitea-runner container's compose/run config, to copy its container.network: traefik setting and registration token flow (repo memory project_gitea_ci records this configuration).
  • Admin access to Gitea (gitea.dohertylan.com) to mint a new runner registration token.

Steps — option (a): second runner instance

  1. On 10.100.0.35, locate the existing gitea-runner container/compose definition and copy its configuration for a new instance (same container.network: traefik, same Docker socket mount if it uses docker-in-docker, a distinct container name/data volume).
  2. In Gitea, generate a new runner registration token (instance-level, since the existing runner is registered at the instance level too — Admin → Actions → Runners, or POST /admin/actions/runners/registration-token).
  3. Register and start the second act_runner instance with that token, pointed at the same Gitea origin.
  4. Confirm both runners show online: instance runner list in the Gitea admin UI, or the equivalent API listing.

Verification

  • Push two branches to mxaccessgw back-to-back (or trigger one mxaccessgw push while an lmxopcua run is in flight) and confirm both runs execute concurrently, not serially — the second run's jobs should start before the first finishes, not queue behind it.
  • GET /repos/dohertj2/mxaccessgw/actions/runners (or the instance runner listing) shows ≥2 runners online.
  • Re-run the TST-25 acceptance push (a plain push to a scratch branch) and confirm queue depth is materially lower than the ~2030 minute baseline observed under a concurrent lmxopcua run.
  • Confirm windows-x86 still resolves gitea:3000 correctly from a job scheduled on the new runner instance (the traefik network property must hold for both instances).

The no-cancel reality does not go away

A second runner relieves contention; it does not add a cancel/delete API — Gitea 1.26 still returns 404/400 for both. A stale or hung run on either runner still holds its slot until it finishes or times out. Two runners just means one stale run blocks at most half the capacity instead of all of it. Do not treat the second runner as a substitute for the escape hatch: a specific commit can still be verified out of band without waiting on either runner via CI_SHA=<sha> scripts/ci/run-windev-ci.sh <build|test|live> (Linux, needs SSH access to windev) or the manual windev worktree flow — see the "Runner capacity is shared and finite" section in docs/GatewayTesting.md.

Optional: workflow-level concurrency group

As belt-and-suspenders against the missing cancel API, .gitea/workflows/ci.yml could add a top-level concurrency group (e.g. keyed on ${{ github.ref }}) so a newer push to the same branch automatically supersedes an in-flight run instead of both running to completion. Verify this Gitea deployment actually honors concurrency and cancels the superseded run before relying on it — Gitea Actions' YAML surface does not track GitHub Actions feature parity release-for-release, and a concurrency block that is silently ignored would look like a working safeguard while doing nothing. If verified working, this is a ci.yml change (not covered by this runbook) and should land as its own small change with its own verification (push twice to the same branch quickly, confirm the first run's jobs cancel).

Done criteria

  • A second act_runner instance (or raised maxParallel) is online on 10.100.0.35 with the same container.network: traefik configuration as the existing runner.
  • GET /repos/dohertj2/mxaccessgw/actions/runners (or the instance listing) shows ≥2 runners.
  • Two concurrent runs (one mxaccessgw, one lmxopcua, or two mxaccessgw pushes) execute in parallel rather than serially.
  • docs/GatewayTesting.md's shared-runner/no-cancel prose and the run-windev-ci.sh bypass remain accurate (they describe the bypass as still valid, which it is regardless of runner count).