The local act_runner on this Mac registered as instance runner id 4 with ubuntu-latest/22.04/20.04 labels, so it competed with the two docker runners on 10.100.0.35 for Linux jobs it had no Docker daemon to run -- 13 of the last 20-run window in historiangw landed on it and all but one failed. Registration deleted; local config kept disabled for re-use with mac-specific labels.
9.0 KiB
TST-30 — Register A Second CI Runner (Operator Runbook)
Executed 2026-08-07 — option (a) shipped; this runbook is now history plus the one correction below.
gitea-runner-2(runner id 5, capacity 2, labelsubuntu-latest/ubuntu-22.04) runs on10.100.0.35from the/opt/giteacompose stack with the samecontainer.network: traefiksetting as the original; its registration token is mounted from a0600file rather than inlined in compose. The existinggitea-runner(id 1, capacity 4) was left untouched, so capacity went 4 → 6 by addition and the change reverts by removing one container. Concurrency was verified by pushing HEAD to two scratch branches while an unrelated run was in flight: jobs from three runs ran simultaneously across both runners, and agitea-runner-2job cloned successfully fromhttp://gitea:3000(the property option (c) was rejected for losing).Correction to the Verification and Done-criteria sections below: they expect
GET /repos/dohertj2/mxaccessgw/actions/runnersto show ≥2 runners. It does not — it still returnstotal_count: 0, correctly, because both runners are registered at the instance level, which is the very condition the "Why" section describes. UseGET /api/v1/admin/actions/runnersinstead (it now lists only id 1 and id 5; id 4 was removed 2026-08-07 — a local macOSact_runnermislabelledubuntu-latest/ubuntu-22.04/ubuntu-20.04, so it captured Linux-labelled jobs it had no Docker daemon to run and failed them. Its config and registration are kept disabled at~/gitea-act-runner.disabled-2026-08-07(launchd plist at~/Library/LaunchAgents/com.dohertj2.gitea-act-runner.plist.disabled-2026-08-07) so it can be re-registered with mac-specific labels if a mac-only job ever needs one). For per-job runner attribution,GET /repos/{owner}/{repo}/actions/runs/{id}/jobsexposesrunner_id/runner_nameon each job; theactions/taskslisting does not.
Operator steps to relieve the single shared Gitea Actions runner that CI depends on. The
repo-side half of TST-30 (documenting the shared-runner/no-cancel reality and the
run-windev-ci.sh bypass) is already landed in docs/GatewayTesting.md; registering the
second runner below is infrastructure work outside this repo's tree and is yours to execute.
Why
All CI for this repo runs on one co-located gitea-runner container on docker host
10.100.0.35 with maxParallel=1. That runner is registered at the instance level, not
scoped to this repo (GET /repos/dohertj2/mxaccessgw/actions/runners returns
total_count: 0), so it is shared with dohertj2/lmxopcua and every job in every run across
both repos executes serially on the single slot. A mxaccessgw push fans out to portable,
java, windows-x86, and an active lmxopcua run blocks all of them — queue depth of
~20–30 minutes was observed during TST-25 acceptance under cross-repo contention. Gitea 1.26
also exposes no run cancel or delete via the API
(POST .../actions/runs/{id}/cancel → 404, DELETE .../actions/runs/{id} → 400), so a
superseded or hung run cannot be cleared and holds the slot until it finishes or times out.
This is not a correctness problem — every job still reports accurately — but it undercuts the
fast-feedback purpose of the TST-25 Windows tier and makes CI fragile to a single host: if
10.100.0.35 wedges or goes down, CI for both repos stops with no failover.
Options (cheapest first)
- (a) Register a second
act_runnerinstance on10.100.0.35— recommended. The host already runsgitea-runner; add a secondact_runnercontainer (or raise the existing runner'smaxParallelwhere the docker-in-docker/resource budget allows) so at least two jobs run concurrently. Cheapest change, and it keeps the runner co-located on thecontainer.network: traefiknetwork that resolvesgitea:3000— the property TST-03 depended on. Use the samecontainer.network: traefikconfig as the existing runner. - (b) Dedicate a labelled runner to
mxaccessgw. Cleaner isolation —lmxopcuaload never blocks this repo — but needs label wiring: register the new runner with a distinct label (e.g.mxgw) and change.gitea/workflows/ci.yml'sruns-on:for this repo's jobs to gate on that label (e.g.runs-on: [ubuntu-latest, mxgw]). Only do this if (a) proves insufficient — it is more moving parts for the same throughput gain, and it meansci.ymlchanges, which is out of scope for the doc-only half of TST-30. - (c) Put the runner on windev / a second host — rejected as the primary fix. windev is
the Windows build target (
10.100.0.48), not a CI host, and co-locating a Linux runner there loses thegitea:3000name resolution TST-03 relies on. Only consider if10.100.0.35genuinely runs out of capacity for a second instance.
Default to (a). Escalate to (b) only if lmxopcua contention persists after a second
instance is online (i.e., (a) is not sufficient because the two repos' combined load exceeds
two slots).
Preconditions
- SSH/docker access to
10.100.0.35. - The existing
gitea-runnercontainer's compose/run config, to copy itscontainer.network: traefiksetting and registration token flow (repo memoryproject_gitea_cirecords this configuration). - Admin access to Gitea (
gitea.dohertylan.com) to mint a new runner registration token.
Steps — option (a): second runner instance
- On
10.100.0.35, locate the existinggitea-runnercontainer/compose definition and copy its configuration for a new instance (samecontainer.network: traefik, same Docker socket mount if it uses docker-in-docker, a distinct container name/data volume). - In Gitea, generate a new runner registration token (instance-level, since the existing
runner is registered at the instance level too — Admin → Actions → Runners, or
POST /admin/actions/runners/registration-token). - Register and start the second
act_runnerinstance with that token, pointed at the same Gitea origin. - Confirm both runners show online: instance runner list in the Gitea admin UI, or the equivalent API listing.
Verification
- Push two branches to
mxaccessgwback-to-back (or trigger onemxaccessgwpush while anlmxopcuarun is in flight) and confirm both runs execute concurrently, not serially — the second run's jobs should start before the first finishes, not queue behind it. GET /repos/dohertj2/mxaccessgw/actions/runners(or the instance runner listing) shows ≥2 runners online.- Re-run the TST-25 acceptance push (a plain push to a scratch branch) and confirm queue depth
is materially lower than the ~20–30 minute baseline observed under a concurrent
lmxopcuarun. - Confirm
windows-x86still resolvesgitea:3000correctly from a job scheduled on the new runner instance (thetraefiknetwork property must hold for both instances).
The no-cancel reality does not go away
A second runner relieves contention; it does not add a cancel/delete API — Gitea 1.26 still
returns 404/400 for both. A stale or hung run on either runner still holds its slot until it
finishes or times out. Two runners just means one stale run blocks at most half the capacity
instead of all of it. Do not treat the second runner as a substitute for the escape hatch: a
specific commit can still be verified out of band without waiting on either runner via
CI_SHA=<sha> scripts/ci/run-windev-ci.sh <build|test|live> (Linux, needs SSH access to
windev) or the manual windev worktree flow — see the "Runner capacity is shared and finite"
section in docs/GatewayTesting.md.
Optional: workflow-level concurrency group
As belt-and-suspenders against the missing cancel API, .gitea/workflows/ci.yml could add a
top-level concurrency group (e.g. keyed on ${{ github.ref }}) so a newer push to the same
branch automatically supersedes an in-flight run instead of both running to completion.
Verify this Gitea deployment actually honors concurrency and cancels the superseded run
before relying on it — Gitea Actions' YAML surface does not track GitHub Actions feature
parity release-for-release, and a concurrency block that is silently ignored would look like
a working safeguard while doing nothing. If verified working, this is a ci.yml change (not
covered by this runbook) and should land as its own small change with its own verification
(push twice to the same branch quickly, confirm the first run's jobs cancel).
Done criteria
- A second
act_runnerinstance (or raisedmaxParallel) is online on10.100.0.35with the samecontainer.network: traefikconfiguration as the existing runner. GET /repos/dohertj2/mxaccessgw/actions/runners(or the instance listing) shows ≥2 runners.- Two concurrent runs (one
mxaccessgw, onelmxopcua, or twomxaccessgwpushes) execute in parallel rather than serially. docs/GatewayTesting.md's shared-runner/no-cancel prose and therun-windev-ci.shbypass remain accurate (they describe the bypass as still valid, which it is regardless of runner count).