Files
mxaccessgw/.gitea/workflows/ci.yml
T
Joseph Doherty d769244ac0
ci / java (push) Successful in 2m2s
ci / windows-x86 (push) Successful in 1m10s
ci / nightly-windev (push) Has been skipped
ci / portable (push) Successful in 8m52s
fix(TST-25,TST-26): restore Windows/x86 CI tier via SSH-driven windev job
The `windows`/`live-mxaccess` CI jobs were removed in abb0930 because Gitea
act_runner host-mode on Windows is broken and a runs-on gate with no runner
wedges the queue. This left the entire x86/net48 Worker + Worker.Tests tier
(and the live-MXAccess smoke) unguarded by CI.

Restore it without a native Windows runner: Linux jobs on ubuntu-latest (which
always schedule) SSH to windev (10.100.0.48), fetch+checkout the SHA under test
in an isolated clone C:\build\mxaccessgw-ci, run the x86 build/tests there, and
propagate the remote exit code back — so a Worker regression turns the job red
and an unreachable host fails loud (never stuck).

- scripts/ci/run-windev-ci.sh: Linux driver (key/known-hosts, UTF-16LE base64
  EncodedCommand bootstrap, ssh, exit-code passthrough).
- scripts/ci/windev-worker-ci.ps1: windev stage — worktree lock, re-fetch/
  checkout under lock, build|test|live modes; PS 5.1-safe, checks $LASTEXITCODE.
- scripts/ci/windev.known_hosts: pinned host keys for StrictHostKeyChecking.
- scripts/ci/README.md: operator bring-up + acceptance checklist.
- ci.yml: windows-x86 (per-push `test`) + nightly-windev (scheduled `live` +
  on-failure Gitea issue); drop the removal note; fix header/cron comments.

TST-26 (same commit, by rule): correct docs/scripts that still describe the
removed jobs — GatewayTesting.md, Contracts.md, check-codegen.ps1 — and
reattribute the Generated/-commit guard (primary = check-codegen diff in the
portable job; secondary = windows-x86 net48 compile).

Mechanism hand-verified on windev: build->0, bogus-SHA->nonzero (lock released),
test->356 passed/0 failed in ~50s, ssh exit-code propagation confirmed. Merge
requires operator bring-up first (dedicated ci@ key + Gitea secrets) or the
per-push job is red every push — see scripts/ci/README.md.
2026-07-13 10:09:49 -04:00

176 lines
8.5 KiB
YAML

# Continuous integration for mxaccessgw (Gitea Actions; origin is Gitea at gitea.dohertylan.com).
#
# Scope note: the x86 Worker (src/ZB.MOM.WW.MxGateway.Worker) and Worker.Tests target
# .NET Framework 4.8 / x86 and need MXAccess COM installed, so they build ONLY on a Windows host.
# They are out of scope for this Linux `portable` job — the `windows-x86` job below builds and
# tests them on windev (10.100.0.48) over SSH (see scripts/ci/run-windev-ci.sh), and the nightly
# `nightly-windev` job additionally runs the live-MXAccess smoke, which is scheduled-only so it
# never gates a push.
name: ci
on:
push:
pull_request:
schedule:
# 06:00 UTC nightly: the `nightly-windev` job re-runs the x86 Worker build + full Worker.Tests on
# windev, then the live-MXAccess smoke (MXGATEWAY_RUN_LIVE_MXACCESS_TESTS=1) and a full-slnx build.
# Push/PR runs skip it — the job is gated `if: github.event_name == 'schedule'`.
- cron: '0 6 * * *'
jobs:
portable:
# Any Linux runner: no MXAccess, no x86. Covers the NonWindows solution, gateway fake-worker
# tests, the codegen/descriptor freshness guards, and the clients that build on Linux.
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-dotnet@v4
with:
dotnet-version: '10.0.x'
- uses: actions/setup-go@v5
with:
go-version: '1.26'
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- uses: dtolnay/rust-toolchain@stable
with:
components: clippy, rustfmt
# protoc is pinned to 34.1 (see docs/ToolchainLinks.md) so the committed client descriptor
# regenerates reproducibly. The freshness check normalizes source_code_info, so it tolerates
# patch drift, but keep the CI toolchain on the pin.
- name: Install protoc 34.1
run: |
curl -sSL -o /tmp/protoc.zip https://github.com/protocolbuffers/protobuf/releases/download/v34.1/protoc-34.1-linux-x86_64.zip
sudo unzip -o /tmp/protoc.zip -d /usr/local bin/protoc 'include/*'
protoc --version
- name: Build NonWindows solution
run: dotnet build src/ZB.MOM.WW.MxGateway.NonWindows.slnx -c Release
# GitHub-hosted runners ship pwsh; the self-hosted act image does not. Install it as a
# .NET global tool (the SDK is already set up above) so the codegen check's `shell: pwsh` works.
- name: Install PowerShell (pwsh)
run: |
dotnet tool install --global PowerShell
echo "$HOME/.dotnet/tools" >> "$GITHUB_PATH"
# IPC-01 / IPC-19 / IPC-20: descriptor set + Contracts/Generated must match the current protos.
- name: Codegen / descriptor freshness
shell: pwsh
run: ./scripts/check-codegen.ps1
- name: Gateway fake-worker tests
run: dotnet test src/ZB.MOM.WW.MxGateway.Tests/ZB.MOM.WW.MxGateway.Tests.csproj -c Release --no-build
- name: .NET client
run: dotnet build clients/dotnet/ZB.MOM.WW.MxGateway.Client.slnx -c Release
- name: Go client
working-directory: clients/go
run: |
test -z "$(gofmt -l .)" || (gofmt -l . && echo 'gofmt needed' && exit 1)
go build ./...
go test ./...
- name: Rust client
working-directory: clients/rust
run: |
cargo fmt --all -- --check
cargo test --workspace
cargo clippy --workspace --all-targets -- -D warnings
- name: Python client
working-directory: clients/python
run: |
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"
python -m pytest
java:
# Java client runs on a JDK-17 Linux runner (the macOS dev box has no JRE). The protobuf gradle
# plugin rewrites MxaccessGateway.java with spurious protobuf-runtime-version churn on every
# build; when no .proto changed, revert that one file so checkGeneratedClean / a dirty tree does
# not fail the build (repo memory project_java_generated_churn).
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '17'
# GitHub-hosted runners ship gradle on PATH; the self-hosted act image does not, and the repo
# has no gradle wrapper. act also can't resolve the gradle/actions monorepo action, so install
# gradle directly (pinned to 9.5.1, matching the local homebrew build in CLAUDE.md/memory).
- name: Install Gradle 9.5.1
run: |
curl -sSL "https://services.gradle.org/distributions/gradle-9.5.1-bin.zip" -o /tmp/gradle.zip
sudo unzip -q -d /opt/gradle /tmp/gradle.zip
echo "/opt/gradle/gradle-9.5.1/bin" >> "$GITHUB_PATH"
- name: Gradle test
working-directory: clients/java
run: gradle test
- name: Revert spurious protobuf-version churn (no .proto changed)
# Both generated aggregates can pick up protobuf-runtime-version churn on regen; revert
# both so verify-clean still catches a real, uncommitted proto/codegen change elsewhere.
run: |
git checkout -- clients/java/src/main/generated/main/java/mxaccess_gateway/v1/MxaccessGateway.java || true
git checkout -- clients/java/src/main/generated/main/java/mxaccess_worker/v1/MxaccessWorker.java || true
- name: Verify generated tree is clean
run: git diff --exit-code -- clients/java/src/main/generated
windows-x86:
# TST-25: x86 / net48 Worker build + Worker.Tests, executed on windev (10.100.0.48) over SSH.
# Runs on a Linux runner so it ALWAYS schedules — act_runner host-mode on Windows is broken (its
# hostexecutor writes each step's script to a path it cannot resolve, and it cannot stage JS
# actions), and a `runs-on` gate with no matching runner wedges the queue in "queued" forever.
# An unreachable windev fails this job RED (loud), never leaves it stuck. A red job can mean the
# Windows tier is down rather than the change — see docs/GatewayTesting.md § Continuous Integration.
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Worker x86 build + tests on windev
env:
WINDEV_SSH_KEY: ${{ secrets.WINDEV_SSH_KEY }}
WINDEV_SSH_KNOWN_HOSTS: ${{ secrets.WINDEV_SSH_KNOWN_HOSTS }}
WINDEV_SSH_USER: ${{ vars.WINDEV_SSH_USER }}
CI_SHA: ${{ github.sha }}
# `test` = x86 build + Worker.Tests (~50s measured on windev). Demote to `build` only if the
# test step ever exceeds ~10 min; the x86 build alone still guards the net48/CS0246/x86 class.
run: ./scripts/ci/run-windev-ci.sh test
nightly-windev:
# TST-25: scheduled deep run — full Worker.Tests + live-MXAccess smoke + full-slnx build on windev.
# Gated to the schedule event via `if:` (not an unsatisfiable `runs-on`), so push/PR runs mark it
# skipped, not queued. Never gates a push. On failure it opens a Gitea issue so a red nightly is
# visible even though nobody watches the Actions page.
if: github.event_name == 'schedule'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Full Worker.Tests + live-MXAccess smoke on windev
env:
WINDEV_SSH_KEY: ${{ secrets.WINDEV_SSH_KEY }}
WINDEV_SSH_KNOWN_HOSTS: ${{ secrets.WINDEV_SSH_KNOWN_HOSTS }}
WINDEV_SSH_USER: ${{ vars.WINDEV_SSH_USER }}
CI_SHA: ${{ github.sha }}
run: ./scripts/ci/run-windev-ci.sh live
- name: Report a red nightly as a Gitea issue
if: failure()
run: |
curl -sSf -X POST \
-H "Authorization: token ${{ github.token }}" \
-H "Content-Type: application/json" \
"${{ github.server_url }}/api/v1/repos/${{ github.repository }}/issues" \
-d "{\"title\":\"nightly-windev failed on ${{ github.sha }}\",\"body\":\"The scheduled windev Worker + live-MXAccess run failed: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }} . A red nightly may mean the Windows tier is down rather than the change — see docs/GatewayTesting.md (Continuous Integration).\"}"
# NOTE: there is intentionally no native `windows` runner job. act_runner v0.6.1 host-mode on
# Windows is broken and Windows containers are impractical for the net48/x86/MXAccess Worker, so the
# Windows tier runs via the SSH-driven `windows-x86` (per-push) and `nightly-windev` (scheduled)
# jobs above. Restore a native job only alongside a Windows runner that actually works.