Skip to content

a Java servant invokes a reference it was handed, and §1's last work … #272

a Java servant invokes a reference it was handed, and §1's last work …

a Java servant invokes a reference it was handed, and §1's last work … #272

Workflow file for this run

name: CI
# Licensing: omniORB, TAO and JacORB are LGPL/GPL/DOC. They are installed into
# the ephemeral runner and used only as (a) external programs whose text output
# we read and (b) separate-process wire peers over TCP. Nothing here links them
# into Orbweaver, and no image or artifact containing them is published —
# publishing would be redistribution. See docs/PLAN.md §10.
#
# Installs are best-effort (`|| true`) on purpose: the harness scripts decide
# whether an oracle or peer is present, because they are the ones that know
# whether its absence is a skip or a failure. An apt step that aborts the job
# tells you a package name was wrong and nothing about the code.
on:
push:
branches: [main]
pull_request:
workflow_dispatch:
# **Only the newest push to a ref runs to completion.** Measured 2026-08-31 over
# this repository's own last 100 runs — all 100 of them `push`, and there has
# never been a pull request: 4,679 wall minutes, of which **35 runs were still
# going when the next push arrived** and 2,105 minutes (45%) were spent finishing
# work a newer commit had already superseded.
#
# This is not a saving at the cost of coverage. Today a superseded run finishes
# AND the newer one runs; with this, the newer one still runs to completion, so
# every ref's final state is verified exactly as before. What stops is paying to
# learn the verdict of a commit nobody will look at again — and, incidentally,
# the thing that made HEAD carry no verdict at all on 2026-08-30, when two runs
# were cut by something else while the older one kept going.
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
env:
CARGO_TERM_COLOR: always
jobs:
# ── Everything that needs no fixture. Fast, and gates the rest. ────────────
rust:
name: build, lint, test
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/cache@v4
with:
path: |
~/.cargo/registry
~/.cargo/git
target
key: cargo-${{ runner.os }}-${{ hashFiles('**/Cargo.lock', '**/Cargo.toml') }}
restore-keys: cargo-${{ runner.os }}-
- name: rustfmt
run: cargo fmt --all --check
- name: clippy
run: cargo clippy --workspace --all-targets -- -D warnings
- name: tests
run: cargo test --workspace
env:
RUSTFLAGS: -D warnings
# NOTICE promises that --no-default-features drops encoding_rs and the
# BSD-3-Clause attribution with it. A promise that is testable is tested.
- name: attribution-free build
run: cargo test -p orbweaver-giop --lib --no-default-features
env:
RUSTFLAGS: -D warnings
# The licensing boundary is the one rule that cannot be renegotiated
# later, so it is checked mechanically rather than trusted to review.
- name: licence boundary
run: |
# The names are NOT written here. They used to be, and on 2026-08-27
# this copy and the harness's had drifted to different patterns: this
# one carried `\btao\b` and the harness's did not — and the harness
# is the copy that runs on the machine where a TAO fixture gets built,
# which D035 approved that day as its second step. Two more copies
# turned up in the same sweep, in `spikes/jacorb_python_servant.sh`
# and `spikes/bindings/java/client-jacorb.sh`, which is what a rule
# restated four times does.
#
# `spikes/licence_boundary.sh` owns the names, the producer-status
# check and the herestring match. Its `--self-test` earns the silence
# before anyone reads it: a pattern that matches nothing over a clean
# tree is indistinguishable from one that cannot match.
if ! ./spikes/licence_boundary.sh --self-test; then
echo "::error::the licence-boundary pattern failed its own self-test"; exit 1
fi
lb_out=$(./spikes/licence_boundary.sh); lb_rc=$?
echo "$lb_out"
if [ "$lb_rc" -eq 1 ]; then
echo "::error::an ORB fixture has become a dependency"; exit 1
elif [ "$lb_rc" -ne 0 ]; then
echo "::error::cargo tree --workspace did not run — the licence boundary was NOT measured"; exit 1
fi
# NOTICE's other promise, which is about a feature rather than a
# fixture and so is not the shared gate's business.
if ! ndf_out=$(cargo tree -p orbweaver-giop --no-default-features); then
echo "::error::cargo tree --no-default-features did not run — NOTICE's promise was NOT measured"; exit 1
fi
if grep -q encoding_rs <<<"$ndf_out"; then
echo "::error::--no-default-features still pulls encoding_rs; NOTICE is wrong"; exit 1
fi
# ── The differential oracle: two independent deployed IDL compilers. ───────
differential:
name: differential conformance (omniidl + JacORB)
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/cache@v4
with:
path: |
~/.cargo/registry
~/.cargo/git
target
key: cargo-diff-${{ runner.os }}-${{ hashFiles('**/Cargo.lock', '**/Cargo.toml') }}
restore-keys: cargo-diff-${{ runner.os }}-
- uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '21'
# TAO is deliberately not installed: Ubuntu has no tao-idl package, which
# the first run of this workflow established rather than assumed. The
# script picks it up if a future runner has it; JacORB's IDL compiler is
# the second oracle in the meantime, and being an independent Java
# implementation it serves the same purpose.
- name: install omniidl
run: sudo apt-get update -qq && sudo apt-get install -y --no-install-recommends omniidl || true
- name: fetch the JacORB IDL compiler
env:
JAVA_HOME_21: ${{ env.JAVA_HOME }}
run: ./spikes/jacorb/setup.sh --jars-only
# --require makes a missing oracle a failure rather than a skip. On a
# laptop one oracle is normal; in CI an unmeasured column is a failure,
# because a harness that reports green on something it did not run is
# worse than no harness.
- name: differential conformance
env:
JAVA_HOME_21: ${{ env.JAVA_HOME }}
run: ./spikes/differential.sh --require omniidl,jacorb_idl
# ── The wire: two independent ORBs, both directions. ──────────────────────
interop:
name: interop harness (omniORB + JacORB)
runs-on: ubuntu-latest
# **A stuck run, not a slow one.** Measured 2026-08-31: the last 25
# runs' successful interop jobs took 22, 23 (x8), 24 and 29 minutes,
# and a failing one ran 48. 30 would leave the slowest success one
# minute of headroom, and a timeout that turns a green run red on a
# slow runner is a false red — which kills a gate as surely as a true
# one nobody reads. 45 clears the slowest observed success by 16
# minutes and still cuts the runaway.
timeout-minutes: 45
# Disk discipline, and it belongs here rather than in a cleanup step: the
# cache below restores `target`, and a restored `target` only ever grows —
# cargo evicts nothing, and `restore-keys` falls back to an older, larger
# cache when the exact key misses. Reclaiming 30 GB at the start does not
# touch that; it just moves the crossing later. So the job stops PRODUCING
# what nobody here reads:
#
# CARGO_INCREMENTAL incremental state is for a developer rebuilding
# the same tree twice. A CI runner never does.
# *_DEBUG=0 nobody attaches a debugger to a CI job, and
# debug info is the bulk of a debug `target`.
#
# Measured 2026-08-26 with neither set: `target` reached **35 G** and the
# job ended at **15 G free (90% used)**. The effect of these three is NOT
# yet measured — the next run is its first reading, and the `df` lines
# around the harness are what will say so.
env:
CARGO_INCREMENTAL: 0
CARGO_PROFILE_DEV_DEBUG: 0
CARGO_PROFILE_TEST_DEBUG: 0
steps:
# This job ran the runner out of disk on 2026-08-25 and stayed that way
# for EIGHT consecutive pushes, including the one tagged v0.7.0. It did
# not fail like a check failing: the runner process itself died with
#
# Unhandled exception. System.IO.IOException: No space left on device
# : '/home/runner/actions-runner/cached/.../Worker_*.log'
#
# so the `harness` step is left reading `in_progress` with a null
# conclusion forever, `gh run view --log-failed` answers `log not found`,
# and the only place the cause is written is the check-run annotation.
# A red that takes an API call to read is a red nobody reads, and this one
# meant the interop harness had not run in CI at all for a day — which by
# this project's own rule is an unmeasured check, and an unmeasured check
# is a failure, never a pass. Local runs stayed green throughout, which is
# exactly why it went unnoticed.
#
# ubuntu-latest ships ~30 GB of toolchains this job never touches. Freeing
# them is the cheap half; the load-bearing half is the reporting below, so
# that the NEXT time this crosses, it crosses as a legible number in the
# log rather than as a runner crash.
- name: reclaim runner disk, and report it
run: |
echo "--- before:"; df -h / | tail -1
sudo rm -rf /usr/share/dotnet /usr/local/lib/android /opt/ghc \
/usr/local/share/boost /usr/local/.ghcup \
/usr/share/swift "$AGENT_TOOLSDIRECTORY" || true
sudo docker image prune -af >/dev/null 2>&1 || true
echo "--- after:"; df -h / | tail -1
# A floor, not a figure: this asserts the reclaim happened, and says
# nothing about whether the job will fit. The number the next reader
# wants is the one `df` prints after the harness, below.
avail=$(df -BG --output=avail / | tail -1 | tr -dc '0-9')
echo "available: ${avail}G"
if [ "$avail" -lt 20 ]; then
echo "::error::only ${avail}G free after reclaiming — the interop job"
echo "::error::has historically needed more than this, and a runner"
echo "::error::that dies on ENOSPC reports no failing step at all"
exit 1
fi
- uses: actions/checkout@v4
- uses: actions/cache@v4
with:
path: |
~/.cargo/registry
~/.cargo/git
target
key: cargo-interop-${{ runner.os }}-${{ hashFiles('**/Cargo.lock', '**/Cargo.toml') }}
restore-keys: cargo-interop-${{ runner.os }}-
- uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '21' # JDK 24 removed java.applet.Applet; JacORB 3.9 needs it
# The fixtures are Python servants, so the omniORB Python binding is the
# load-bearing piece. Two runs of this workflow established, rather than
# assumed, that Ubuntu packages no such binding: `apt-cache search omni`
# on noble returns omniORB's C++ libraries, omniidl and the nameserver,
# and nothing else. omniORBpy is therefore built from source against the
# packaged C++ core, at a version matched to it.
- name: install the omniORB peer
run: |
sudo apt-get update -qq
# apt suggests omniidl-python, the IDL compiler's Python back-end,
# but noble carries no installation candidate for it — it has been
# dropped from the archive along with the Python runtime binding.
# The back-end therefore comes out of the omniORBpy source tree in
# the next step.
sudo apt-get install -y --no-install-recommends \
omniidl omniorb-idl omniorb-nameserver \
libomniorb4-dev python3-dev build-essential
dpkg-query -W -f='${Version}\n' libomniorb4-dev
echo "--- where omniidl keeps its back-ends:"
dpkg -L omniidl | grep -iE "omniidl_be" || echo " (omniidl ships no omniidl_be)"
- name: build omniORBpy against the packaged core
run: |
set -x
ver=$(dpkg-query -W -f='${Version}' libomniorb4-dev | sed 's/[+-].*//')
echo "omniORB core is $ver; building the matching omniORBpy"
curl -sfL --max-time 180 -o /tmp/omniORBpy.tar.bz2 \
"https://downloads.sourceforge.net/project/omniorb/omniORBpy/omniORBpy-${ver}/omniORBpy-${ver}.tar.bz2"
mkdir -p /tmp/opy && tar xf /tmp/omniORBpy.tar.bz2 -C /tmp/opy --strip-components=1
cd /tmp/opy
# omniORBpy compiles omniORB's own IDL (corbaidl.idl, ir.idl,
# boxes.idl, pollable.idl) and looks for it under
# $(IMPORT_TREES)/idl, which --with-omniorb=/usr makes /usr/idl.
# Debian ships those files at /usr/share/idl/omniORB, so the build
# failed with "couldn't find idl file corbaidl.idl" until they were
# put where its makefiles look.
sudo mkdir -p /usr/idl
sudo cp -r /usr/share/idl/omniORB /usr/idl/
ls /usr/idl/omniORB | head
./configure --prefix=/usr --with-omniorb=/usr
make -j"$(nproc)"
sudo make install
python3 -c "import omniORB; print('omniORB python at', omniORB.__file__)"
# omniidl loads its back-ends as submodules of one omniidl_be
# package. Debian's sysconfig inserts "local" into the install path,
# so omniORBpy's python back-end lands under /usr/local while the apt
# omniidl imports omniidl_be from /usr/lib/python3/dist-packages. The
# two directories shadow each other rather than merging, and every
# fixture died on "Could not import back-end 'python'". Merge them.
# `make install` did not put omniidl's Python back-end anywhere the
# compiler looks — a filesystem-wide search for omniidl_be after the
# previous round came back empty. The source tree definitely has it,
# so install it from there into every place omniidl might look, and
# print what was found either way.
echo "--- omniidl_be anywhere on the filesystem:"
find / -maxdepth 8 -name 'omniidl_be' 2>/dev/null || true
echo "--- in the source tree:"
ls /tmp/opy/python3/omniidl_be
for be in $(dpkg -L omniidl | grep -oE '.*/omniidl_be' | sort -u) \
/usr/lib/python3/dist-packages/omniidl_be; do
sudo mkdir -p "$be"
sudo cp -r /tmp/opy/python3/omniidl_be/. "$be"/
echo "installed the python back-end into $be"
done
# Not verified with `python3 -c "import omniidl_be.python"`: omniidl
# keeps its own modules in /usr/lib/omniidl, which is not on the
# default sys.path — the compiler puts it there itself. Importing the
# back-end standalone therefore fails even when omniidl works, which
# is how the previous round failed a step whose actual work had
# succeeded. The real check is running the compiler, in the step below.
# The fixtures call omniORB.importIDL, which shells out to omniidl with
# the python back-end. Proving that works here turns a confusing cascade
# of "fixture did not publish an IOR" into one clear failure.
- name: check the fixtures can compile their IDL
run: |
# -C, not -o: omniidl's output-directory flag. (tao_idl uses -o,
# which is why differential.sh spells it differently.)
mkdir -p /tmp/idlcheck
omniidl -bpython -C /tmp/idlcheck spikes/echo.idl
python3 -c "import sys; sys.path.insert(0, '/tmp/idlcheck'); import spike; print('spike stubs generated')"
- name: build the JacORB fixture
env:
JAVA_HOME_21: ${{ env.JAVA_HOME }}
run: ./spikes/jacorb/setup.sh
# The cache can restore `target/*/incremental` from before CARGO_INCREMENTAL
# was set. It regenerates on demand and is never read here, so dropping it
# is free; `|| true` because a cache miss means it does not exist.
- name: drop restored incremental state, and report the margin
run: |
before=$(df -BG --output=avail / | tail -1 | tr -dc '0-9')
du -sh target/*/incremental 2>/dev/null || echo "no incremental state restored"
rm -rf target/*/incremental || true
after=$(df -BG --output=avail / | tail -1 | tr -dc '0-9')
echo "available before the harness: ${after}G (was ${before}G)"
# No floor here on purpose. This is the reading a later run compares
# against, not a gate: the number that decides whether the job fits is
# the one printed AFTER the harness, and inventing a threshold for the
# before-figure would be a number nobody could defend.
# The disk steps above have a sibling this job never had, and its absence
# cost the same kind of blindness twice.
#
# Locally, 2026-08-27: the machine froze at 15:17:50 with no panic report,
# and the only reason the cause could be named afterwards was that the
# kernel happens to stamp `memorystatus_available_pages` onto unrelated
# log lines — they showed available memory falling from 6.42 GB to
# 0.75 GB in 33 seconds and staying under 0.4 GB for eight minutes. That
# is a lucky reconstruction, not a measurement.
#
# On a runner the same blindness has a different face: a job whose runner
# runs out of memory does not report a failing step, it reports
# `The operation was canceled`, and there is nothing in the log to say
# which. So the figure is printed here, and the harness's own trace is
# kept as an artifact by a step that runs even when this one did not
# finish. A report, not a gate — the same reason the `df` lines are.
- name: memory before the harness
run: |
free -m || true
echo "--- MemAvailable:"; awk '/^MemAvailable:/{printf "%d MB\n",$2/1024}' /proc/meminfo
nproc
- name: harness
env:
JAVA_HOME_21: ${{ env.JAVA_HOME }}
# A fixed path so the `always()` steps below can find it whatever the
# harness did. It is the same default the harness picks on its own;
# naming it here means the workflow and the script cannot drift into
# two different files, one of which nobody uploads.
ORBWEAVER_MEMLOG: /tmp/orbweaver-memory.log
run: ./spikes/run_checks.sh
# `always()`, for the same reason as the disk step: the reading that
# matters is the one taken after a run that died, and a cancelled job
# skips everything that is not marked.
- name: what memory the job ran in
if: always()
run: |
echo "--- final:"; free -m || true
echo "--- kernel memory kills during this job:"
./spikes/memlog.sh kills --since "$(( $(date +%s) - 3600 ))" || \
echo "(the kernel's record is unreadable here — unmeasured, not clean)"
echo "--- the harness's own trace, tightest samples last:"
if [ -s /tmp/orbweaver-memory.log ]; then
./spikes/memlog.sh summary --out /tmp/orbweaver-memory.log || true
tail -20 /tmp/orbweaver-memory.log
else
echo "(no trace — the harness did not get far enough to record one)"
fi
- name: keep the memory trace
if: always()
uses: actions/upload-artifact@v4
with:
name: memory-trace-${{ github.run_id }}-${{ github.run_attempt }}
path: |
/tmp/orbweaver-memory.log
/tmp/orbweaver-memory.log.prev
if-no-files-found: ignore
retention-days: 14
# `always()`, because the reading that matters is the one taken after a
# run that died. This is the figure that was missing for eight pushes:
# the failure mode leaves no failing step and no fetchable log, so the
# margin has to be printed by a step that runs even when the one before
# it did not finish. A report, not a gate — the gate is the floor above,
# and this is what tells the next reader whether that floor still holds.
- name: what the job left on the disk
if: always()
run: |
df -h / | tail -1
echo "--- the four that grow:"
du -sh target ~/.cargo /tmp/opy spikes/jacorb 2>/dev/null || true