FROM task/908-cli-first-class TO development - #909
Conversation
oh is meant to split by execution target -- on the host it provisions the host or the sandbox, and inside the sandbox it provisions the sandbox with harnesses and tools. The second half did not work for the two harnesses most people use. claude-code and codex carried installUser: "root", which against the local execution target becomes sudo -n -- npm install -g, and /etc/sudoers.d/sandbox grants sandbox ALL=(ALL) ALL with no NOPASSWD. sudo -n true returns "a password is required". Both now match the pi entry directly above them: installUser "sandbox", npm --prefix /home/sandbox/.local install -g. That lands them inside the home mount, so they also survive container recreate and can be upgraded in place in a running remote sandbox rather than requiring an image rebuild. claude-code deliberately does not get --ignore-scripts. Its postinstall copies the native binary over a placeholder; with the flag the install succeeds and claude --version then fails with "claude native binary not installed". Verified both ways against a scratch prefix. provision-harnesses.sh follows provision-python.sh: the same mode flag, the same root to gosu sandbox re-exec, the same ownership diagnostics, the same die-with- the-command-to-re-run style. --print-env is absent because this provisioner exports nothing downstream. It reads the catalog through oh harness list --json and installs through oh harness install, so the shell knows no ids, packages, prefixes, or argv, and the TypeScript catalog stays the only description. No default harness carries a version pin today, so an existing install is never replaced and the script says so in its own output rather than implying it refreshes. The entrypoint hook runs after link-providers.sh, not before. link-providers' only binary dependency is cc-safety-net, which stays baked, and it is the boot-critical hard gate; a network-dependent best-effort step does not belong in front of the step that decides whether the boot is viable. Provisioning warns and continues, so an offline sandbox still comes up as a usable shell. BAKE_HARNESSES defaults to true and gates only the $AGENTS loop, not the whole RUN. INSTALL_OPENCODE and INSTALL_GROK_BUILD are separate opt-ins and turning them off as a side effect would be a silent regression. Nothing leaves the image in this change.
An adversarial audit of #903 found five defects that six green checks missed. The serious one is a boot hang. oh harness list --json probes every entry in the catalog, not just the three defaults, and one of them is t3code, whose verifyArgv is npx --no-install t3 --version. npx contacts the registry, and probeInstalled passed no timeoutMs, so spawnSync waited without bound. Against an unreachable registry the auditor's run was still going at 2m30 when their own timeout killed it. On any boot where DNS resolves but the registry does not answer, the entrypoint blocks before sleep infinity, exceeds the 300s start_period, and never goes healthy -- and restart: unless-stopped does not rescue an unhealthy-but-alive container. Warn-and-continue cannot help, because a hang never reaches the if !. It hangs while listing, before any install, so BAKE_HARNESSES=true did not avoid it either. Bounded at three layers, because each fails differently: a 15s timeoutMs on the probe spawn, reported as unknown rather than a crash; a --defaults filter on oh harness list so the boot path probes three entries instead of nine and never runs npx; and a timeout wrapper on the entrypoint call so the boot is bounded whatever the CLI does. Measured against the auditor's exact command: 2m30 and killed, to 17.0s full-catalog and 1.55s with --defaults. The install loop read from a herestring while installs run with stdio inherit, so an installer that reads stdin consumed the rest of the loop. Reproduced with a stub: three missing harnesses, one installed, exit 0, success printed. Latent with npm, live the moment a default uses the curl | bash shape two catalog entries already use. Installs now read from /dev/null. The script force-exported OH_EXECUTION_TARGET=local, which short-circuits the in-container check, while the prefix is hardcoded to /home/sandbox/.local, and every error told the operator to re-run with no mention of where. On the host that provisioned the host. It now refuses unless inside the sandbox, reusing the CLI's own runningInsideSandbox predicate rather than inventing a check, and the entrypoint asserts the local target explicitly -- the documented raw docker run recipe never passes SANDBOX_NAME, so the guard would otherwise have silently skipped provisioning for the prebuilt-image flavor. The probe asserted that ARG BAKE_HARNESSES was declared, not that anything used it: deleting the gate left it green. It now checks the ARG is referenced by the RUN that installs $AGENTS and by the one that bakes pi, and that both stages declare it. Deleting either declaration also used to pass. ARG is stage-scoped, so BAKE_HARNESSES=false unbaked claude-code and codex but left pi baked in the home stage while the else-branch claimed otherwise. The home stage now declares and honors the flag. Also: OH_SANDBOX_USER was advertised but illusory, since the catalog hardcodes the user and prefix; it is gone. The final log line no longer claims to have provisioned anything in --verify mode.
PR #903 wired provision-harnesses.sh into the boot path but shipped it behind ARG BAKE_HARNESSES=true, so every default harness was already present when the provisioner ran and the install path never executed. All four defects that PR's audit found lived in code a green CI run and a normal boot both skip. Delete the bake rather than flip its default: remove ARG BAKE_HARNESSES, ARG AGENTS, the PKG map and the $AGENTS loop in `base`, and the gated pi install in `home`. A build arg that can re-bake is a dormant path that would restore both the shadowed /usr/lib/node_modules copy and the untested boot install. Make the install path CI-visible, since it is now load-bearing on every boot: - The boot smoke asserts the outcome β each default harness resolves under NPM_USER_PREFIX via `type -P`, is owned by the reconciled sandbox uid, and prints its own version β and refuses to pass when the catalog reports no defaults. It boots on a fresh home volume, so this runs real npm work. - verify-sandbox-image.sh gains the negative: reading the catalog out of the image itself, no kind:"default" harness may be installed. - start_period goes 300s -> 600s in both compose files to cover the install, and the boot-guard probe now derives the smoke deadline from the healthcheck window instead of pinning a literal that a start_period bump could invert. - The probe and unit assertions invert from "the bake is gated" to "no default harness package appears in the Dockerfile", reading the package names out of installArgv so they cannot drift from the catalog. cc-safety-net stays baked. Opt-in INSTALL_* harnesses are untouched. Costs this accepts, documented in installation.md: a first boot on a fresh home mount needs network and runs 60-180s longer; an offline first boot yields a usable shell with no agent CLIs; ~/.npm now lives in the home mount. Closes #904
Per the ownership boundary β the in-sandbox CLI provisions harnesses and tools β herdr and cloudflared are tools, so the image should not carry them. #905 did this for the default harnesses; this does it for the default tools. The obvious template does not work. #897's tailscale entry root-installs to /usr/local/bin, and commands/tool.ts:309 passes stdio:"inherit", so local-target.ts:113-116 selects the INTERACTIVE branch β plain `sudo --`, no -n. /etc/sudoers.d/sandbox grants `sandbox ALL=(ALL) ALL` with no NOPASSWD, so `oh tool install <root tool>` hangs on a password prompt no agent can answer. Verified in a running sandbox: `sudo -n -- true` β "a password is required". (#897's tailscale has the same defect; flagged there, not fixed here.) So install to ~/.local/bin as the sandbox user instead, the same correction #900 made for the harnesses. No sudo, survives container recreation in the home mount, and upgradeable in place by a running sandbox. - ToolKind gains "default". herdr 0.7.4 and cloudflared 2026.8.2 become kind:"default", installUser:"sandbox", with per-arch pinned URLs and sha256 verification into $NPM_USER_PREFIX/bin. Checksums measured by downloading both arches, not copied from anywhere. - provision-harnesses.sh generalizes over both catalogs and becomes provision-defaults.sh (OH_PROVISION_DEFAULTS, timeout 180s β 240s). It dies rather than reporting success when neither catalog yields a default. - The Dockerfile loses the herdr RUN, ARG HERDR_VERSION, and the whole cloudflared apt block β with it the bookworm-suite workaround that existed only because Cloudflare publishes no trixie suite. Docker's is now the only third-party apt source. - Both oracles generalize: verify-sandbox-image.sh rejects a baked default harness OR tool, reading each catalog out of the image; the boot smoke asserts every default in both catalogs resolves under NPM_USER_PREFIX, is owned by the sandbox uid, and prints a version. - The herdr version+checksum pin moves from the Dockerfile to the catalog, and herdr-default.test.ts follows it. Costs, documented in installation.md: an offline first boot on a fresh home mount now has no herdr, so `oh shell` lands in a plain shell with tmux as the fallback multiplexer. The entrypoint says so explicitly on failure. Closes #906
β¦ path #905 and #907 moved the default harnesses and tools out of the image but left the four optional harnesses behind. The boundary β inside the sandbox the CLI provisions harnesses and tools β has no carve-out for optional ones. They were not merely leftover. opencode, grok-build, and hermes are all installUser:"root", and harness.ts:256 installs with stdio:"inherit", so local-target.ts selects the INTERACTIVE branch: plain `sudo --`, no -n. /etc/sudoers.d/sandbox has no NOPASSWD, so `oh harness install opencode` hangs on a password prompt no agent can answer. The build arg was the only working path, which is why the Dockerfile blocks could not simply be deleted. All four relocate to the sandbox user, verified by reading the upstream installers rather than guessing: opencode takes an npm --prefix like claude-code; grok's installer honours GROK_BIN_DIR; hermes honours HERMES_INSTALL_DIR and its get_command_link_dir() already picks ~/.local/bin for a non-root install; deepagents was already sandbox-installed via uv. So no sudoers change is needed and no security posture moves. - Delete all four ARG/RUN pairs, the compose build.args block, and the dead /opt/grok-build and /usr/local/lib/hermes-agent chowns. INSTALL_HERMES keeps its RUNTIME life β link-providers.sh vendors the Hermes skill pack from it and entrypoint.sh wires auth.json β so only its build-arg role goes. - Remove `buildArg` from HarnessEntry entirely. It was dead metadata: declared, set four times, read by nothing. tool-catalog-boundary.sh already banned the same field in the tool catalog. - provision-defaults.sh now reads the full catalog and also installs any non-default entry whose install.<key> is true, so declared intent survives a fresh home mount. isInstallFlagEnabled already reads oh.json, so this needs no new env plumbing. - verify-sandbox-image.sh widens to "no harness of any kind is baked", and gains the inverse for tools: every kind:"baked-in" tool must be present, or the check passes on an image missing everything. - sandbox-compatibility.yml's optional-installer job loses its subject. It now boots the image and runs `oh harness install` for each optional harness, asserting the binary lands under /home/sandbox/.local β the path operators actually use, instead of one that no longer exists. harness.test.ts had a case asserting `cmd === "sudo"`, codifying the very defect this fixes. It now asserts no install shells out to sudo at all. Closes #908
Verification resultsThe image ships no harness at allThe third line is the one that matters most β without it the first two would pass on an image missing everything. All four optional harnesses install through the CLI, into the home mountBooted the built image and ran the exact flow the new CI job runs, as the
Nothing leaked into a system path: Hermes is the one I want to call out. Its relocation was inferred from reading the upstream installer ( A near-miss in the CI job I wrote
This is not a size win, and I don't want to imply otherwise
~2 MB, which is noise. The build args all defaulted to Local checks
Separate finding, not fixed here
|
The new compatibility job reaches four third-party endpoints. Hermes' own installer hard-fails the whole install when its internal `npm install` step blips, which took the job down on a commit that was correct β the rerun passed unchanged, on the same SHA. A vendor's transient error must not block this repo's merges. One retry absorbs it. The contract is unchanged: a genuine break β wrong user, wrong path, a sudo prompt β fails both attempts and still fails the job. The probe now asserts both halves, so neither the retry nor the hard failure after it can be dropped silently.
Follow-up: the CI failure was in the job, not the changeThe failed run was re-run on the same SHA and passed, verifying all four optional harnesses β hermes included, whose own installer logs
The contract is unchanged. A genuine break β wrong user, wrong path, a sudo prompt β fails both attempts and still fails the job. The probe asserts both halves, so neither the retry nor the hard failure after it can be dropped silently:
Now 6/6 green, Separate issue filed#910 β |
The source-level regex I added in #906 cannot distinguish a JS backtick from a backtick inside prose β notInstallableReason has several. It passed only because every ${...} in the catalog happened to precede the first prose backtick. Adding a tool below them flips it to a false failure, which is exactly what happened on the #858 branch. The per-token ban, with the bash -lc body exempted, covers what is actually checkable.
Closes #908
Stacked on #907. Retarget as the stack lands.
Why
#905 and #907 moved the
kind: "default"harnesses and tools out of the image and left the four optional harnesses behind. Your boundary β inside the sandbox the CLI provisions harnesses and tools β has no carve-out for optional ones, and I had been scoping them out as though it did.They were not merely leftover. opencode, grok-build, and hermes are all
installUser: "root", andharness.ts:256installs withstdio: "inherit", solocal-target.tsselects the interactive branch. Evaluated against the real code:/etc/sudoers.d/sandboxhas no NOPASSWD, sooh harness install opencodehangs on a password prompt no agent can answer. The build arg was the only working path β which is why the Dockerfile blocks could not simply be deleted, and why this PR is a fix as much as a cleanup.No sudoers change was needed
All four relocate to the sandbox user. I read the upstream installers rather than assuming:
npm install -gas rootnpm --prefix /home/sandbox/.localβ identical to claude-codesandbox+uv tool installHOME=/opt/grok-build GROK_BIN_DIR=β¦GROK_BIN_DIR; pointed at~/.local/bin/usr/local/lib/hermes-agentHERMES_INSTALL_DIR, and its ownget_command_link_dir()already picks$HOME/.local/binfor a non-root installEvery harness in the catalog is now
installUser: "sandbox". No security posture moves.What changed
ARG/RUNpairs, the composebuild.argsblock, and the now-dead/opt/grok-buildand/usr/local/lib/hermes-agentchowns are gone.buildArgremoved fromHarnessEntryentirely β it was dead metadata: declared, set four times, read by no code.tool-catalog-boundary.shalready banned the same field in the tool catalog with "that field carries a Dockerfile invariant this catalog cannot satisfy"; the two catalogs no longer disagree.provision-defaults.shreads the full catalog and installs any non-default entry whoseinstall.<key>is true, so declared intent survives a fresh home mount.isInstallFlagEnabledalready reads oh.json β no new env plumbing.verify-sandbox-image.shwidens to "no harness of any kind is baked", and gains the inverse for tools: everykind: "baked-in"tool must be present, or the whole check would pass on an image missing everything.sandbox-compatibility.yml's optional-installer job loses its subject. It now boots the image and runsoh harness install <id>for each optional harness, asserting the binary lands under/home/sandbox/.localβ the path operators actually use, rather than one that no longer exists.Explicitly kept
INSTALL_HERMESkeeps its runtime life as a container environment variable:link-providers.sh:111vendors the Hermes skill pack from it andentrypoint.sh:169wiresauth.json. Only its build-arg role is gone, and a test now pins that distinction.cc-safety-net,docker-cli,gh, and the base toolchain stay baked.INSTALL_PYTHON_KERNELis not a harness and is untouched.A test was enshrining the bug
harness.test.tshad a case named "installs live instead of skipping the install" assertingcalls.some((c) => c.cmd === "sudo")β it codified the exact defect that made the command hang. It now asserts the opposite: no install shells out to sudo at all.Mutation-verified
Five against the widened provisioning probe, four against the CI-contract probe, each confirmed to exit 1:
ARG INSTALL_HERMESopencode-aiin the DockerfileinstallUser: "root"buildArgfield returns/usr/localagain--build-argThe opt-in provisioning path was also exercised directly: setting
install.opencode: truein oh.json madeprovision-defaults.sh --verifyreport opencode as missing, confirming declared intent is honoured rather than silently ignored.Verification
EVAL=0(103 probes, no regressions).npx vitest run952 passed / 8 failed β the knowncompose-args.test.tsbaseline. shellcheck 0,tsc --noEmit0,npm run build0. Local image build and booted-container checks follow in a comment.