test/, all passingAbsolute, and the whole basis of the exercise:
What may be read is exactly what any other reader gets:
| Source | Why it is legitimate |
|---|---|
contract/help-json.json | the machine-readable command catalog — literally designed as the contract |
contract/guide.json | the embedded mental model, concepts, gotchas and examples |
contract/llms.txt | the public front door |
contract/spec-*.md | the agent-first CLI specs the tool is built to |
test/ and examples/ | the assertions and the userland scripts, both published |
| the live HTTP surface | what any caller can observe |
A second implementation earns its keep by finding things the first one cannot see about itself.
bkn's suites were written against a deployment whose fixtures had been
seeded by hand, so a fresh instance scores nonsense — t-files read
1/12 for want of two namespaces. dog.sh also takes its
admin token from $SP/admin.tok rather than the environment, so a mismatch
turns every write into a 403 and cascades into failures that look like missing
features. The seed is now written down.
Three of them were invisible until the fixtures were there, because
without seed data every suite failed for the wrong reason. Two were the same
shape and both presented as "the query found nothing": a
where clause that only understood the flat string form, and a
find() that dropped its criteria entirely and answered with whatever
happened to be first in the collection.
#672 —
sqlite_query returns [] for a query that failed,
indistinguishable from one that matched nothing; it silently broke three features in
a day. #673 — a goroutine
satisfies its own unbuffered channel send, so the keeper-lock idiom (the only way to
write a lock when every channel is unbuffered) spins at 100% CPU.
Six commands were fully implemented and simply absent from
help-json. Since that catalog is how an agent discovers the surface, a
working-but-unlisted command does not exist for any caller — and the contract suite is
blind to it, because it drives HTTP. A test now checks all three directions: everything
listed runs, everything the guide teaches is listed, everything listed is taught.
contract/ from a current bkn turned up four
commands added since the last sync — backup,
files sign, store access and
store count. Nothing was wrong with either build; the
snapshot had gone quietly out of date, and a score measured against
it read better than the truth. A second implementation is the thing that
notices, provided its copy of the contract is refreshed rather than trusted.
All four are now implemented, and files sign brought signed
links with it: a URL that is its own credential, so a private namespace can
be opened without handing out a token.
Needs the machin compiler and libssl-dev. QuickJS is vendored and
linked as a static archive, so the result is one file with no runtime dependencies.
There is a live one at
machin-bkn.vps1.intrane.fr,
so you can point bkn's own suite at it rather than take any of this on trust.
It holds the test fixtures and nothing else; every admin route is behind a
token, so the public surface is the guide, the health probe and the seeded
hooks.
./test/deploy-gate.sh --on-host runs all 191
against that deployment in one command, and answers with a number rather than
a wall of output:
Two parts, because t-scriptaccess starts its own server on its
own throwaway database: a script can only be created by the CLI on the machine
holding the data, so running it against the live instance would mean leaving a
publicly-runnable fixture script on a public host — and its last act closes a
script it opened, so the gate would rewrite the deployment every time. The
claim is therefore exact: the instance answers 172 over HTTPS and the
binary passes the other 19.
--on-host, and the mistake it came from.
Run from a laptop this gate flakes: t-access gave 31/31, 16/15,
28/3 and 30/1 on consecutive runs. The symptom is an empty response
body rather than an error status, and when the empty one lands on a
request that yields a token or a record id, every later assertion fails with
it — which is why the count swings between one and fifteen. It looks exactly
like a protocol bug, and I had a tidy explanation ready.
--on-host runs the HTTPS
half from the server: same TLS, same proxy, same routing, minus the flaky
hop.
To run it against your own build instead, seed the fixtures first or the score means nothing:
And t-scriptaccess.sh, which owns its own server for the reason
above — point it at the binary with BKN_BIN.
Built to the agent-first CLI spec family. This is a verification instrument rather than a distributed tool, so the three specs that exist to ship and support a product are deliberately out of scope.
| Spec | Status |
|---|---|
| cli-guide-spec | Yes — bkn guide [--human] (embedded, never fetched), GET /guide, GET /llms.txt |
| cli-output-spec | Yes — stdout = data, help-json, a version on every success, and typed errors whose code equals the exit code: 80 invalid/unknown, 85 validation, 92 not found, 95 already exists, 110 internal, each with type, recoverable and, where there is one, the command that fixes it |
| cli-daemon-spec | Yes — serve --host --port (loopback default, announced on stderr), GET /_health with a real pid, POST /_shutdown that answers before it exits and needs a token when bound off-loopback, and idempotent daemon start|stop|status |
| cli-update-spec | Out of scope — nothing here is distributed, so there is nothing to self-update |
| cli-feedback-spec | Out of scope — feedback on the contract belongs on bkn |
| cli-telemetry-spec | Out of scope — a test instrument counting its own runs measures nothing |