What it shows
A buyer’s first question about research software is whether it can be run at all, and what has been checked about running it.
The record states the result this way:
The published record says, word for word (an excerpt)
The estate is evaluable as a platform — API, MCP tool surface, buyer-verify script and a Dockerized appliance — and the three zero-dependency replay verifiers cited to buyers all pass; the wider buyer_verify.sh suite is a different object, carried as red and NOT re-run this round.
In plain words: the software can be reached through an API, a tool interface for AI agents, a verification script and a packaged appliance. The three small replay checkers a buyer is pointed to pass. The wider verification suite was failing when last run, and was not run again.
Why it matters
Buyers should know which parts of a delivery have been checked and which have not, before relying on any of it. The record says which is which, and calls this its weakest evidence.
What is ours, and what is not
APIs, agent tool interfaces and container appliances are standard ways to deliver software (see the prior art below). What is ours is this delivery surface and its stated state.
Who should care
- Technical evaluators who want to run the software themselves.
- Reviewers. The record names what has no check.
The limits, in the record’s words
The published record says, word for word (an excerpt)
The API asset has NO verify_*.py gate at all and the provenance linter has NO gate AND NO committed record. `make verify-buyer` covers ONLY the three stdlib replay verifiers, NOT the full `bash scripts/audit/buyer_verify.sh` suite, which is a different and still-red object.
The published record says, word for word (an excerpt)
is separately exactly 100 routes stale.
In plain words: the API has no check of its own, and its endpoint documentation, which the quoted line refers to, is out of date by 100 routes. The wider suite remains failing. Treat this as a description of what can be run today, not as evidence that it works end to end.
Open source for this step
Tools and datasets we publish for the package step of building a multi-chip package. They are the checkers around this work, not a copy of the result itself.
- physics-lint: One command that checks a folder of physics models against a fixed set of named physical rules, with findings straight into CI.
- maxwell-lint: Flags a coupling extractor whose answers no passive set of conductors could produce.
- sparam-lint: Is your signal-response model physically possible? Five physical laws checked from the command line.
- interval-core: The interval arithmetic core behind our proofs over whole families of layouts.
- touchstone-tools: Read, write and convert Touchstone files, the standard text files that record how signals pass through a package's connections, and refuse to write one that cannot be read back.
- physics-lint-mcp: The physics checks, callable by an AI agent.
- physics-lint-action: A GitHub Action that fails the build when a model breaks one of a fixed set of named physical rules.
- Signal-response validity corpus: A labelled corpus of physically invalid signal-response networks, and a scorer that grades any checker against it.
- screening-ceiling: The screening-ceiling family as an open dataset.