What it shows
A checker that has never been seen to fail proves little: it may accept anything. The usual answer is to run it on altered inputs and show that it rejects them.
The record states the result this way. In its words, "a buyer" means anyone given the kit, since we have no customer yet, and "verifies" means the kit file passes the checker’s tests, not that its models are right:
The published record says, word for word (an excerpt)
A buyer verifies the compiled PDK with zero dependencies, and can make the verifier generate 26 forgeries of the object and reject every one — a verifier that passes everything scores 0/26.
In plain words: the checker runs with nothing installed beyond Python itself. On request it writes 26 altered copies of the kit and must reject every one; a checker that accepted everything would reject 0 of them. The alterations were written by separate attempts to slip a forged kit past earlier versions of the checker, each of which had succeeded before the checker was fixed.
Why it matters
A buyer evaluating a design kit should not have to trust the seller’s software to check it. A checker with no dependencies can be read in full and run anywhere, and the forged kits give the buyer a way to see it fail on demand.
What is ours, and what is not
Design kits are standard, and testing a test by altering its inputs is a known practice (see the prior art below). What is ours is this compiled kit, the checker, and the set of forgeries it generates.
Who should care
- Buyers and evaluators who receive a compiled design kit and want to check it without installing our software.
- Reviewers. The record states how far the kit’s own accuracy claims are supported, below.
The limits, in the record’s words
The published record says, word for word (an excerpt)
Its 26 mutations were written by seven independent adversarial passes, each of which had already forged a PDK past the then-current gate. The compiler's anchor solver is 2-D quasi-static, measured median 18.2% / p95 22.6% against the Palace Driven full-wave reference below 110 GHz, while this PDK compiles at 10 GHz and advertises a tighter envelope. Realised holdout coverage is 0.8810 against a 0.95 target (40 seeds show the estimator unbiased at mean 0.9494 — a 10th-percentile draw), and the guarantee is marginal over the corpus, not on the narrower fab window (~0.906). Its gate verify_pdk_compiler_gate.py is NOT in the counted suite.
In plain words: passing the checker is a statement about the kit file, not about whether the kit’s models are right. The solver the kit is compiled from is a simplified two-dimensional one. Against Palace, an outside full-wave solver, below 110 GHz, its median error was 18.2%, and the kit, compiled at 10 GHz, states a tighter accuracy band than that. The kit’s error bands covered 0.8810 of held-out cases against a target of 0.95, and that coverage is an average over the whole data set, not a promise for a narrower manufacturing window.
Outside comparison
Compared with: Palace, the open-source finite-element electromagnetics solver from AWS Labs. Retrieved 2026-10-05. Palace, its public page · the retrieval record
Open source for this step
Tools and datasets we publish for the package step of building a multi-chip package. They are the checkers around this work, not a copy of the result itself.
- physics-lint: One command that checks a folder of physics models against a fixed set of named physical rules, with findings straight into CI.
- maxwell-lint: Flags a coupling extractor whose answers no passive set of conductors could produce.
- sparam-lint: Is your signal-response model physically possible? Five physical laws checked from the command line.
- interval-core: The interval arithmetic core behind our proofs over whole families of layouts.
- touchstone-tools: Read, write and convert Touchstone files, the standard text files that record how signals pass through a package's connections, and refuse to write one that cannot be read back.
- physics-lint-mcp: The physics checks, callable by an AI agent.
- physics-lint-action: A GitHub Action that fails the build when a model breaks one of a fixed set of named physical rules.
- Signal-response validity corpus: A labelled corpus of physically invalid signal-response networks, and a scorer that grades any checker against it.
- screening-ceiling: The screening-ceiling family as an open dataset.