What it shows
A fast stand-in that is right most of the time is still dangerous if nothing says when it is not. Conformal methods give a bound, from a held-out sample, on how often an accepted answer is wrong, and let the system decline the rest.
The record states the result this way:
The published record says, word for word (an excerpt)
Surrogate predictions are wrapped in distribution-free error control and uncertain cases escalate, rather than silently returning an overconfident fast answer.
In plain words: the stand-in’s answers are accepted only when the error control admits them, and the rest go to the full solver. In the record’s own sample, accepted answers were wrong 0.958% of the time, against a target of 1%; the bound that is certified is the finite-sample one, and its interval reaches above the target.
Why it matters
A design check that uses a fast stand-in needs a stated rate at which the stand-in’s answers can be wrong, and a rule for what happens otherwise. Without one, a fast answer looks the same whether or not it can be trusted.
What is ours, and what is not
Conformal prediction and risk control are published methods (see the prior art below). What is ours is their use around our stand-in, the routing of unsure cases to the full solver, and the recorded weaknesses.
Who should care
- Teams that put learned stand-ins into design checks.
- Reviewers. The record states each place where the bounds are weaker than they first look.
The limits, in the record’s words
The published record says, word for word (an excerpt)
Gate 114's certification is the finite-sample BOUND, not the point estimate — the confidence interval straddles the target rate. Gate 70 holds under the STATED deployment prior, pre-registered JEDEC HBM4 / UCIe datasheet nominals, not tuned to the residual.
The published record says, word for word (an excerpt)
Gate 114 certifies E[false-pass] ≤ 1% (split-CRC) and P(rate > 1%) ≤ 5% (RCPS) with observed 0.958% (23/2,400) and 90% CP [0.66, 1.36]%. Gate 70 found the uniform-corpus quantile UNDER-covers at 86.0% against 95% nominal in EXTREME_TIGHT; weighting restores 95.9% and the deployed rule is the coverage-safe max of the two. Gate 115 repairs a real correctness gap: the active-learning flywheel selects via the model, breaking exchangeability, so plain coverage 0.883 becomes 0.965 under Fannjiang-2019. Gate 123's field-disjoint coverage 0.0529 → 0.9102 was bought largely with WIDTH — mean half-width 1.16 → 40.2 dB, 83.8% of grid score mass in >20 dB cells. It converts a FALSE 90% guarantee into an HONEST one; it does not make the operator accurate on grids.
In plain words: the guarantee is about how often accepted answers are wrong, under the stated assumptions about the inputs, not about how accurate the stand-in is. In one setting the bands covered only 86.0% of cases against a stated 95% until they were re-weighted, and on new kinds of layout the bands reach their coverage mostly by becoming very wide. All of this is measured against our own solver, in simulation.
Open source for this step
Tools and datasets we publish for the package step of building a multi-chip package. They are the checkers around this work, not a copy of the result itself.
- physics-lint: One command that checks a folder of physics models against a fixed set of named physical rules, with findings straight into CI.
- maxwell-lint: Flags a coupling extractor whose answers no passive set of conductors could produce.
- sparam-lint: Is your signal-response model physically possible? Five physical laws checked from the command line.
- interval-core: The interval arithmetic core behind our proofs over whole families of layouts.
- touchstone-tools: Read, write and convert Touchstone files, the standard text files that record how signals pass through a package's connections, and refuse to write one that cannot be read back.
- physics-lint-mcp: The physics checks, callable by an AI agent.
- physics-lint-action: A GitHub Action that fails the build when a model breaks one of a fixed set of named physical rules.
- Signal-response validity corpus: A labelled corpus of physically invalid signal-response networks, and a scorer that grades any checker against it.
- screening-ceiling: The screening-ceiling family as an open dataset.