Skip to content

Packaging

Proved checks on a neural stand-in, and the one that fails

The result

Proofs that a neural stand-in for our impedance solver stays in range, moves the physically right way and stays smooth, with the one check that fails stated.

Limit Properties of one secondary network’s shape, not proof its answers are correct; the range check misses most planted corruptions.

Our software can use a small neural network as a fast stand-in for a slower solver that computes the impedance of vertical wires through a glass carrier. We ran a set of mathematical checks on one such network: that its output stays inside a proved range, moves the physically right way when a dimension changes, never changes too sharply, can be worked backwards, and agrees across frequencies. One more check, agreement with the slower solver, turned out to be broken, and we say so. None of this shows the network’s answers are correct.

A dotted magenta underline marks a number read straight from a published file when this page was built.

On this page
  1. What it shows
  2. Why it matters
  3. Who should care
  4. The limits, in the record’s words

What it shows

A neural network trained to imitate a solver answers quickly, but nothing in its training stops it from answering in a way the physics forbids, for example predicting that a wire’s impedance falls when the physics says it must rise. Testing it on sample inputs cannot rule that out. A proof over a whole range of inputs can.

The record states the result this way:

The published record says, word for word (an excerpt)

The Gate-71 substrate-agnostic surrogate carries six certificate studies — output range, solver agreement, monotonicity, Lipschitz tolerance, inverse reachability and frequency consistency — of which FIVE are discriminative and are re-verified live by ONE command (6 gates green, output range on both shipped models). The sixth, solver agreement (Gate 150), is EXCLUDED from the battery rather than counted: it is anti-correlated with truth — a 1.5% output-layer degradation makes agreement 5× worse, 1.53% → 7.72%, while all three of its criteria still pass.

In plain words: five of the checks are proved properties of the network and are re-run together by one command. Each is proved over a range of inputs, not at sample points; the record states the range for the direction check (the whole range of via diameter and pitch it calls manufacturable), not for every check. The sixth, agreement with the slower solver, is left out of the count because it points the wrong way: when we damaged the network slightly, by 1.5%, agreement got worse by a factor of 5, from 1.53% to 7.72%, while the check still passed.

Why it matters

Teams that put a fast neural stand-in into a design check need more than an accuracy figure on a test set. They need to know the stand-in cannot do something physically impossible on inputs nobody tested. These checks are about that shape: range, direction, smoothness, invertibility and consistency across frequencies.

What is ours, and what is not

Proving bounds and slope directions for a trained network is a known field (see the prior art below). What is ours is this set of checks on a network we ship, the single command that re-runs them, and the record of the check that failed.

Who should care

  • Packaging software teams that put machine-learned stand-ins into design checks.
  • Reviewers of our results. The record names the broken check and the corrections behind it, rather than leaving them to be found.

The limits, in the record’s words

The published record says, word for word (an excerpt)

ONE CHECK IN THIS BATTERY IS BROKEN AND STAYS FAILING. Gate 150's ≤~58% sound solver-agreement bound, with 100% escalation and 0 wrong signoffs, is an honest result that was corrected twice: a /100 units bug produced an 'auto-decides 97.6%' figure, which was removed, and a claimed tight 'validated ≤1.56%' bound was removed after the signoff test found 34 or more wrong signoffs. But the check itself is anti-correlated with truth: a 1.5% output-layer degradation makes agreement 5× worse, 1.53% → 7.72%, while all three of its criteria pass. It is still live and is stated here rather than left to be found. Gate 152's bound is loose: elasticity about 2× the ideal-coax value of about 1. Gate 89's detectability is low: only 19 of 193 single-row ×10 corruptions are band-detectable, so the output-range check misses most deliberately corrupted models, and corruptions that preserve the network's function are undetectable by any sound output-range method. The model these checks cover is the Gate-71 substrate-agnostic surrogate, not the network the serving path uses by default. None of this shows that the network's answers are correct.

In plain words: these checks prove properties of the shape of one network, not that its answers match the solver, and not anything about a built package. The range check catches only 19 of 193 planted corruptions, and no range check of this kind can catch a corruption that leaves the network’s behaviour unchanged. The network checked is a secondary one, not the network our software uses by default.

Open source for this step

Tools and datasets we publish for the package step of building a multi-chip package. They are the checkers around this work, not a copy of the result itself.

  • physics-lint: One command that checks a folder of physics models against a fixed set of named physical rules, with findings straight into CI.
  • maxwell-lint: Flags a coupling extractor whose answers no passive set of conductors could produce.
  • sparam-lint: Is your signal-response model physically possible? Five physical laws checked from the command line.
  • interval-core: The interval arithmetic core behind our proofs over whole families of layouts.
  • touchstone-tools: Read, write and convert Touchstone files, the standard text files that record how signals pass through a package's connections, and refuse to write one that cannot be read back.
  • physics-lint-mcp: The physics checks, callable by an AI agent.
  • physics-lint-action: A GitHub Action that fails the build when a model breaks one of a fixed set of named physical rules.
  • Signal-response validity corpus: A labelled corpus of physically invalid signal-response networks, and a scorer that grades any checker against it.
  • screening-ceiling: The screening-ceiling family as an open dataset.

Ask about a result, or check one yourself

Founder: Nick Harris. AI agents do our research and engineering. Each result page says how it was checked: against an outside solver, by an interval-arithmetic proof, by a Lean-checked step, or against our own simulator; these checks ran on our own machines. Who we are · How the work is checked

Every result on this site links to the file it comes from. Acquisition, licensing and partnership enquiries go to one address, nick@chipletos.com, and a person reads it.

Write to us Read the results

Each number links to the file it comes from; every file is listed, with its checksum, on Published files.

When a number is left off

We leave a number off a page, or mark it, when

  • its file has not loaded yet
  • nobody has looked into it yet
  • a search for it found nothing
  • its file holds no value for it
  • its file is missing or altered
  • files disagree on what it describes
  • its sample is too small for the claim
  • two files give different values
  • its file cannot be published
  • it was measured over ninety days ago
  • the question does not apply here
  • the program behind it stopped with an error