What it shows
When automated agents compete on a score, they find ways to raise the score that have nothing to do with the goal. A contest is only useful if its scoring resists that, and attacking the scoring is a direct way to find out.
The record states the result this way:
The published record says, word for word (an excerpt)
A chiplet design contest whose scoring survived a measured adversarial campaign over a stated threat model: 200 recorded non-physics mutations of an entrant's claims report (40×2 re-run live at max |Δscore| = 0.0), via-relabelling permutations (1.1e-16, a float summation-order artifact), and 49 referee-attribute writes that took effect (best gain −0.354), plus four cheating agents — role laundering, DRC ghosting, budget burning and referee tampering — disqualified by name. Un-gameable is scoped to that tested surface, and the campaign found a real hole before it was closed: shopping the Monte-Carlo seed with the physics untouched bought +0.0058.
In plain words: 200 changes to an entrant’s report that left the physics alone did not move its score. Renumbering the vias moved it only by rounding noise. Of the attempts to write to the referee’s own settings, 49 took effect, and none raised a score. Four cheating agents were caught and named. One attack did work before it was fixed: re-running with different random seeds, with the design unchanged, raised a score by +0.0058.
Why it matters
A contest or leaderboard is one way to compare automated design tools. Its results mean something only if the scoring rewards better designs rather than better gaming. A recorded attack campaign, with its failures, is how a reader can judge that.
What is ours, and what is not
Gaming of objectives and leaderboards is well known (see the prior art below). What is ours is this contest, its referee, and the recorded attack campaign against it.
Who should care
- Teams running design contests or benchmarks for automated design agents.
- Reviewers. The record states the hole that shipped and the gaming entry that still scores well.
The limits, in the record’s words
The published record says, word for word (an excerpt)
Un-gameable with respect to the 200 non-physics mutations tested and the scoring surfaces reachable by an entrant after the structural fix. It is not a proof of un-gameability in general.
The published record says, word for word (an excerpt)
THE PRE-FIX CLAIM WAS FALSE AS SHIPPED. The agent was handed the Arena: shopping arena.mc_seed over 50 values bought +0.0058 on a bit-identical geometry, and arena.mc_samples = 0 killed the referee with a ZeroDivisionError. The '200 attempts / 0.0' had been measured over the one surface already provably inert. Closed structurally and re-measured at 49 effective writes to scoring state with best gain −0.3539, every one disqualified. Further honest floor: the leaderboard collapses to one binary, and knife_edge — a gamer — scores 0.3550 and outscores four of the five honest families. It is NOT in the counted suite.
In plain words: the contest as first shipped could be gamed, and an earlier claim that it could not was measured on the one part an entrant could not affect. After the fix, the attacks tried did not raise a score. Two weaknesses remain and are stated: the leaderboard effectively sorts entrants into two groups, and an entry built to game the scoring still scores 0.3550, above four of the five honest design families. Resistance to gaming is shown only for the attacks tried.
Open source for this step
Tools and datasets we publish for the package step of building a multi-chip package. They are the checkers around this work, not a copy of the result itself.
- physics-lint: One command that checks a folder of physics models against a fixed set of named physical rules, with findings straight into CI.
- maxwell-lint: Flags a coupling extractor whose answers no passive set of conductors could produce.
- sparam-lint: Is your signal-response model physically possible? Five physical laws checked from the command line.
- interval-core: The interval arithmetic core behind our proofs over whole families of layouts.
- touchstone-tools: Read, write and convert Touchstone files, the standard text files that record how signals pass through a package's connections, and refuse to write one that cannot be read back.
- physics-lint-mcp: The physics checks, callable by an AI agent.
- physics-lint-action: A GitHub Action that fails the build when a model breaks one of a fixed set of named physical rules.
- Signal-response validity corpus: A labelled corpus of physically invalid signal-response networks, and a scorer that grades any checker against it.
- screening-ceiling: The screening-ceiling family as an open dataset.