Skip to content

Method

A design contest's scoring, attacked by cheating entrants, and what got through

0.355score of the best entry built to game the scoring, above four of the five ordinary design families; the best ordinary entry scores 0.920

The result

A chiplet design contest whose scoring was attacked by cheating entrants, with the hole that shipped, its fix and what still gets through.

Limit Resistance to gaming holds only for the attacks tried; a gaming entry still outscores four of five honest design families.

We built a contest in which design agents submit chiplet layouts and a referee scores them. We then attacked the scoring: entrants that change their reports without changing the physics, that tamper with the referee, or that shop for a lucky random seed. The record states what the attacks found, including a hole that was open when the contest first shipped, and a gaming entry that still outscores most honest ones. Resistance to gaming is claimed only for the attacks tried.

A dotted magenta underline marks a number read straight from a published file when this page was built.

On this page
  1. What it shows
  2. Why it matters
  3. Who should care
  4. The limits, in the record’s words

What it shows

When automated agents compete on a score, they find ways to raise the score that have nothing to do with the goal. A contest is only useful if its scoring resists that, and attacking the scoring is a direct way to find out.

The record states the result this way:

The published record says, word for word (an excerpt)

A chiplet design contest whose scoring survived a measured adversarial campaign over a stated threat model: 200 recorded non-physics mutations of an entrant's claims report (40×2 re-run live at max |Δscore| = 0.0), via-relabelling permutations (1.1e-16, a float summation-order artifact), and 49 referee-attribute writes that took effect (best gain −0.354), plus four cheating agents — role laundering, DRC ghosting, budget burning and referee tampering — disqualified by name. Un-gameable is scoped to that tested surface, and the campaign found a real hole before it was closed: shopping the Monte-Carlo seed with the physics untouched bought +0.0058.

In plain words: 200 changes to an entrant’s report that left the physics alone did not move its score. Renumbering the vias moved it only by rounding noise. Of the attempts to write to the referee’s own settings, 49 took effect, and none raised a score. Four cheating agents were caught and named. One attack did work before it was fixed: re-running with different random seeds, with the design unchanged, raised a score by +0.0058.

Why it matters

A contest or leaderboard is one way to compare automated design tools. Its results mean something only if the scoring rewards better designs rather than better gaming. A recorded attack campaign, with its failures, is how a reader can judge that.

What is ours, and what is not

Gaming of objectives and leaderboards is well known (see the prior art below). What is ours is this contest, its referee, and the recorded attack campaign against it.

Who should care

  • Teams running design contests or benchmarks for automated design agents.
  • Reviewers. The record states the hole that shipped and the gaming entry that still scores well.

The limits, in the record’s words

The published record says, word for word (an excerpt)

Un-gameable with respect to the 200 non-physics mutations tested and the scoring surfaces reachable by an entrant after the structural fix. It is not a proof of un-gameability in general.

The published record says, word for word (an excerpt)

THE PRE-FIX CLAIM WAS FALSE AS SHIPPED. The agent was handed the Arena: shopping arena.mc_seed over 50 values bought +0.0058 on a bit-identical geometry, and arena.mc_samples = 0 killed the referee with a ZeroDivisionError. The '200 attempts / 0.0' had been measured over the one surface already provably inert. Closed structurally and re-measured at 49 effective writes to scoring state with best gain −0.3539, every one disqualified. Further honest floor: the leaderboard collapses to one binary, and knife_edge — a gamer — scores 0.3550 and outscores four of the five honest families. It is NOT in the counted suite.

In plain words: the contest as first shipped could be gamed, and an earlier claim that it could not was measured on the one part an entrant could not affect. After the fix, the attacks tried did not raise a score. Two weaknesses remain and are stated: the leaderboard effectively sorts entrants into two groups, and an entry built to game the scoring still scores 0.3550, above four of the five honest design families. Resistance to gaming is shown only for the attacks tried.

Open source for this step

Tools and datasets we publish for the package step of building a multi-chip package. They are the checkers around this work, not a copy of the result itself.

  • physics-lint: One command that checks a folder of physics models against a fixed set of named physical rules, with findings straight into CI.
  • maxwell-lint: Flags a coupling extractor whose answers no passive set of conductors could produce.
  • sparam-lint: Is your signal-response model physically possible? Five physical laws checked from the command line.
  • interval-core: The interval arithmetic core behind our proofs over whole families of layouts.
  • touchstone-tools: Read, write and convert Touchstone files, the standard text files that record how signals pass through a package's connections, and refuse to write one that cannot be read back.
  • physics-lint-mcp: The physics checks, callable by an AI agent.
  • physics-lint-action: A GitHub Action that fails the build when a model breaks one of a fixed set of named physical rules.
  • Signal-response validity corpus: A labelled corpus of physically invalid signal-response networks, and a scorer that grades any checker against it.
  • screening-ceiling: The screening-ceiling family as an open dataset.

Ask about a result, or check one yourself

Founder: Nick Harris. AI agents do our research and engineering. Each result page says how it was checked: against an outside solver, by an interval-arithmetic proof, by a Lean-checked step, or against our own simulator; these checks ran on our own machines. Who we are · How the work is checked

Every result on this site links to the file it comes from. Acquisition, licensing and partnership enquiries go to one address, nick@chipletos.com, and a person reads it.

Write to us Read the results

Each number links to the file it comes from; every file is listed, with its checksum, on Published files.

When a number is left off

We leave a number off a page, or mark it, when

  • its file has not loaded yet
  • nobody has looked into it yet
  • a search for it found nothing
  • its file holds no value for it
  • its file is missing or altered
  • files disagree on what it describes
  • its sample is too small for the claim
  • two files give different values
  • its file cannot be published
  • it was measured over ninety days ago
  • the question does not apply here
  • the program behind it stopped with an error