What it shows
A bet only counts as lost if its criteria were fixed before the result was known. Otherwise the criteria can drift until every bet looks like a win.
The record states the result this way:
The published record says, word for word (an excerpt)
Six pre-registered bets that lost are published with their pre-registered criteria intact, and the eval seeds are derived from the commit sha so they are untunable by construction.
In plain words: each of the six had its pass criteria written down first, and each failed them. The evaluation seeds come from the code version itself, so they cannot be picked after the fact to make a result look better. One example: certified training made its bounds 3.06 times tighter, clearing its bar, but made accuracy 2.98 times worse against a budget of 2 times, so nothing was promoted.
Why it matters
Published losses with their criteria are what make the published wins checkable. A buyer can see that the same discipline that produced the positives also rejected these.
What is ours, and what is not
Pre-registration is an established practice (see the prior art below). What is ours is these six bets, their criteria and their published results.
Who should care
- Reviewers and buyers judging how the positive results were selected.
- Teams training certified models, for the recorded accuracy cost.
The limits, in the record’s words
The published record says, word for word (an excerpt)
These are NEGATIVES. Their value is that the criteria were fixed first and the losses published, which is what makes the positives in this document legible.
The published record says, word for word (an excerpt)
DAD amortized experiment design is STATISTICALLY WORSE than exact greedy: Δ = −0.53 nats, 95% CI [−0.836, −0.210] excluding 0, winning 26.2% of 512 campaigns.
In plain words: these are failures, published as failures. The amortised experiment design won only 26.2% of 512 campaigns against the plain method. All of it is our own software and our own solver, in simulation.
Open source for this step
Tools and datasets we publish for the package step of building a multi-chip package. They are the checkers around this work, not a copy of the result itself.
- physics-lint: One command that checks a folder of physics models against a fixed set of named physical rules, with findings straight into CI.
- maxwell-lint: Flags a coupling extractor whose answers no passive set of conductors could produce.
- sparam-lint: Is your signal-response model physically possible? Five physical laws checked from the command line.
- interval-core: The interval arithmetic core behind our proofs over whole families of layouts.
- touchstone-tools: Read, write and convert Touchstone files, the standard text files that record how signals pass through a package's connections, and refuse to write one that cannot be read back.
- physics-lint-mcp: The physics checks, callable by an AI agent.
- physics-lint-action: A GitHub Action that fails the build when a model breaks one of a fixed set of named physical rules.
- Signal-response validity corpus: A labelled corpus of physically invalid signal-response networks, and a scorer that grades any checker against it.
- screening-ceiling: The screening-ceiling family as an open dataset.