Skip to content

Method

Six bets that lost, published with the criteria fixed before they ran

The result

Six approaches that failed criteria written down before they ran, published with the criteria and the numbers.

Limit Failures in our own software and simulation, published as failures.

A record that publishes only wins tells a reader little about how the wins were chosen. The record publishes six pre-registered bets that lost, each with the criteria fixed before it ran and the result against them. Among them: a certified-training attempt that tightened its bounds enough but lost too much accuracy, and an amortised experiment design that did worse than the plain method it was meant to replace.

A dotted magenta underline marks a number read straight from a published file when this page was built.

On this page
  1. What it shows
  2. Why it matters
  3. Who should care
  4. The limits, in the record’s words

What it shows

A bet only counts as lost if its criteria were fixed before the result was known. Otherwise the criteria can drift until every bet looks like a win.

The record states the result this way:

The published record says, word for word (an excerpt)

Six pre-registered bets that lost are published with their pre-registered criteria intact, and the eval seeds are derived from the commit sha so they are untunable by construction.

In plain words: each of the six had its pass criteria written down first, and each failed them. The evaluation seeds come from the code version itself, so they cannot be picked after the fact to make a result look better. One example: certified training made its bounds 3.06 times tighter, clearing its bar, but made accuracy 2.98 times worse against a budget of 2 times, so nothing was promoted.

Why it matters

Published losses with their criteria are what make the published wins checkable. A buyer can see that the same discipline that produced the positives also rejected these.

What is ours, and what is not

Pre-registration is an established practice (see the prior art below). What is ours is these six bets, their criteria and their published results.

Who should care

  • Reviewers and buyers judging how the positive results were selected.
  • Teams training certified models, for the recorded accuracy cost.

The limits, in the record’s words

The published record says, word for word (an excerpt)

These are NEGATIVES. Their value is that the criteria were fixed first and the losses published, which is what makes the positives in this document legible.

The published record says, word for word (an excerpt)

DAD amortized experiment design is STATISTICALLY WORSE than exact greedy: Δ = −0.53 nats, 95% CI [−0.836, −0.210] excluding 0, winning 26.2% of 512 campaigns.

In plain words: these are failures, published as failures. The amortised experiment design won only 26.2% of 512 campaigns against the plain method. All of it is our own software and our own solver, in simulation.

Open source for this step

Tools and datasets we publish for the package step of building a multi-chip package. They are the checkers around this work, not a copy of the result itself.

  • physics-lint: One command that checks a folder of physics models against a fixed set of named physical rules, with findings straight into CI.
  • maxwell-lint: Flags a coupling extractor whose answers no passive set of conductors could produce.
  • sparam-lint: Is your signal-response model physically possible? Five physical laws checked from the command line.
  • interval-core: The interval arithmetic core behind our proofs over whole families of layouts.
  • touchstone-tools: Read, write and convert Touchstone files, the standard text files that record how signals pass through a package's connections, and refuse to write one that cannot be read back.
  • physics-lint-mcp: The physics checks, callable by an AI agent.
  • physics-lint-action: A GitHub Action that fails the build when a model breaks one of a fixed set of named physical rules.
  • Signal-response validity corpus: A labelled corpus of physically invalid signal-response networks, and a scorer that grades any checker against it.
  • screening-ceiling: The screening-ceiling family as an open dataset.

Ask about a result, or check one yourself

Founder: Nick Harris. AI agents do our research and engineering. Each result page says how it was checked: against an outside solver, by an interval-arithmetic proof, by a Lean-checked step, or against our own simulator; these checks ran on our own machines. Who we are · How the work is checked

Every result on this site links to the file it comes from. Acquisition, licensing and partnership enquiries go to one address, nick@chipletos.com, and a person reads it.

Write to us Read the results

Each number links to the file it comes from; every file is listed, with its checksum, on Published files.

When a number is left off

We leave a number off a page, or mark it, when

  • its file has not loaded yet
  • nobody has looked into it yet
  • a search for it found nothing
  • its file holds no value for it
  • its file is missing or altered
  • files disagree on what it describes
  • its sample is too small for the claim
  • two files give different values
  • its file cannot be published
  • it was measured over ninety days ago
  • the question does not apply here
  • the program behind it stopped with an error