> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mzizi.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# The Phase 0 benchmark

> Phase 0's one goal: to build Mzizi as a programming language, measured against the best existing language for each kind of task, TypeScript, Python, Go, C++ and Rust among them. The arms, the task families, the benchmark harness, why one input is held out, and the rule that keeps the project forkable. Two pilots have run; the kill-criterion run has not.

<Warning>
  **Two pilots have run, and neither showed an advantage for Mzizi.** Both ran on
  2026-09-27, on two or three small tasks. With a frontier model the arms tied; with a \~7B
  open-weight model Mzizi did worse on all three metrics. Neither is the run that tests the
  kill criterion, and that run has not happened. See [the pilot results](/pilots).
</Warning>

## The goal

**Phase 0 has one goal: to build Mzizi as a programming language, measured against the best
existing language for each kind of task.** The measure: an agent writes the same task in Mzizi
and in each incumbent language, under the same benchmark harness, and the result is scored
per task family, never pooled. This is
[RFC-0009](https://github.com/mzizi-dev/mzizi/blob/main/design/RFC-0009-comparison-benchmark.md),
a draft whose kill criterion (§6) and publication rule (§7) are owner decisions. Parts of its
runner work are built (below); nothing has been measured.

The pilots measured a few registry UI components against Dioxus. That was a test inside
Phase 0, not its goal.

## The task families

| Family | Input the agent sees | Role |
| - | - | - |
| `ui-spec` | `spec.md`, language-neutral | **Gating**, on held-out tasks |
| `backend` | `spec.md`, language-neutral; scored by HTTP probes | **Gating**, on held-out tasks |
| `ui-port` | `spec.tsx`, as both pilots used | Development and comparison with the pilots |
| Public suites | Each suite's own prompt | Comparability with the field; never gating |

`ui-spec` exists because handing every arm a React component would make a React arm a copy
task. The backend tasks are grounded in the one real backend Mzizi operates, the API
gateway, and every fact is checked at the HTTP boundary, so no language's syntax is scored.

## The arms

What Mzizi is measured against, from RFC-0009 §1, with each arm's state as
[`benchmarks/arms/`](https://github.com/mzizi-dev/mzizi/tree/main/benchmarks/arms) holds it at
language `main` `6da2170`. The benchmark's design, not a result: nothing has been measured
against TypeScript, Python, Go, C++ or a Rust backend.

| Arm | Family | Language and framework | State |
| - | - | - | - |
| `mzizi` | UI | Mzizi | Exists |
| `dioxus` | UI | Rust, Dioxus `=0.7.10` | Exists; both pilots ran it |
| `leptos` | UI | Rust, Leptos `=0.8.21` | Exists; never run |
| `react` | UI | TypeScript, React, `class-variance-authority` | Exists; never run. Its pins are provisional until taken from the registry's lockfile |
| `mzizi-be` | backend | Mzizi, a `service` (RFC-0011) | Exists, with the probe crate `mzprobe` and task B1; never run: the runner does not score with probes yet |
| `ts` | backend | TypeScript, Hono on Node | Not added yet |
| `python` | backend | Python, FastAPI with pydantic v2, served by uvicorn | Not added yet |
| `go` | backend | Go, `net/http` from the standard library | Not added yet |
| `cpp` | backend | C++20, `cpp-httplib` and `nlohmann/json` | Not added yet |
| `rust` | backend | Rust, axum on tokio, serde | Not added yet |

Task B1's reference implementations, one in Mzizi and one in plain axum, both hold all 59 of
its facts; a reference is not an arm. The public suites (MultiPL-E, EvalPlus, Aider polyglot,
BaxBench) are blocked: Mzizi has no functions yet ([what still has to be built](/tracker)).

## What it measures

Three metrics, for every arm in a family:

| Metric | Which failure modes it tests |
| - | - |
| **Tokens consumed** | FM-6 (ceremony) and FM-8 (formatting entropy) |
| **Iterations to a clean compile** | FM-2 (anonymous delimiters) and FM-5 (round-trip starvation) |
| **Defect rate** — compiles but wrong | FM-1, FM-3, FM-7, and contracts being in-language |

The stopping rule, as RFC-0001 §6 first wrote it:

> If Mzizi doesn't beat both on at least two of three metrics, the thesis is wrong and Phase
> 1 does not start.

<Note>
  **Owner decision, 29 September 2026: the kill criterion is now Mzizi against the best
  existing language for each kind of task**, with the aim of ranking with the top languages.
  It is [RFC-0009](https://github.com/mzizi-dev/mzizi/blob/main/design/RFC-0009-comparison-benchmark.md)
  §6, and charter v0.4 §4 states it. Within each gating task family, Mzizi must beat the best
  existing language's value on at least two of the three metrics, on held-out tasks, with a
  bootstrap 95% interval that excludes zero. It has not been measured. See
  [the pilot results](/pilots#the-kill-criterion-the-run-will-test) for the rule in full.
</Note>

And from RFC-0002 §5, an addition that matters because the whole design target changed:
**the benchmark must include a small open-weight model arm.** The frontier arm alone cannot
validate the thesis, because the frontier model is the one the design helps least.

## The UI tasks, and what counts as a defect

Both resolved in the charter (§6, dated 2026-08-23), and widened by RFC-0009.

**The UI tasks come from Mzizi's own components.** The registry's 577 components, a growing
number of them with Rust siblings, are the ground truth for the UI families, with their
existing `.tsx` and `.rs` implementations as references. Explicitly **not** a port of an
external library — not shadcn, not a generic primitive set. The nine
[primitives](/primitives) and the two component examples in the repository are hand ports of
registry components. Supplying these tasks is one of the components' jobs; they are built
to support the language as its UI layer, and they are not the benchmark's goal. The
`backend` family's tasks come from the API gateway instead.

**A defect is code that checks cleanly but is behaviourally wrong.** It passes its check and
fails a fact against the reference implementation: a UI fact, or an HTTP probe for a backend
task. This
mirrors the contract-test pattern already used for the `.tsx` → `.rs` ports: the reference
is read from disk, and disagreement is the new code's fault unless it is a documented,
deliberate divergence.

Note what is *not* a defect: a syntax or compile error. That is the normal friction the
compile-error-density design goal is trying to minimise. The defect rate measures what gets
**past** the compiler wrong.

## The benchmark harness that runs it

The public half of the benchmark is built, and lives in
[`benchmarks/`](https://github.com/mzizi-dev/mzizi/tree/main/benchmarks) in `mzizi-dev/mzizi`:

| Piece | What it does |
| - | - |
| `benchmarks/runner` | `mzbench`, the episode runner. It drives an author through compile iterations and records every one |
| `benchmarks/harness` | The scorer. It diffs a candidate's variants, defaults, touch heights and, per task, its `data-slot` set against the Rust reference |
| `benchmarks/arms` | Every arm as an `arm.toml`: `mzizi`, `dioxus`, `leptos`, `react` (a strict `tsc` check) and `mzizi-be` (`serve.sh` lowers and serves a candidate) |
| `benchmarks/probe` | `mzprobe`, which sends a backend task's HTTP probes and judges each fact |
| `benchmarks/prompts` | One guide per arm, built by the same prompt builder |
| `benchmarks/tasks` | The public tasks: `button`, `badge`, `card`, `mzizi-changelog-renderer`, and backend task `b1-routing` |
| `benchmarks/kill-criterion` | The driver for the real run, a held-out task checker, and the plan. Refuses to run without a `PLAN.md` |
| `benchmarks/openweight` | Setup and drivers for a local open-weight model |
| `benchmarks/results` | Every pilot run, raw episodes included |
| `benchmarks/READINESS.md` | The audit of pilot 2's fixes, and what the kill-criterion run still waits on |

CI runs the harness and runner tests with the rest of the workspace. The held-out task set
does not exist yet (see below), so every run so far has used public tasks.

## Where the benchmark is going

<Note>
  **Partly built.** RFC-0009's new arms ([above](#the-arms)) are the goal's comparison: a
  React arm beside Dioxus and Leptos in `ui-spec`, and TypeScript, Python, Go, C++ and Rust in
  `backend`. The React arm and the Mzizi backend arm are added, and neither has run; the
  Hono, FastAPI, Go, C++ and axum arms are not added yet. Until they are, and until the runner
  scores backend episodes with probes, the `backend` family cannot run.
</Note>

## Results are always published

The owner's rule: **every run is published, whichever way it falls.** Each lands in
[`benchmarks/results/`](https://github.com/mzizi-dev/mzizi/tree/main/benchmarks/results)
with its raw episodes, and RFC-0009 §7 adds a `PLAN.md` committed before a run's first
episode, so a run that never reported is visible as exactly that. That includes the two pilots, which did not show an advantage. For
the held-out set, scores are published immediately; the task texts are published only when a
task set is retired, so that publishing them does not contaminate a set still in use.

## Why one input is held out

This is [RFC-0004](/rfcs), and it opens by correcting its own premise rather than building
on it.

The prompt for that RFC was SQLite — public-domain code, proprietary TH3 test harness, and
the intuition that the split makes a project harder to attack. Two corrections:

**SQLite's split is commercial, not defensive.** TH3 is proprietary because it is sold and
carries DO-178B avionics certification evidence with its own licensing constraints. SQLite
simultaneously ships a very large *public* test suite.

**Private tests are close to worthless as a security control.** An attacker has the source, a
fuzzer and symbolic execution. A test suite mostly documents what a project already handles
correctly — one of the least useful artifacts an attacker could be handed. Meanwhile hiding
it costs real things: contributors cannot verify their own work, and coverage gaps become
invisible to the people best placed to point them out. *"Security that depends on the
mechanism being secret is the failure mode Kerckhoffs named in 1883."*

So no test in Mzizi is private because privacy makes it stronger. One reason is decisive:

### Benchmark contamination

If the task set and its expected outputs are public, they get scraped into training data.
After that the benchmark measures **memorisation**, not the language, and reports a
flattering number for exactly the wrong reason. This is not hypothetical — it is what
happened to GSM8K, HumanEval and most public LLM benchmarks.

It is also the one failure that cannot be detected from inside: a contaminated benchmark
looks like a successful one. The charter gives Phase 0 a kill criterion, and **a kill
criterion that cannot fire is not a criterion.** A held-out set is therefore a
correctness requirement, not a secrecy preference.

Three narrower reasons also qualify: fuzzing seed corpora (a curated set of inputs that once
crashed the compiler is a head start against unpatched forks — but never the fix or the
regression test, both of which stay public); embargo windows, which are temporary by design;
and certification evidence, which is not applicable today and is named only so the boundary
is already drawn.

## The split

<CardGroup cols={2}>
  <Card title="Public — everything below" icon="unlock">
    The language, compiler, primitives and RFCs. The **entire correctness suite** — unit
    tests, contract tests, parse gates, measured IR properties — permanently. The
    **benchmark harness**: runner, metric definitions, scoring code. A published **fixture
    format**, so anyone can write their own task set and run it.
  </Card>

  <Card title="Private — a small annex" icon="lock">
    The held-out task set and its expected outputs. Fuzzing seed corpora. Embargoed security
    tests, for the length of the embargo only.
  </Card>
</CardGroup>

RFC-0004 states the ratio plainly: *"this is a small private annex to a large public suite,
not a public shell around private testing. If someone cannot verify Mzizi's correctness from
the public repository alone, the split has been drawn wrongly."*

## The dependency rule

> **Private consumes public. Public never consumes private.**

Public CI must be completely self-contained: a fork with no secrets, no access and no
relationship to the private repository must be able to run the full public suite and get a
green result. The moment public CI needs a private token, every outside contributor's CI
fails and the project is open source in name only.

Three consequences the RFC calls non-negotiable:

1. **A missing private result is `neutral`, never `failure`.** On a fork PR, a Dependabot
   PR, whenever the secret is absent, the check reports "skipped". *"A red X that an outside
   contributor is structurally unable to turn green is a wall, not a gate."*
2. **The private result is advisory to outsiders, blocking only for maintainers.** Branch
   protection may require it on `main`; it must not be required to *propose* a change.
3. **A private failure must be reportable in public without leaking the test.** The check
   reports the *shape* of the failure — which metric regressed, by how much, against which
   component — not the task input.

## The mechanism, and its current state

The public half exists and is inert.
`.github/workflows/mzizi-lang-benchmark-dispatch.yml` POSTs a `repository_dispatch` to the
private repository carrying four public facts: commit SHA, ref, source repository, run id.

It is **inert unless two things are configured** — the `MZIZI_HELDOUT_REPO` repository
variable and the `MZIZI_DISPATCH_TOKEN` secret (`contents:write` on the private repository,
nothing else). Neither exists, so today the job runs and reports "not configured". That is
the intended steady state, not a failure. It is additionally guarded on `github.repository`
so a fork skips it, and it **never** fails the build — a dispatch error is a warning, because
a missed notification does not mean the commit is bad.

The companion rule: the private runner must also **poll** for public commits it has not
measured. The dispatch is a latency optimisation; the poll is the correctness guarantee.

### The private repository does not exist, deliberately

RFC-0004 settles where it will live — `mzizi-dev`, so read access to the held-out set is
governed by one org's membership, which is the Mzizi maintainer set and nothing wider — and
then argues it should not be created yet, for a design reason rather than a scheduling one:

> Its whole job is to run the *public* harness against a private input, so if it exists
> before the harness does, the harness ends up shaped around the private runner — the exact
> inversion §3 forbids.

Plus an argument about incentives worth quoting, because it is the kind of thing usually
left unsaid:

> An empty private repository is an attractive nuisance. Today every test in this project is
> public, which is correct. The moment the repository exists, the marginal cost of filing a
> test there drops to zero; each individual "this one is easier to keep private" is
> defensible, and the aggregate is the public-shell-around-private-testing outcome §2 rules
> out.

**Trigger: the first held-out task. Not before.**

## What this deliberately does not claim

* It does not make Mzizi harder to attack. Security comes from the public suite, the
  fuzzing, the `-D warnings` gate and review — all public.
* It does not hide the compiler's behaviour. Every assertion about what Mzizi *does* is
  public; only a held-out measurement of *how well an agent uses it* is not.
* It is not permanent for security tests. Embargoed tests move to public at disclosure, and
  a test still private after its embargo has expired is a bug in the process.

## Still unsolved

**Held-out set rotation.** A held-out set leaks slowly through published results and needs a
refresh policy — probably a fraction rotated per reported run.

**Third-party verification.** If an outside party needs to reproduce a benchmark claim there
has to be a path, most likely a time-limited grant under an agreement not to publish. *"Worth
solving before any number is published."*


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.