> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mzizi.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# The Phase 0 pilots

> Two small pilots of the Phase 0 benchmark ran on 2026-09-27. Neither showed an advantage for Mzizi, and neither is the run that tests the kill criterion.

<Warning>
  **Neither pilot showed an advantage for Mzizi.** With a frontier model the two arms tied,
  and Mzizi used about 8% fewer tokens. With a \~7B open-weight model, the size the language
  is designed for, Mzizi did **worse** on all three metrics. Both pilots are tiny, and
  neither is the charter's measurement. The run that tests the kill criterion has not
  happened.
</Warning>

Every number on this page comes from the write-ups committed in
[`benchmarks/results/`](https://github.com/mzizi-dev/mzizi/tree/main/benchmarks/results) in
`mzizi-dev/mzizi`. The raw episodes are there too: every prompt, reply, candidate,
diagnostic and score. Read those, not this summary, if the detail matters to you.

## Why pilots, and why they were published

The charter makes Phase 0 a gate: Phase 1 does not start until a benchmark run finds a
measurable advantage. Phase 0 has one goal: to build Mzizi as a programming language, measured
against the best existing language for each kind of task, TypeScript, Python, Go, C++ and
Rust among them. The pilots were tests
inside Phase 0, not that goal: they compared Mzizi with Dioxus only, on a few registry UI
components. See [the benchmark](/benchmark#the-arms) for every arm.

A pilot tests the harness before it tests the thesis. It asks whether the runner, the arms
and the scorer produce numbers that mean something. Results are always published, whichever
way they fall: that is an owner decision, and it applied to both pilots. The Phase 1
decision stays the owner's after reading them.

## Pilot 1: frontier only

[`2026-09-27-pilot/RUN.md`](https://github.com/mzizi-dev/mzizi/blob/main/benchmarks/results/2026-09-27-pilot/RUN.md)

| Field | Value |
| - | - |
| Code | `ded425a`, the `main` branch at the time |
| Tasks | `button`, `badge`, `mzizi-changelog-renderer` (then `nyuchi-…`) |
| Arms | Mzizi and Dioxus |
| Author | One frontier model, one seed, five compile iterations |
| Tokens | Not measured: no tokenizer was running |

Every episode on both arms compiled clean on the first iteration. After a scoring fix
(below), both arms scored 0 defects. The only difference was class fidelity: Mzizi ports
kept fewer of the reference's Tailwind classes.

The write-up names the reasons this says nothing about the thesis:

* **n = 3** tasks, one seed, one model.
* **No small-model arm**, which RFC-0002 §5.4 says the frontier arm cannot replace.
* **No token data.**
* **The Mzizi checker at `ded425a` could not fail on names or types.** An undefined type
  compiled with 0 errors, so the Mzizi arm's first-time-clean result was measured against a
  checker that could not say no. Name and type resolution landed afterwards
  ([RFC-0008](https://github.com/mzizi-dev/mzizi/blob/main/design/RFC-0008-types-collections-records.md)).
* **The Mzizi changelog candidate solved a smaller problem.** Mzizi had no list type then,
  so it rendered one entry where the Dioxus candidate rendered the whole feed. The scorer
  only reads enums, so it could not see the difference.

The scoring fix: both arms first scored 2 identical "defects" on the changelog task. The
task's spec and its Rust reference name the same four variants differently, and both
authors followed the spec. The harness now pairs renamed variants by class string, but only
for a task that opts in with a stated reason.

## Pilot 2: the open-weight arm

[`2026-09-27-pilot-2/RUN.md`](https://github.com/mzizi-dev/mzizi/blob/main/benchmarks/results/2026-09-27-pilot-2/RUN.md)

The same code, `ded425a`, with a \~7B open-weight coding model running locally on CPU and
measured tokens. The frontier arm was re-run with three seeds per task. The write-up names
both models.

| Model | Arm | Clean compile | Iterations to clean | Tokens (mean) | Defect rate |
| - | - | - | - | - | - |
| Frontier | Mzizi | 6/6 | 1.17 | 4,172 | 0/6 |
| Frontier | Dioxus | 6/6 | 1.00 | 4,545 | 0/6 |
| \~7B open-weight | Mzizi | **2/6** | 1.50 (n=2) | 8,608 (n=5) | **2/2** |
| \~7B open-weight | Dioxus | **4/6** | 2.00 (n=4) | 7,782 | **0/4** |

Two scored tasks (`button` and `badge`, 11 facts) × 3 seeds per arm. One 7B Mzizi episode ran
out of context. The table counts it as not clean, because counting it the other way would
flatter Mzizi.

### What the write-up concludes

* **Frontier: the tasks were too easy to separate the arms.** Every episode compiled and
  none had a defect. The one real difference was tokens: Mzizi was 8.2% cheaper, because the
  ported file is smaller.
* **7B: the repair loop did not converge on Mzizi's diagnostics.** In 15 of 16 Mzizi repair
  attempts, the model resubmitted a byte-identical file. On Dioxus, 6 of 12. The charter bets
  that dense compiler errors make that loop faster. For the target model size, in this
  pilot, they did not.
* **React idioms with no Mzizi form.** Every badge episode stalled on the spec's `...props`
  spread and an `else` branch. Mzizi had neither, and the diagnostics did not say what to
  write instead.
* **The language's own variant table invited a defect.** Both clean 7B Mzizi buttons wrote a
  touch height that disagreed with their class. RFC-0006 names this failure: one fact
  written twice drifts.

### Threats to validity, in both directions

* **Tiny n.** One episode moves a rate by 17 points. Nothing here is statistically
  significant, and nothing is claimed to be.
* **A harness bug that penalised Mzizi.** `mz check` printed full absolute paths in every
  diagnostic, and the Dioxus check script shortened its paths. That cost the 7B Mzizi arm
  about 13,500 tokens, more than its whole token deficit.
* **Public tasks.** The tasks and results are public, so a later run on these tasks may be
  contaminated.
* **Shared authorship.** The RFCs and guides were written by the same model family as the
  frontier author, which may find Mzizi unusually legible for that reason.

## What has to change before the real run

Pilot 2 lists the fixes to make before a run that tests the kill criterion.
[`benchmarks/READINESS.md`](https://github.com/mzizi-dev/mzizi/blob/main/benchmarks/READINESS.md)
audits every one of them: each item was checked by running it on `main` at `a9c928d`, then
fixed where the fix was code. Those fixes are on `main` (checked at `6da2170`, 4 October
2026\). The tables below follow its audit. Read `READINESS.md` itself for the commands and file paths.

### Pilot 2's list, in its order

| # | Fix | State on `main` |
| - | - | - |
| 1 | Relative paths in the check output | **Fixed** in the runner, for every arm: the default `file-name` normaliser replaces the candidate's path with `candidate.mz` before the model sees it. Recorded in `meta.json` |
| 2 | Diagnostics for React idioms: the `...props` spread and `asChild` | **Fixed.** A spread is one `MZ0106` with an `exact` fix; `as_child` is the warning `MZ0312`. The 7B badge file that gave 13 errors now gives 4, and `mz fix` takes it to 1 |
| 3 | Derive `height` from `h-N` / `size-N`, or check both | **Fixed, both ways.** A row without `height` takes the height its class renders. A row whose `height` disagrees with its class is `MZ0313`, with the rendered number as the `exact` fix |
| 4 | A task set that can fail: more facts, more components, more seeds | **Partly.** A `slot_set` fact, a `card` task that uses it, and a driver that defaults to 5 seeds. The set that counts has to be held out: see 5 |
| 5 | Fresh, unpublished tasks | **Open; scaffolding done.** The plan and a task checker are in `benchmarks/kill-criterion/`. The private repository has not been created |

### Compiler behaviour that disagreed with the RFCs

| # | Divergence | State on `main` |
| - | - | - |
| 1 | `else` in a view rejected | **Fixed** by RFC-0008, before the audit |
| 2 | `mz fix` did not exist | **Fixed.** `mz fix` applies every `exact` fix and checks again. See [the compiler](/compiler#mz-fix-apply-every-exact-fix) |
| 3 | Unknown type accepted silently | **Fixed** before the audit: `MZ0701` |
| 4 | A missing `=` gave four errors, none on the wrong line | **Fixed.** One `MZ0406` on that line, with `=` as the fix |
| 5 | `if` in a view accepted silently | **Fixed.** `MZ0407`, with `when` as the `exact` fix. Anything after an element word on its line is `MZ0409` |
| 6 | `MZ0602` without its `exact` fix | **The RFC was wrong**, and RFC-0006 §8.2 is corrected: the fix needs an operand, and the clause had none |
| 7 | Closer fixes that replaced only `end`, and conflicting ones | **Fixed.** Every closer fix rewrites the whole closer, and open blocks at `end component` are one `MZ0204` |

### Threats to validity

| Threat | State on `main` |
| - | - |
| Tiny n | **Open.** It waits on the held-out set. The plan is at least 10 tasks and 30 facts per gating family, and 5 seeds on every arm |
| Path strings penalised Mzizi | **Fixed** (item 1) |
| Public tasks | **Open** (item 5) |
| The frontier model and the language share an author | **Open, mitigated in the plan.** The \~7B model is the headline arm, and the held-out tasks are not to be drafted by the same model family. The owner is to confirm who writes them |
| Frontier tokens are a proxy | **Open**, inherent to how the frontier arm runs |
| Seeds at temperature 0.2 repeated | **Planned:** temperature 0.7 with seeds 1–5, plus one seed at 0.2 as a check against the pilots. See below |
| A context overflow left an episode unfinished | **Fixed.** Such an episode now ends as `context_exceeded` and counts as not clean |
| The Leptos arm has never run | **Open.** The driver includes it by default |

`READINESS.md` also re-checks the pilots' candidates against the new compiler. Every Mzizi
candidate that compiled cleanly still does, except the two 7B buttons that `MZ0313` now
rejects, which is the defect pilot 2 found. Applying `mz fix` never raises a candidate's
error count.

## What remains before the run counts

**Phase 0 is not complete, and the kill-criterion run has not happened.** The fixes above make
the run possible; they are not a result. The gating task families are `ui-spec` and
`backend` ([RFC-0009](https://github.com/mzizi-dev/mzizi/blob/main/design/RFC-0009-comparison-benchmark.md)
§2), and Phase 0 passes only if both pass. What remains, in `READINESS.md`'s order:

1. **The runner and its input. Built, except the React pins.** Every arm is read from
   `arm.toml`, `ui-spec` hands the author a language-neutral `spec.md`, a `react` arm exists
   (a strict `tsc` check in an offline sandbox), and the four UI guides are rebalanced to
   within 5% of one token budget. Open: the React arm's pins are provisional until taken from
   the registry's lockfile. Neither the React arm nor the Leptos arm has run end to end.
2. **Held-out tasks** for each gating family, in a private repository that does not exist
   yet. Who writes them is open: maintainers by hand, or another vendor's model with a
   person reviewing each task, and not drafted by the model family that wrote the RFCs. The
   owner is to confirm.
3. **A pre-registered `PLAN.md`**, committed before the first episode and before any held-out
   score is seen. The driver refuses to start without one.
4. **A Leptos arm run**, end to end, at least once. Neither pilot ran it.
5. **The `backend` family.** The language work in RFC-0009 §6.4 is now mostly built
   ([RFC-0011](/rfcs)): `mz check` checks a `service`, `mz contract` runs it, and `mz build`
   lowers it to an axum package that serves locally. The `mzizi-be` arm exists, with the probe
   crate `mzprobe` and B1 as its first task, whose Mzizi and axum references both hold all 59
   of its facts. Still open: the runner does not score an episode with probes, tasks B2–B5 and
   the other-language backend arms do not exist, and nothing has been measured. Until they
   do, the `backend` family has not passed, and so neither has Phase 0.

### The settings the plan is to register

`benchmarks/kill-criterion/README.md` records the settings `PLAN.md` is to fix. They are
proposed, not registered: no `PLAN.md` has been written, and the settings are pending the
owner's confirmation.

* temperature 0.7 with seeds 1–5, plus one seed at 0.2 as a check against the pilots, which
  ran at 0.2;
* a 32,768-token context for the \~7B model's server, where pilot 2 ran at 16,384 and two 7B
  Mzizi episodes overflowed it.

### The kill criterion the run will test

The rule is RFC-0009 §6, an owner decision of 29 September 2026, and it is fixed before any
held-out score is seen:

* Within each gating task family, and for each of the three metrics (tokens, iterations to
  a clean check, defect rate), the bar is the **best** value any existing language in that
  family reached. The best on tokens and the best on defects can be different languages.
* Mzizi passes the family when it beats that bar on **at least two of the three** metrics, on
  the headline \~7B model, on held-out tasks.
* A win on a metric counts only if a paired bootstrap over tasks and seeds gives a **95%
  interval** for the difference that excludes zero in Mzizi's favour.
* **Everything is published**, whichever way it falls. A held-out run publishes its scores
  and a hash of its raw data on the day, and the task texts when the set is retired.

RFC-0009 §6.5 states the likely outcome before the run: on the only small-model data there
is, Mzizi did worse, so a loss on today's compiler is the expected result.

## How to talk about these results

"Designed for small models" is accurate. Calling Mzizi "faster", "cheaper" or "better" is not:
nothing has shown it. The honest summary is that Mzizi is a design with an argument behind
it, that the first two pilots did not show an advantage, and that the measurement that
decides the project has not run.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.