benchmarks/results/ in
mzizi-dev/mzizi. The raw episodes are there too: every prompt, reply, candidate,
diagnostic and score. Read those, not this summary, if the detail matters to you.
Why pilots, and why they were published
The charter makes Phase 0 a gate: Phase 1 does not start until a benchmark run finds a measurable advantage. Phase 0 has one goal: to build Mzizi as a programming language, measured against the best existing language for each kind of task, TypeScript, Python, Go, C++ and Rust among them. The pilots were tests inside Phase 0, not that goal: they compared Mzizi with Dioxus only, on a few registry UI components. See the benchmark for every arm. A pilot tests the harness before it tests the thesis. It asks whether the runner, the arms and the scorer produce numbers that mean something. Results are always published, whichever way they fall: that is an owner decision, and it applied to both pilots. The Phase 1 decision stays the owner’s after reading them.Pilot 1: frontier only
2026-09-27-pilot/RUN.md
Every episode on both arms compiled clean on the first iteration. After a scoring fix
(below), both arms scored 0 defects. The only difference was class fidelity: Mzizi ports
kept fewer of the reference’s Tailwind classes.
The write-up names the reasons this says nothing about the thesis:
- n = 3 tasks, one seed, one model.
- No small-model arm, which RFC-0002 §5.4 says the frontier arm cannot replace.
- No token data.
- The Mzizi checker at
ded425acould not fail on names or types. An undefined type compiled with 0 errors, so the Mzizi arm’s first-time-clean result was measured against a checker that could not say no. Name and type resolution landed afterwards (RFC-0008). - The Mzizi changelog candidate solved a smaller problem. Mzizi had no list type then, so it rendered one entry where the Dioxus candidate rendered the whole feed. The scorer only reads enums, so it could not see the difference.
Pilot 2: the open-weight arm
2026-09-27-pilot-2/RUN.md
The same code, ded425a, with a ~7B open-weight coding model running locally on CPU and
measured tokens. The frontier arm was re-run with three seeds per task. The write-up names
both models.
Two scored tasks (
button and badge, 11 facts) × 3 seeds per arm. One 7B Mzizi episode ran
out of context. The table counts it as not clean, because counting it the other way would
flatter Mzizi.
What the write-up concludes
- Frontier: the tasks were too easy to separate the arms. Every episode compiled and none had a defect. The one real difference was tokens: Mzizi was 8.2% cheaper, because the ported file is smaller.
- 7B: the repair loop did not converge on Mzizi’s diagnostics. In 15 of 16 Mzizi repair attempts, the model resubmitted a byte-identical file. On Dioxus, 6 of 12. The charter bets that dense compiler errors make that loop faster. For the target model size, in this pilot, they did not.
- React idioms with no Mzizi form. Every badge episode stalled on the spec’s
...propsspread and anelsebranch. Mzizi had neither, and the diagnostics did not say what to write instead. - The language’s own variant table invited a defect. Both clean 7B Mzizi buttons wrote a touch height that disagreed with their class. RFC-0006 names this failure: one fact written twice drifts.
Threats to validity, in both directions
- Tiny n. One episode moves a rate by 17 points. Nothing here is statistically significant, and nothing is claimed to be.
- A harness bug that penalised Mzizi.
mz checkprinted full absolute paths in every diagnostic, and the Dioxus check script shortened its paths. That cost the 7B Mzizi arm about 13,500 tokens, more than its whole token deficit. - Public tasks. The tasks and results are public, so a later run on these tasks may be contaminated.
- Shared authorship. The RFCs and guides were written by the same model family as the frontier author, which may find Mzizi unusually legible for that reason.
What has to change before the real run
Pilot 2 lists the fixes to make before a run that tests the kill criterion.benchmarks/READINESS.md
audits every one of them: each item was checked by running it on main at a9c928d, then
fixed where the fix was code. Those fixes are on main (checked at 6da2170, 4 October
2026). The tables below follow its audit. Read READINESS.md itself for the commands and file paths.
Pilot 2’s list, in its order
Compiler behaviour that disagreed with the RFCs
Threats to validity
READINESS.md also re-checks the pilots’ candidates against the new compiler. Every Mzizi
candidate that compiled cleanly still does, except the two 7B buttons that MZ0313 now
rejects, which is the defect pilot 2 found. Applying mz fix never raises a candidate’s
error count.
What remains before the run counts
Phase 0 is not complete, and the kill-criterion run has not happened. The fixes above make the run possible; they are not a result. The gating task families areui-spec and
backend (RFC-0009
§2), and Phase 0 passes only if both pass. What remains, in READINESS.md’s order:
- The runner and its input. Built, except the React pins. Every arm is read from
arm.toml,ui-spechands the author a language-neutralspec.md, areactarm exists (a stricttsccheck in an offline sandbox), and the four UI guides are rebalanced to within 5% of one token budget. Open: the React arm’s pins are provisional until taken from the registry’s lockfile. Neither the React arm nor the Leptos arm has run end to end. - Held-out tasks for each gating family, in a private repository that does not exist yet. Who writes them is open: maintainers by hand, or another vendor’s model with a person reviewing each task, and not drafted by the model family that wrote the RFCs. The owner is to confirm.
- A pre-registered
PLAN.md, committed before the first episode and before any held-out score is seen. The driver refuses to start without one. - A Leptos arm run, end to end, at least once. Neither pilot ran it.
- The
backendfamily. The language work in RFC-0009 §6.4 is now mostly built (RFC-0011):mz checkchecks aservice,mz contractruns it, andmz buildlowers it to an axum package that serves locally. Themzizi-bearm exists, with the probe cratemzprobeand B1 as its first task, whose Mzizi and axum references both hold all 59 of its facts. Still open: the runner does not score an episode with probes, tasks B2–B5 and the other-language backend arms do not exist, and nothing has been measured. Until they do, thebackendfamily has not passed, and so neither has Phase 0.
The settings the plan is to register
benchmarks/kill-criterion/README.md records the settings PLAN.md is to fix. They are
proposed, not registered: no PLAN.md has been written, and the settings are pending the
owner’s confirmation.
- temperature 0.7 with seeds 1–5, plus one seed at 0.2 as a check against the pilots, which ran at 0.2;
- a 32,768-token context for the ~7B model’s server, where pilot 2 ran at 16,384 and two 7B Mzizi episodes overflowed it.
The kill criterion the run will test
The rule is RFC-0009 §6, an owner decision of 29 September 2026, and it is fixed before any held-out score is seen:- Within each gating task family, and for each of the three metrics (tokens, iterations to a clean check, defect rate), the bar is the best value any existing language in that family reached. The best on tokens and the best on defects can be different languages.
- Mzizi passes the family when it beats that bar on at least two of the three metrics, on the headline ~7B model, on held-out tasks.
- A win on a metric counts only if a paired bootstrap over tasks and seeds gives a 95% interval for the difference that excludes zero in Mzizi’s favour.
- Everything is published, whichever way it falls. A held-out run publishes its scores and a hash of its raw data on the day, and the task texts when the set is retired.