The goal
Phase 0 has one goal: to build Mzizi as a programming language, measured against the best existing language for each kind of task. The measure: an agent writes the same task in Mzizi and in each incumbent language, under the same benchmark harness, and the result is scored per task family, never pooled. This is RFC-0009, a draft whose kill criterion (§6) and publication rule (§7) are owner decisions. Parts of its runner work are built (below); nothing has been measured. The pilots measured a few registry UI components against Dioxus. That was a test inside Phase 0, not its goal.The task families
ui-spec exists because handing every arm a React component would make a React arm a copy
task. The backend tasks are grounded in the one real backend Mzizi operates, the API
gateway, and every fact is checked at the HTTP boundary, so no language’s syntax is scored.
The arms
What Mzizi is measured against, from RFC-0009 §1, with each arm’s state asbenchmarks/arms/ holds it at
language main 6da2170. The benchmark’s design, not a result: nothing has been measured
against TypeScript, Python, Go, C++ or a Rust backend.
Task B1’s reference implementations, one in Mzizi and one in plain axum, both hold all 59 of
its facts; a reference is not an arm. The public suites (MultiPL-E, EvalPlus, Aider polyglot,
BaxBench) are blocked: Mzizi has no functions yet (what still has to be built).
What it measures
Three metrics, for every arm in a family:
The stopping rule, as RFC-0001 §6 first wrote it:
If Mzizi doesn’t beat both on at least two of three metrics, the thesis is wrong and Phase 1 does not start.
Owner decision, 29 September 2026: the kill criterion is now Mzizi against the best
existing language for each kind of task, with the aim of ranking with the top languages.
It is RFC-0009
§6, and charter v0.4 §4 states it. Within each gating task family, Mzizi must beat the best
existing language’s value on at least two of the three metrics, on held-out tasks, with a
bootstrap 95% interval that excludes zero. It has not been measured. See
the pilot results for the rule in full.
The UI tasks, and what counts as a defect
Both resolved in the charter (§6, dated 2026-08-23), and widened by RFC-0009. The UI tasks come from Mzizi’s own components. The registry’s 577 components, a growing number of them with Rust siblings, are the ground truth for the UI families, with their existing.tsx and .rs implementations as references. Explicitly not a port of an
external library — not shadcn, not a generic primitive set. The nine
primitives and the two component examples in the repository are hand ports of
registry components. Supplying these tasks is one of the components’ jobs; they are built
to support the language as its UI layer, and they are not the benchmark’s goal. The
backend family’s tasks come from the API gateway instead.
A defect is code that checks cleanly but is behaviourally wrong. It passes its check and
fails a fact against the reference implementation: a UI fact, or an HTTP probe for a backend
task. This
mirrors the contract-test pattern already used for the .tsx → .rs ports: the reference
is read from disk, and disagreement is the new code’s fault unless it is a documented,
deliberate divergence.
Note what is not a defect: a syntax or compile error. That is the normal friction the
compile-error-density design goal is trying to minimise. The defect rate measures what gets
past the compiler wrong.
The benchmark harness that runs it
The public half of the benchmark is built, and lives inbenchmarks/ in mzizi-dev/mzizi:
CI runs the harness and runner tests with the rest of the workspace. The held-out task set
does not exist yet (see below), so every run so far has used public tasks.
Where the benchmark is going
Partly built. RFC-0009’s new arms (above) are the goal’s comparison: a
React arm beside Dioxus and Leptos in
ui-spec, and TypeScript, Python, Go, C++ and Rust in
backend. The React arm and the Mzizi backend arm are added, and neither has run; the
Hono, FastAPI, Go, C++ and axum arms are not added yet. Until they are, and until the runner
scores backend episodes with probes, the backend family cannot run.Results are always published
The owner’s rule: every run is published, whichever way it falls. Each lands inbenchmarks/results/
with its raw episodes, and RFC-0009 §7 adds a PLAN.md committed before a run’s first
episode, so a run that never reported is visible as exactly that. That includes the two pilots, which did not show an advantage. For
the held-out set, scores are published immediately; the task texts are published only when a
task set is retired, so that publishing them does not contaminate a set still in use.
Why one input is held out
This is RFC-0004, and it opens by correcting its own premise rather than building on it. The prompt for that RFC was SQLite — public-domain code, proprietary TH3 test harness, and the intuition that the split makes a project harder to attack. Two corrections: SQLite’s split is commercial, not defensive. TH3 is proprietary because it is sold and carries DO-178B avionics certification evidence with its own licensing constraints. SQLite simultaneously ships a very large public test suite. Private tests are close to worthless as a security control. An attacker has the source, a fuzzer and symbolic execution. A test suite mostly documents what a project already handles correctly — one of the least useful artifacts an attacker could be handed. Meanwhile hiding it costs real things: contributors cannot verify their own work, and coverage gaps become invisible to the people best placed to point them out. “Security that depends on the mechanism being secret is the failure mode Kerckhoffs named in 1883.” So no test in Mzizi is private because privacy makes it stronger. One reason is decisive:Benchmark contamination
If the task set and its expected outputs are public, they get scraped into training data. After that the benchmark measures memorisation, not the language, and reports a flattering number for exactly the wrong reason. This is not hypothetical — it is what happened to GSM8K, HumanEval and most public LLM benchmarks. It is also the one failure that cannot be detected from inside: a contaminated benchmark looks like a successful one. The charter gives Phase 0 a kill criterion, and a kill criterion that cannot fire is not a criterion. A held-out set is therefore a correctness requirement, not a secrecy preference. Three narrower reasons also qualify: fuzzing seed corpora (a curated set of inputs that once crashed the compiler is a head start against unpatched forks — but never the fix or the regression test, both of which stay public); embargo windows, which are temporary by design; and certification evidence, which is not applicable today and is named only so the boundary is already drawn.The split
Public — everything below
The language, compiler, primitives and RFCs. The entire correctness suite — unit
tests, contract tests, parse gates, measured IR properties — permanently. The
benchmark harness: runner, metric definitions, scoring code. A published fixture
format, so anyone can write their own task set and run it.
Private — a small annex
The held-out task set and its expected outputs. Fuzzing seed corpora. Embargoed security
tests, for the length of the embargo only.
The dependency rule
Private consumes public. Public never consumes private.Public CI must be completely self-contained: a fork with no secrets, no access and no relationship to the private repository must be able to run the full public suite and get a green result. The moment public CI needs a private token, every outside contributor’s CI fails and the project is open source in name only. Three consequences the RFC calls non-negotiable:
- A missing private result is
neutral, neverfailure. On a fork PR, a Dependabot PR, whenever the secret is absent, the check reports “skipped”. “A red X that an outside contributor is structurally unable to turn green is a wall, not a gate.” - The private result is advisory to outsiders, blocking only for maintainers. Branch
protection may require it on
main; it must not be required to propose a change. - A private failure must be reportable in public without leaking the test. The check reports the shape of the failure — which metric regressed, by how much, against which component — not the task input.
The mechanism, and its current state
The public half exists and is inert..github/workflows/mzizi-lang-benchmark-dispatch.yml POSTs a repository_dispatch to the
private repository carrying four public facts: commit SHA, ref, source repository, run id.
It is inert unless two things are configured — the MZIZI_HELDOUT_REPO repository
variable and the MZIZI_DISPATCH_TOKEN secret (contents:write on the private repository,
nothing else). Neither exists, so today the job runs and reports “not configured”. That is
the intended steady state, not a failure. It is additionally guarded on github.repository
so a fork skips it, and it never fails the build — a dispatch error is a warning, because
a missed notification does not mean the commit is bad.
The companion rule: the private runner must also poll for public commits it has not
measured. The dispatch is a latency optimisation; the poll is the correctness guarantee.
The private repository does not exist, deliberately
RFC-0004 settles where it will live —mzizi-dev, so read access to the held-out set is
governed by one org’s membership, which is the Mzizi maintainer set and nothing wider — and
then argues it should not be created yet, for a design reason rather than a scheduling one:
Its whole job is to run the public harness against a private input, so if it exists before the harness does, the harness ends up shaped around the private runner — the exact inversion §3 forbids.Plus an argument about incentives worth quoting, because it is the kind of thing usually left unsaid:
An empty private repository is an attractive nuisance. Today every test in this project is public, which is correct. The moment the repository exists, the marginal cost of filing a test there drops to zero; each individual “this one is easier to keep private” is defensible, and the aggregate is the public-shell-around-private-testing outcome §2 rules out.Trigger: the first held-out task. Not before.
What this deliberately does not claim
- It does not make Mzizi harder to attack. Security comes from the public suite, the
fuzzing, the
-D warningsgate and review — all public. - It does not hide the compiler’s behaviour. Every assertion about what Mzizi does is public; only a held-out measurement of how well an agent uses it is not.
- It is not permanent for security tests. Embargoed tests move to public at disclosure, and a test still private after its embargo has expired is a bug in the process.