Skip to main content
Two pilots have run, and neither showed an advantage for Mzizi. Both ran on 2026-09-27, on two or three small tasks. With a frontier model the arms tied; with a ~7B open-weight model Mzizi did worse on all three metrics. Neither is the run that tests the kill criterion, and that run has not happened. See the pilot results.

The goal

Phase 0 has one goal: to build Mzizi as a programming language, measured against the best existing language for each kind of task. The measure: an agent writes the same task in Mzizi and in each incumbent language, under the same benchmark harness, and the result is scored per task family, never pooled. This is RFC-0009, a draft whose kill criterion (§6) and publication rule (§7) are owner decisions. Parts of its runner work are built (below); nothing has been measured. The pilots measured a few registry UI components against Dioxus. That was a test inside Phase 0, not its goal.

The task families

ui-spec exists because handing every arm a React component would make a React arm a copy task. The backend tasks are grounded in the one real backend Mzizi operates, the API gateway, and every fact is checked at the HTTP boundary, so no language’s syntax is scored.

The arms

What Mzizi is measured against, from RFC-0009 §1, with each arm’s state as benchmarks/arms/ holds it at language main 6da2170. The benchmark’s design, not a result: nothing has been measured against TypeScript, Python, Go, C++ or a Rust backend. Task B1’s reference implementations, one in Mzizi and one in plain axum, both hold all 59 of its facts; a reference is not an arm. The public suites (MultiPL-E, EvalPlus, Aider polyglot, BaxBench) are blocked: Mzizi has no functions yet (what still has to be built).

What it measures

Three metrics, for every arm in a family: The stopping rule, as RFC-0001 §6 first wrote it:
If Mzizi doesn’t beat both on at least two of three metrics, the thesis is wrong and Phase 1 does not start.
Owner decision, 29 September 2026: the kill criterion is now Mzizi against the best existing language for each kind of task, with the aim of ranking with the top languages. It is RFC-0009 §6, and charter v0.4 §4 states it. Within each gating task family, Mzizi must beat the best existing language’s value on at least two of the three metrics, on held-out tasks, with a bootstrap 95% interval that excludes zero. It has not been measured. See the pilot results for the rule in full.
And from RFC-0002 §5, an addition that matters because the whole design target changed: the benchmark must include a small open-weight model arm. The frontier arm alone cannot validate the thesis, because the frontier model is the one the design helps least.

The UI tasks, and what counts as a defect

Both resolved in the charter (§6, dated 2026-08-23), and widened by RFC-0009. The UI tasks come from Mzizi’s own components. The registry’s 577 components, a growing number of them with Rust siblings, are the ground truth for the UI families, with their existing .tsx and .rs implementations as references. Explicitly not a port of an external library — not shadcn, not a generic primitive set. The nine primitives and the two component examples in the repository are hand ports of registry components. Supplying these tasks is one of the components’ jobs; they are built to support the language as its UI layer, and they are not the benchmark’s goal. The backend family’s tasks come from the API gateway instead. A defect is code that checks cleanly but is behaviourally wrong. It passes its check and fails a fact against the reference implementation: a UI fact, or an HTTP probe for a backend task. This mirrors the contract-test pattern already used for the .tsx → .rs ports: the reference is read from disk, and disagreement is the new code’s fault unless it is a documented, deliberate divergence. Note what is not a defect: a syntax or compile error. That is the normal friction the compile-error-density design goal is trying to minimise. The defect rate measures what gets past the compiler wrong.

The benchmark harness that runs it

The public half of the benchmark is built, and lives in benchmarks/ in mzizi-dev/mzizi: CI runs the harness and runner tests with the rest of the workspace. The held-out task set does not exist yet (see below), so every run so far has used public tasks.

Where the benchmark is going

Partly built. RFC-0009’s new arms (above) are the goal’s comparison: a React arm beside Dioxus and Leptos in ui-spec, and TypeScript, Python, Go, C++ and Rust in backend. The React arm and the Mzizi backend arm are added, and neither has run; the Hono, FastAPI, Go, C++ and axum arms are not added yet. Until they are, and until the runner scores backend episodes with probes, the backend family cannot run.

Results are always published

The owner’s rule: every run is published, whichever way it falls. Each lands in benchmarks/results/ with its raw episodes, and RFC-0009 §7 adds a PLAN.md committed before a run’s first episode, so a run that never reported is visible as exactly that. That includes the two pilots, which did not show an advantage. For the held-out set, scores are published immediately; the task texts are published only when a task set is retired, so that publishing them does not contaminate a set still in use.

Why one input is held out

This is RFC-0004, and it opens by correcting its own premise rather than building on it. The prompt for that RFC was SQLite — public-domain code, proprietary TH3 test harness, and the intuition that the split makes a project harder to attack. Two corrections: SQLite’s split is commercial, not defensive. TH3 is proprietary because it is sold and carries DO-178B avionics certification evidence with its own licensing constraints. SQLite simultaneously ships a very large public test suite. Private tests are close to worthless as a security control. An attacker has the source, a fuzzer and symbolic execution. A test suite mostly documents what a project already handles correctly — one of the least useful artifacts an attacker could be handed. Meanwhile hiding it costs real things: contributors cannot verify their own work, and coverage gaps become invisible to the people best placed to point them out. “Security that depends on the mechanism being secret is the failure mode Kerckhoffs named in 1883.” So no test in Mzizi is private because privacy makes it stronger. One reason is decisive:

Benchmark contamination

If the task set and its expected outputs are public, they get scraped into training data. After that the benchmark measures memorisation, not the language, and reports a flattering number for exactly the wrong reason. This is not hypothetical — it is what happened to GSM8K, HumanEval and most public LLM benchmarks. It is also the one failure that cannot be detected from inside: a contaminated benchmark looks like a successful one. The charter gives Phase 0 a kill criterion, and a kill criterion that cannot fire is not a criterion. A held-out set is therefore a correctness requirement, not a secrecy preference. Three narrower reasons also qualify: fuzzing seed corpora (a curated set of inputs that once crashed the compiler is a head start against unpatched forks — but never the fix or the regression test, both of which stay public); embargo windows, which are temporary by design; and certification evidence, which is not applicable today and is named only so the boundary is already drawn.

The split

Public — everything below

The language, compiler, primitives and RFCs. The entire correctness suite — unit tests, contract tests, parse gates, measured IR properties — permanently. The benchmark harness: runner, metric definitions, scoring code. A published fixture format, so anyone can write their own task set and run it.

Private — a small annex

The held-out task set and its expected outputs. Fuzzing seed corpora. Embargoed security tests, for the length of the embargo only.
RFC-0004 states the ratio plainly: “this is a small private annex to a large public suite, not a public shell around private testing. If someone cannot verify Mzizi’s correctness from the public repository alone, the split has been drawn wrongly.”

The dependency rule

Private consumes public. Public never consumes private.
Public CI must be completely self-contained: a fork with no secrets, no access and no relationship to the private repository must be able to run the full public suite and get a green result. The moment public CI needs a private token, every outside contributor’s CI fails and the project is open source in name only. Three consequences the RFC calls non-negotiable:
  1. A missing private result is neutral, never failure. On a fork PR, a Dependabot PR, whenever the secret is absent, the check reports “skipped”. “A red X that an outside contributor is structurally unable to turn green is a wall, not a gate.”
  2. The private result is advisory to outsiders, blocking only for maintainers. Branch protection may require it on main; it must not be required to propose a change.
  3. A private failure must be reportable in public without leaking the test. The check reports the shape of the failure — which metric regressed, by how much, against which component — not the task input.

The mechanism, and its current state

The public half exists and is inert. .github/workflows/mzizi-lang-benchmark-dispatch.yml POSTs a repository_dispatch to the private repository carrying four public facts: commit SHA, ref, source repository, run id. It is inert unless two things are configured — the MZIZI_HELDOUT_REPO repository variable and the MZIZI_DISPATCH_TOKEN secret (contents:write on the private repository, nothing else). Neither exists, so today the job runs and reports “not configured”. That is the intended steady state, not a failure. It is additionally guarded on github.repository so a fork skips it, and it never fails the build — a dispatch error is a warning, because a missed notification does not mean the commit is bad. The companion rule: the private runner must also poll for public commits it has not measured. The dispatch is a latency optimisation; the poll is the correctness guarantee.

The private repository does not exist, deliberately

RFC-0004 settles where it will live — mzizi-dev, so read access to the held-out set is governed by one org’s membership, which is the Mzizi maintainer set and nothing wider — and then argues it should not be created yet, for a design reason rather than a scheduling one:
Its whole job is to run the public harness against a private input, so if it exists before the harness does, the harness ends up shaped around the private runner — the exact inversion §3 forbids.
Plus an argument about incentives worth quoting, because it is the kind of thing usually left unsaid:
An empty private repository is an attractive nuisance. Today every test in this project is public, which is correct. The moment the repository exists, the marginal cost of filing a test there drops to zero; each individual “this one is easier to keep private” is defensible, and the aggregate is the public-shell-around-private-testing outcome §2 rules out.
Trigger: the first held-out task. Not before.

What this deliberately does not claim

  • It does not make Mzizi harder to attack. Security comes from the public suite, the fuzzing, the -D warnings gate and review — all public.
  • It does not hide the compiler’s behaviour. Every assertion about what Mzizi does is public; only a held-out measurement of how well an agent uses it is not.
  • It is not permanent for security tests. Embargoed tests move to public at disclosure, and a test still private after its embargo has expired is a bug in the process.

Still unsolved

Held-out set rotation. A held-out set leaks slowly through published results and needs a refresh policy — probably a fraction rotated per reported run. Third-party verification. If an outside party needs to reproduce a benchmark claim there has to be a path, most likely a time-limited grant under an agreement not to publish. “Worth solving before any number is published.”