Back to archive
Field noteMay 20, 202616 min read

ZeroLang Field Notes: An Agent-First Language Meets Dijkstra

I cloned Vercel Labs Zero, built the native compiler, wrote Dijkstra in Zero and Python, and intentionally broke both. The result is promising for agent tooling, but the current executable subset still makes simple abstractions expensive.

ZeroLang Field Notes: An Agent-First Language Meets Dijkstra

Vercel Labs released Zero, an experimental programming language whose stated target user is not primarily a human typing in an editor, but an agent learning, checking, repairing, and shipping code through structured tooling.

That claim is interesting enough to test directly. Not with hello world. I wanted a small algorithm with real control flow, mutation, arrays, sentinel values, and enough surface area for mistakes to matter.

So I cloned the repo, built the compiler, wrote Dijkstra twice, then deliberately broke both versions.

Setup

I tested the repository at:

The first surprise was that the checked-out bin/zero wrapper was not immediately usable:

zero: native compiler is not built yet run: make -C native/zero-c

So the real setup was:

git clone --depth=1 https://github.com/vercel-labs/zerolang.git cd zerolang make -C native/zero-c bin/zero --version bin/zero run examples/add.0

The build succeeded, with one C warning about an ignored system() return value. The smoke example printed:

math works

So far, reasonable for a pre-1 compiler repo.

The Test Program

The task: shortest paths on a six-node weighted graph. In Python, this is boring in the best way:

import heapq N = 6 INF = 1_000_000 weights = [ [0, 7, 9, 0, 0, 14], [7, 0, 10, 15, 0, 0], [9, 10, 0, 11, 0, 2], [0, 15, 11, 0, 6, 0], [0, 0, 0, 6, 0, 9], [14, 0, 2, 0, 9, 0], ] def dijkstra(source: int) -> list[int]: dist = [INF] * N dist[source] = 0 queue = [(0, source)] while queue: cost, u = heapq.heappop(queue) if cost != dist[u]: continue for v, weight in enumerate(weights[u]): if weight == 0: continue candidate = cost + weight if candidate < dist[v]: dist[v] = candidate heapq.heappush(queue, (candidate, v)) return dist

It is compact because Python gives me lists, nested lists, heapq, dynamic allocation, and a rich standard library by default.

Zero pushed me in the opposite direction. The version that actually ran had to be flattened into main, use fixed arrays, avoid helper functions taking Span<i32>, avoid top-level constants, and represent visited as 0/1 integers instead of booleans:

pub fun main(world: World) -> Void raises { let weights: [36]i32 = [ 0, 7, 9, 0, 0, 14, 7, 0, 10, 15, 0, 0, 9, 10, 0, 11, 0, 2, 0, 15, 11, 0, 6, 0, 0, 0, 0, 6, 0, 9, 14, 0, 2, 0, 9, 0 ] let mut dist: [6]i32 = [1000000, 1000000, 1000000, 1000000, 1000000, 1000000] let mut used: [6]i32 = [0, 0, 0, 0, 0, 0] dist[0] = 0 let mut round: usize = 0 while round < 6 { let mut best: usize = 6 let mut bestDist: i32 = 1000000 let mut i: usize = 0 while i < 6 { if used[i] == 0 { if dist[i] < bestDist { best = i bestDist = dist[i] } } i = i + 1 } if best == 6 { round = 6 } else { used[best] = 1 let mut v: usize = 0 while v < 6 { let weight = weights[(best * 6) + v] if weight > 0 { if used[v] == 0 { let candidate = dist[best] + weight if candidate < dist[v] { dist[v] = candidate } } } v = v + 1 } round = round + 1 } } if dist[4] == 20 { if dist[5] == 11 { check world.out.write("dijkstra ok\n") } else { check world.out.write("dijkstra wrong\n") } } else { check world.out.write("dijkstra wrong\n") } }

It ran:

dijkstra ok

But getting there was the experiment.

Cost Of The Working Versions

The rough surface area:

VersionLinesBytesResult
Python38835runs
Zero, runnable subset591598runs
Zero, modular attempt781878checks, does not run

The Python version is shorter and more expressive because the heap is a one-line import. The Zero version is more explicit, more static, and more inspectable, but today that comes with a large ergonomic tax once you leave the compiler's direct-backend MVP subset.

This is not a criticism of the language idea. It is exactly what a pre-1 language looks like when the checker is ahead of the executable backend.

Clean Runtime Measurement

The first versions above are about ergonomics. For runtime, I used a separate benchmark so the measurement did not include compilation or zero run startup behavior.

Method:

  • build Zero once with zero build --emit exe --profile release-fast dijkstra_bench.0;
  • run the compiled executable directly;
  • run an equivalent flattened Python version, not the prettier heapq version, so both execute the same O(N²) fixed-graph algorithm;
  • repeat Dijkstra 200,000 times per process;
  • verify a checksum after every process run;
  • run 10 processes, discard the first as warmup, report the median of the remaining 9.

The benchmark programs were intentionally boring: no input parsing, no printing inside the loop, no file IO, no timing inside the measured program.

RuntimeMedian wall timePer DijkstraNotes
Zero compiled exe, release-fast0.042829 s0.214 µs3.4 KiB executable
Python 3 flattened loop0.835812 s4.179 µsinterpreter + Python loops

That is a 19.5x runtime advantage for the compiled Zero executable on this tiny benchmark.

The full run spread was also tight:

Zero release-fast: 0.042408–0.045452 s Python 3: 0.826273–0.858486 s

I would not overinterpret the absolute numbers. This graph has six nodes, everything is hot in cache, and the Python version is deliberately written in Python loops rather than NumPy. But the direction is real: once I forced the Zero program into the executable subset, it produced a tiny native binary and ran like native code. The price was paid earlier, in source shape and lost abstraction.

What Went Wrong In Zero

My first Zero version was the version I wanted to write: helper functions, a Span<i32> matrix parameter, mutable output spans, and a global INF.

zero check --json accepted one modular version:

{ "ok": true, "targetReadiness": { "ok": false, "languageOk": true, "buildable": false } }

That distinction is important. For an agent, this is useful structured information: the language accepted the program, but the target backend could not lower it.

Then zero run failed:

dijkstra.0:21:14 CGEN004: direct backend parameter type is unsupported expected: direct ELF64 object MVP subset actual: Span<i32> help: choose a supported direct target or restrict this program to exported primitive integer arithmetic functions

Earlier versions also hit:

CGEN004: direct backend MVP does not support declarations other than functions

That one was caused by a top-level constant. Again: a reasonable language feature, but outside the current executable subset.

The strangest failure came from a boolean visited array indexed with a variable. A tiny repro around used[i] produced:

{ "code": "PAR100", "message": "expected '}' after block", "line": 9, "column": 9, "repair": { "id": "repair-syntax", "summary": "Repair the syntax at the reported parser span, then rerun zero check." } }

The diagnostic is structured, which is good for machines. But the actual explanation is not yet high-level enough. The repair hint says "syntax", while the useful human action was "avoid this pattern; use an integer marker or simplify the expression."

What Went Wrong In Python

I broke Python in two boring ways:

heapq.heappushh(queue, (1, 1))

Python gave a traceback with a direct suggestion:

AttributeError: module 'heapq' has no attribute 'heappushh'. Did you mean: 'heappush'?

Then I forced a type error:

candidate = "7" if candidate < dist[0]: dist[0] = candidate

Python said:

TypeError: '<' not supported between instances of 'str' and 'int'

Python's diagnostics are not designed as a formal repair protocol, but for these mistakes they are excellent. They point to the line, show the expression, and name the runtime type mismatch.

Token-ish Cost Of Mistakes

I measured the byte size of the error payloads as a crude proxy for how much context an agent has to read before it can repair:

FailureOutput bytesMy repair effort
Zero parser issue on dynamic bool access514high: message was structured but misleading
Zero backend rejection for Span<i32> parameter283medium: message was clear, required redesign
Python heapq.heappushh typo257low: suggestion included
Python string/int comparison236low: traceback named both types

The interesting result is not that Python errors are shorter. The interesting result is that Zero's backend error was more useful to an agent than its parser error. CGEN004 told me the unsupported feature and the target subset. PAR100 gave me JSON, span, repair ID, and still sent me in the wrong direction.

Structured diagnostics are only valuable if the semantic level is right.

What Zero Does Well

The best part of Zero is not the syntax. It is the surrounding contract:

  • zero check --json gives machine-readable diagnostics.
  • targetReadiness separates "valid language" from "buildable for this target."
  • zero graph --json exposes capabilities, imports, symbols, and target support.
  • zero size --json exposes retained functions, runtime shims, stack bytes, rodata bytes, and helper retention.
  • The compiler ships version-matched skills via zero skills get zero --full.

That is genuinely agent-first. A normal language usually makes an agent scrape stderr, docs, and build logs. Zero is trying to make the compiler explain itself as structured data.

On the runnable Dijkstra, zero size --json reported:

  • one retained function: main
  • no heap allocation
  • world.out as the retained import
  • bounds-check and stdio runtime shims
  • 396 bytes of text
  • 29 bytes of rodata
  • 256 bytes of stack

That is a beautiful surface for an agent. It can reason about what got retained without guessing.

Where Python Still Wins

Python wins the experiment if the metric is "write the correct algorithm quickly."

It has:

  • heapq
  • dynamic lists
  • nested matrices
  • great tracebacks for common mistakes
  • no backend subset distinction

The cost is that Python is much less explicit. It does not tell me the memory floor, retained helpers, or target capability story. It lets me write the algorithm, but it does not give an agent a structured model of the program.

For human productivity, Python wins. For agent observability, Zero is more interesting.

Where Zero Could Win Later

Zero becomes compelling if the current split closes:

  1. The language subset accepted by check should be much closer to what run can execute.
  2. Parser diagnostics need semantic repair hints, not just structured spans.
  3. Common algorithmic patterns need examples: arrays, spans, mutable slices, helpers, small containers.
  4. The stdlib needs a few "agent obvious" data structures before agents stop reinventing flat arrays.

If those land, the payoff is real. An agent could:

  • ask the compiler for a graph of effects and capabilities;
  • get exact target buildability;
  • repair from structured diagnostics;
  • inspect binary size and retained helpers;
  • avoid dependency search for common tasks.

That is a different value proposition from Python. Python optimizes for humans and existing libraries. Zero is trying to optimize for agents and verifiable tooling loops.

Verdict

Zero is not ready for production code, and the README says as much. I would not use it today for a serious numerical kernel.

But as an experiment in agent-first tooling, it is worth watching. The language is currently less important than the protocol around it. check --json, graph --json, size --json, target readiness, repair IDs, and versioned skills are exactly the surfaces agents need.

My practical conclusion:

  • For "solve Dijkstra now", use Python.
  • For "teach an agent to inspect, repair, and reason about a program", Zero is the more interesting research object.
  • The current cost is paid in verbosity and backend limitations.
  • The future upside is lower repair ambiguity, if diagnostics become more semantic.

The most revealing moment was not the successful run. It was this split:

languageOk: true buildable: false

That is frustrating for a programmer. For an agent, it is gold.