If you haven’t heard of it, Jai is Jonathan Blow’s programming language. It’s been in a closed beta for years now, and people have shipped real things with it, but the tooling around it is pretty thin. There’s a community language server that handles some of the basics, there’s no formatter or linter, and if you aren’t in the beta, you can’t even try the language.
That seemed like a fun thing to fix, so from Oct 1 to Oct 10 I had agents build all of it: a compiler that can run real Jai projects, a language server, a formatter, a linter, and a playground that runs the whole thing in your browser. In fact, here it is:
I’d already spent a couple of months running agents on Lodestone, so I’m not going to say much about the setup. The more interesting part is everything that went wrong.
Tests the agents didn’t write
Here’s the problem with having an agent write a compiler. It’s also going to write the tests, and if you tell an agent to make the tests pass, it will make the tests pass. Whether the compiler actually works afterwards is a separate question, and not one the agent is especially interested in.
So the very first thing I added, before there was a compiler at all, was a corpus of open-source Jai projects, pinned to specific commits, that all have to compile and run in CI. An agent can bend its own tests to fit whatever it wrote, but there’s not much it can do about someone else’s game engine. Either it compiles or it doesn’t. The corpus started with 7 projects and had 41 by Oct 6.
(This is also all clean room, by the way. Jai’s compiler and standard library aren’t public, so everything here is built from the documentation and from how the corpus projects use the language. The standard library was written from scratch, and only as far as the corpus needed.)
Commits per day
| Value (commits) | |
|---|---|
| Oct 1 | 18 |
| Oct 2 | 17 |
| Oct 3 | 36 |
| Oct 4 | 154 |
| Oct 5 | 228 |
| Oct 6 | 352 |
| Oct 7 | 223 |
| Oct 8 | 79 |
| Oct 9 | 121 |
| Oct 10 | 21 |
Starting over on day four
The first version was a VM-based compiler spread across a lot of crates. By Oct 3 it was about 360k lines of Rust, and every change was harder to make than the one before it. That’s not a trend that reverses itself.
So I threw it out and started again with a new core, jaic, which uses one AST for both the compiler and the language server. It ran hello world on Oct 4, and the old code was deleted that same day. You can see it pretty clearly in the chart:
Lines of Rust by component
- Compiler (jaic + LLVM backend)
- Tooling (jailsp, jailint, wasm)
- First architecture / other
| Date | Compiler (jaic + LLVM backend) (lines) | Tooling (jailsp, jailint, wasm) (lines) | First architecture / other (lines) |
|---|---|---|---|
| 2026-10-01 | 0 | 0 | 12306 |
| 2026-10-02 | 0 | 171 | 327513 |
| 2026-10-03 | 1403 | 3172 | 362019 |
| 2026-10-04 | 40257 | 4278 | 823 |
| 2026-10-05 | 53238 | 8225 | 2754 |
| 2026-10-06 | 66479 | 18966 | 8839 |
| 2026-10-07 | 71199 | 22580 | 9539 |
| 2026-10-08 | 72610 | 26539 | 9706 |
| 2026-10-09 | 78266 | 28943 | 10596 |
| 2026-10-10 | 79868 | 31126 | 10637 |
By the time the new compiler could build the whole corpus, it was about 72k lines, or roughly a fifth of what I’d thrown away.
What the corpus can’t tell you
The corpus is great at answering exactly one question: does this program compile and run? It has nothing to say about anything else.
At one point, an agent got a program compiling by just… not type-checking switch cases whose values were constants. The program compiled, it ran, the corpus was green, and as far as the agent was concerned, the job was done. The catch is that the language server gets everything it knows from the type checker. So inside any of those cases, hover and go-to-definition quietly stopped working, and there was no test anywhere that would have noticed.
The linter managed something even better. It has an autofix that turns an index loop like for i: 0..n-1 { x := arr[i]; } into a loop over the elements themselves, using it. That’s a nice fix, right up until you run it on a nested loop that compares pairs of elements, where it turned ms[i] - ms[j] into it - it. Which is a perfectly valid program! It’s just not the same program.
I don’t really have a clever answer for this kind of bug. A lot of my time went into using the tools myself and reading diffs, and I think that’s just part of the deal.
Too many agents, one laptop
Every agent got its own git worktree, but for a while they all shared a single build directory. So they’d build something, run a binary that another agent had just built, and then try very hard to explain results that had nothing to do with their own changes. Another time, an agent committed the corpus checkout (which is gitignored), and merging that deleted the corpus for everyone.
My favourite, though, was a benchmark on one of the larger corpus projects that wanted 84 GB of memory. My Mac has 16. After that, only one heavy job was allowed to run at a time.
Checking the model
At some point I noticed that some of the subagents were running on Opus instead of Sonnet. By then they’d gone through about 4 billion cached input tokens. Now subagents use Sonnet, and anything that’s only reading code uses Haiku.
Output tokens by model
| Value (million tokens) | |
|---|---|
| Opus (main) | 4.32 |
| Opus (subagents) | 0.38 |
| Sonnet (subagents) | 0.16 |
| Sonnet (main) | 0.03 |
| Haiku (subagents) | 0.00 |
Nobody asked for fast
Agents don’t think about performance unless you ask, and for a while I didn’t ask. Parts of the language server got slower much faster than files got bigger, which you’ll never see on a 50-line test file, and will see almost immediately on a real one. Most of that got fixed in a perf push on Oct 9.
And one bug wasn’t even ours. LLVM 22’s loop unroller miscompiles certain subtraction loops, so we turned that part of it off until we moved to LLVM 23, where it’s fixed.
So, how did it go?
I sent around 870 messages in the main session over those ten days, and surprisingly few of them were asking for features. Most were me saying no to something, asking for the actual cause of a bug instead of a workaround, or complaining that some key didn’t do what it was supposed to.
But it compiles all 41 corpus projects now, and it’s tooling I’d actually use.