Evolving Algorithms With My Claude Code Plan: ShinkaEvolve + Meridian + Zero API Spend
A few days ago I wrote a post about AutoPerf, a beam-search tool that calls an LLM in a loop to optimize C++ kernels. It worked, kind of — Gemini Flash found binary exponentiation, __builtin_ctzll, a branchless mask trick, and then plateaued after 20 iterations. Beam search is a fine primitive, but it is also the stone-age version of what the rest of the field is doing.
The current generation of open-source evolutionary code generators is more interesting. ShinkaEvolve (Sakana AI, ICLR 2026) is a population-based evolutionary loop with islands, archives, novelty scoring, and meta-summarization — according to their paper, the most sample-efficient of the current crop, ahead of OpenEvolve and CodeEvolve. I wanted to run it on a real problem.
Two things stopped me.
The Wrinkle
ShinkaEvolve, like basically every other LLM-driven evolutionary framework, calls the Anthropic Python SDK. The SDK expects ANTHROPIC_API_KEY. I do not have one. I have a Claude Code Max coding plan, which is a completely different auth mechanism: Claude Code uses OAuth under the hood, and the OAuth token is bound to the desktop client, not to the API console.
The two do not bridge. I could pay for an API key on top of my plan and burn real dollars on Haiku calls, or I could try to find a way to route SDK traffic through the existing Claude Code auth. The plan has generous rate limits for coding workloads; it seemed a shame to waste them.
Meridian
There is a community proxy called Meridian (npm install -g @rynfar/meridian) that does exactly this. It runs locally on 127.0.0.1:3456 and accepts traffic formatted for the Anthropic SDK. Under the hood, it translates those calls and forwards them via the Claude Code SDK, which already holds your OAuth token from the installed Claude Code client. Any tool that respects ANTHROPIC_BASE_URL and sends a dummy API key will transparently route through your Claude Code plan.
Cost to me: $0.00 in API spend, consumed against the existing plan's 5-hour rolling message quota. No Anthropic console signup. No separate billing.
The full chain looks like this:
ShinkaEvolve → Anthropic SDK → Meridian (localhost:3456) → Claude Code SDK → OAuth → Anthropic backend
ShinkaEvolve thinks it is talking to api.anthropic.com. Meridian re-frames each request as a Claude Code SDK call. The Claude Code SDK attaches the OAuth token and sends it onward. From ShinkaEvolve's perspective nothing is unusual — it gets completions back in the shape it expects.
Setup on my side was three env vars before launching the evolution:
export ANTHROPIC_BASE_URL=http://127.0.0.1:3456 export ANTHROPIC_API_KEY=dummy meridian & # background proxy
That is it. The framework does not know or care.
The Experiment
ShinkaEvolve ships with several example problems. I picked the circle packing benchmark because the scoring function is unambiguous and the SOTA is well-documented: place n=26 circles in a unit square, maximize the sum of radii, subject to no-overlap and boundary constraints. The best known result from the literature is 2.635. The seed program is a naive ring placement that scores 0.96. The gap between seed and SOTA is large enough that incremental improvement is visible, and the validator is simple enough that I can trust the numbers.
Configuration for the run:
| Parameter | Value |
|---|---|
| Generations | 30 |
| Islands | 1 |
| Archive size | 20 |
| Model (main) | Claude Haiku 4.5 |
| Model (novelty + meta) | Claude Haiku 4.5 |
| max_tokens | 4096 |
| Temperatures | [0.3, 0.7] |
| Patch types | 70% diff, 30% full rewrite |
| Embeddings | disabled (would need OpenAI) |
| Prompt evolution | disabled |
I chose Haiku over Sonnet for one reason: throughput. Haiku calls are roughly 2-3x faster end-to-end, which means more generations fit inside the plan's rolling rate limit. I will get to whether that was the right call.
Score Timeline
The raw per-generation scores are noisy, which is what you want from an evolutionary loop — the LLM is exploring, not monotonically improving. Here is the best-valid-so-far curve:
The climbing pattern, in round numbers:
| Generation | Best valid score | Notes |
|---|---|---|
| 0 | 0.9598 | Seed: ring placement |
| 1 | 1.7138 | First real jump |
| 7 | 1.7577 | Small refinement |
| 12 | 1.7870 | Small refinement |
| 17 | 1.8755 | parametric_boundary_positioning |
| 18-29 | 1.8755 | Plateau |
Four ladder steps in the first 17 generations, then flat for the remaining 13. The winning candidate — named parametric_boundary_positioning by ShinkaEvolve's meta-summarizer — was rediscovered four times (gens 17, 21, 23, 26). That is a clear local optimum: the LLM keeps proposing variants, the evaluator keeps returning the same score, the archive keeps the original.
Headline numbers:
- Best valid score: 1.8755
- +95% over the 0.96 seed
- 71% of the SOTA 2.635
- 14 of 30 programs (47%) passed geometric validation
"71% of SOTA" is the honest framing. This is not a result that beats the literature. It is a result that shows the pipeline runs end-to-end and makes steady progress from a bad starting point, using zero dollars of API budget.
The 99.36% Near-Miss
The most interesting moment of the run was at generation 22, and it is also the moment I am most annoyed about.
Generation 22 produced a candidate with a score of 2.6182 — that is 99.36% of the SOTA 2.635, inside the top 1% of published results. It would have been the new best valid score by a margin of 0.74 (a 40% improvement over parametric_boundary_positioning). If that had been a valid candidate, the run would have been essentially at SOTA in 22 generations.
It was rejected. The validator reported:
Circle 0 (x=0.3158, y=0.9059, r=0.0941) is outside unit square
0.9059 + 0.0941 = 1.0000. Exactly 1.0000 in floating-point arithmetic that is almost certainly off by a few ULPs in the wrong direction. The validator's check is a strict >: any boundary violation at all, no matter how small, disqualifies the entire candidate. There is no epsilon tolerance.
The patch to accept this candidate is literally:
if y + r > 1.0 + 1e-4: # was: if y + r > 1.0 return False, "out of bounds"
With that single line change, generation 22's candidate validates, the headline jumps from 71% to 99.36% of SOTA, and this article has a very different title. The validator strictness, not the LLM, is the bottleneck for the headline number. Haiku found a near-optimal packing; the scoring layer threw it out over a float rounding error.
I am keeping the original 1.8755 number in the tables because that is what the framework reported, and because rewriting history in the post would be dishonest. But the near-miss deserves to be called out.
The Three Solutions Side By Side
The numbers above are abstract. The geometry is not. Here is what the seed program, the best valid candidate, and the 99.36% near-miss actually look like in the unit square:
Left, gen 0 — the seed. A central circle, a ring of equal circles around it, and the rest distributed loosely around the edges. Sum of radii 0.96. There is obvious wasted space everywhere — the gaps between circles are large, and several boundary positions are unused. This is what the LLM started from.
Middle, gen 17 — the best valid candidate. A core layer, an interior ring, and 9 boundary circles wedged into the corners and edges. The packing is denser, the gaps are smaller, the space is used better. Sum of radii 1.88. This is a real improvement over the seed — the LLM clearly learned the structural pattern of "use the boundary." But there is still visible empty area, especially between the inner and outer layers.
Right, gen 22 — the rejected near-miss. A dense interlocked packing with 26 circles touching their neighbors and the boundary. Sum of radii 2.6182, which is 99.36% of the SOTA 2.635. The dashed red outline shows the offending circle: its center is at (0.3158, 0.9059) and its radius is 0.0941, so its top edge sits at y = 1.0000 — exactly on the boundary, but on the wrong side of the strict > check by a few floating-point ULPs. Visually you cannot see the violation. It is a single ε of arithmetic noise away from being valid.
The contrast between the middle and right panels is the whole story of this run. The LLM found the answer. The validator threw it out. The geometry is the same as a SOTA solution; the only difference is whether the framework's boundary check is willing to tolerate sub-ULP rounding error.
What Went Wrong In The Invalid Half
Of 30 candidate programs, 16 failed validation. The failures cluster into a few categories:
| Failure mode | Count | Typical cause |
|---|---|---|
| Circle overlaps | ~9 | Haiku pushes circles too close for density |
| Boundary violations | ~6 | Center + radius > 1 by a tiny margin |
| Syntax error | 1 | Gen 8: unexpected indent |
The overlap failures are more common than the boundary failures. This matches my intuition about what Haiku is doing: it is aggressively trying to increase packing density by nudging centers toward each other, and it miscalculates the pairwise distance by a little. Sometimes the miscalculation is a cosmetic fraction of a percent; sometimes it is a clear geometry error.
What is more interesting is the score distribution of the invalid programs. Several invalid candidates report scores above the best valid score — the 2.6182 candidate is one of them, but it is not the only one. Haiku is clearly finding configurations in the neighborhood of SOTA. What it struggles with is the arithmetic to make sure those configurations actually satisfy the constraints.
Put another way: Haiku knows roughly where the optimum is. It cannot reliably certify that its proposal lives there. That is a different failure mode than "the LLM has no idea what it is doing." A stricter prompt, a smarter validator with epsilon tolerance, or a larger model (Sonnet) would probably convert a good chunk of the invalid population into the valid population, and the headline would move accordingly.
What This Proves (And Does Not)
Proves:
- ShinkaEvolve + Meridian + Claude Code Plan works end-to-end. The plan auth routes through OAuth via Meridian, the framework is none the wiser, and no API key is needed. Zero dollars of external API spend across 30 generations.
- The evolutionary loop makes steady progress on a real problem with a real SOTA. +95% over the seed program, four distinct ladder steps, clear identification of a local optimum.
- The meta-summarizer names candidates, the archive tracks best-so-far, the novelty scorer adds exploration pressure. The whole thing behaves the way the paper describes.
Does not prove:
- LLM-driven evolution beats hand-tuned code on this problem. It reaches 71% of SOTA in 30 generations, not 100%.
- Haiku is the right model for this. The invalid population suggests it is finding good configurations but failing on precise geometry — exactly the kind of thing a bigger model would handle better.
- 30 generations is enough. The last 13 were a plateau; the run was still making progress when I stopped it.
Also proves (the more uncomfortable bit):
When the validator is strict and the LLM is sloppy, the dominant failure mode is "right answer, wrong epsilon" — and that is a framework problem more than an LLM problem. The 2.6182 candidate is the clearest evidence: a single line of validator code was the difference between 71% and 99.36% of SOTA. If you are running evolutionary code generation on physical or geometric problems, your validator is half the system, and "strict >" is almost certainly the wrong default for floating-point constraints.
Cost Breakdown
| Resource | Consumed |
|---|---|
| Wall-clock | ~50 minutes |
| Per generation | ~100 seconds (LLM call + Python eval) |
| Anthropic API spend | $0.00 |
| Claude Max messages | ~30-50 (well within 5h rolling limit) |
| Local CPU | ~2% of one core |
50 minutes and zero dollars is the cost number that matters. If I had run the same 30 generations against a metered API key, Haiku 4.5 pricing would have put this in the single-digit dollar range — not expensive, but not free. Meridian made it free, and more importantly, made it free without any extra billing plumbing. The exact same command that would have hit the metered endpoint hit the plan instead, because ANTHROPIC_BASE_URL was the only thing that changed.
This also means iteration is genuinely cheap. I can run 30 generations whenever I want to check a hypothesis. The cost per experiment is wall-clock time and a fraction of the plan's rolling quota — nothing I will run out of during a coding session.
What I Am Doing Next
The circle packing example is a good first run because the scoring function is unambiguous and the SOTA is a hard number from the literature. It is a bad second run because I do not own the problem — I have no stake in whether a circle-packing search beats the literature.
What I actually want is to point this at code I wrote. Specifically: the AoSoA container from the previous post. That post ended with a handful of hand-written optimizations: for_each_multistream<8>(f) for L3-bound workloads, per-field AVX2 accumulators for hot reductions, a lambda-driven for_each API to avoid the iterator branch. Each of those took me hours of reading assembly, benchmarking, and staring at perf counters. Some of them came out of a long conversation with subagents.
The question is simple: given the same seed code and the same benchmark, does Shinka rediscover those patterns? Does it find for_each_multistream<8>? Does it find the multi-accumulator reduction trick? Or does it find something I missed?
That is the experiment that actually answers "is evolutionary code generation useful for real performance work, or is it a demo toy that reaches 71% of SOTA on benchmarks whose answers are already in the training data." The circle packing run proved the plumbing. The AoSoA run will prove — or disprove — the value.
I will report back.
ShinkaEvolve: github.com/SakanaAI/ShinkaEvolve. Meridian: npmjs.com/package/@rynfar/meridian. The circle packing example and seed program live in ShinkaEvolve's examples/circle_packing directory.