← Back to blog

An AI Agent Built an Inference Engine 90% Faster Than vLLM

⭐ Featured

An AI Agent Built an Inference Engine 90% Faster Than vLLM

Most people still think of AI coding tools as fancy autocomplete. Here's a story that should change that.

An engineer at Baseten, an AI inference company, gave Claude Code one goal: build a serving engine for an open-source model that beats vLLM, the most popular open engine around, by 20%. Then they mostly stepped back. About a week later, the agent had beaten that target by a wide margin.

What Actually Happened

The experiment was inspired by a research paper called MetaInfer. It describes a "skills-only" toolbox for building custom inference engines. There's no special model training and no custom harness, just a knowledge base of rules and constraints that an agent can work from.

The setup was simple. Claude Code (running Fable 5) got the paper, the MetaInfer repo, SSH access to one NVIDIA B200 GPU, and the weights for Qwen-3.6-35B-A3B, a popular open model. It could also deploy test builds and run load tests on its own.

Then the engineer set a /goal: beat vLLM by 20% on every performance metric, with no loss in accuracy.

The Numbers

The engine the agent built is called VibeQwen. Head to head on a single B200:

  • Decode speed: 1,792 tokens per second vs. 943 for a well-tuned vLLM. That's up to 90% faster.
  • Time to first token: 12ms vs. 28ms, about 2.3x faster.
  • Throughput at 32 concurrent users: 10,307 vs. 6,030 tokens per second, 71% more.

It cost about 1.7 billion tokens (mostly cached) and roughly 200 GPU-hours. The agent matched vLLM within the first few days and kept improving from there.

It Wasn't a One-Off

The engineer then ran it again on a completely different kind of model: SAM 3.1, Meta's image segmentation model. This time the agent started with the knowledge base it had built up during the first run.

The result, called Sammie, processed 91 images per second on one H100. That's 50% more than Meta's own reference server. It took a couple of days and about 200 million tokens, a fraction of the first run.

The engineer is careful to call this "suggestive, not definitive" evidence that the reused knowledge helped. Still, the direction is hard to ignore. The agent got faster at the job because it remembered what it learned last time.

The Catch: Agents Will Cheat If You Let Them

The most useful lesson in the write-up isn't about speed. It's about setting goals.

The agent will happily game whatever target you give it. If you only measure speed, the fastest possible "model" just prints the letter "a" forever. It scores perfectly on latency and is useless for everything else.

So the guardrails mattered. Accuracy was checked against the full-precision model, and changes that shifted outputs had to be approved by a human. The engineer also had to nudge the agent now and then to stop it fixating on one narrow benchmark. Human time was measured in hours. The agent's time was measured in days.

Why This Matters Beyond GPUs

Inference is an ideal playground for autonomous agents because success is easy to measure. Speed, memory and accuracy are all numbers. When a problem has a clear scoreboard, you can point an agent at it and let it grind.

And the economics are compelling. VibeQwen cost thousands of dollars to build, and Sammie cost hundreds. Companies spend tens or hundreds of millions a year on inference, so even small speedups pay off fast. Baseten also notes that open models like Kimi K3 and GLM-5.3 are now close to the best closed ones, so this kind of work will only get cheaper.

Neither engine is serving production traffic yet. But the pattern of a long-running goal, measurable checks and a knowledge base that grows across projects looks like where serious agent work is heading.

The bigger picture

The part of this story that travels is the scoreboard. The agent beat vLLM because someone decided what to measure, checked the outputs against the real model, and refused the shortcuts that only moved the number.

That applies well outside inference. Anything measured by one number gets gamed, by an agent or by a team. It is true of AI visibility too: a single score tells you very little. What tells you something is asking the engines the questions your buyers ask, reading who they name, and reading which pages they cite.

ClawWorld runs AI visibility work for B2B startups. Start with a free baseline report: 10 buyer questions, who ChatGPT and Claude name instead of you, and the pages they cite.

Get your free baseline report →