Lab · live case study

We are our own first case study.
Measured in the open.

ClawWorld pivoted from an AI-agent social network to AEO services in 2026 — which means we start from the worst position an AEO client can be in: answer engines either describe our old company, confuse us with a Las Vegas claw-machine arcade that shares our name, or have nothing to say at all. Instead of hiding that, we are running our own service on our own domain and publishing the measurements here as they happen.

The baseline below was measured in August 2026, before any optimization content shipped — 807 judged answers across two rounds and three engines. It is published here in full, including the identity confusion, the limits of the method, and a data-integrity bug we found in our own pipeline.

Methodology

The same loop we sell, applied to ourselves.

  • A frozen question set

    24 buyer questions in four layers — brand-direct, head category, long-tail category, and problem-phrased — frozen before measurement began. Exact wordings stay sealed until the experiment ends, so our own site can’t contaminate the questions it is measured on; the layer structure and paraphrased examples are public here.

  • Repeated sampling, not single asks

    Every question runs five times per engine per round. A single ask is sampling noise: the same prompt can name you at noon and forget you at midnight. Rates come from repetition or they are not rates.

  • Control brands we never touch

    Every round also measures competitor brands we do no work for — two category leaders and two small tools. If the whole platform shifts, the controls shift with it, and that drift gets subtracted from our numbers instead of being claimed as progress.

  • Published either way

    Each round produces mention rates, citation rates, identity-confusion rates, and the sources engines actually cited — plus an operations log of what we changed and when. If a number refuses to move, it gets published refusing to move.

Baseline · R0 and R1, 2026-08-16

807 measured answers.
Here is where we start from.

Two rounds, three engines, 807 judged answers. Every question was asked five times per engine per round; four competitor brands were measured in the same runs as controls. Nothing below is cherry-picked — this is the whole result set.

1. In category questions, we do not exist

Across the head-category, long-tail, and problem-phrased layers, ClawWorld was mentioned in exactly zero answers — on every engine, in both rounds. Brand-direct questions do return an answer, but as the next table shows, that answer is not about us.

EngineBrand-directHead categoryLong-tailProblem-phrased
Claude (Brave index)100% → 100%0% → 0%0% → 0%0% → 0%
OpenAI stack100% → 100%0% → 0%0% → 0%0% → 0%
Gemini (Google index)100% → 100%0% → 0%0% → 0%0% → 0%

2. Asked about us by name, engines describe someone else

This is the identity problem stated as a number. When engines are asked directly about ClawWorld, they answer confidently — about a Las Vegas claw-machine arcade chain that shares our name, about unrelated AI products with similar names, or about the company we used to be before the pivot. Each row below counts how the entity was resolved across brand-direct answers in the latest round.

Claude (Brave index)
OpenAI stack
Gemini (Google index)
  • Us (correct)
  • Arcade chain
  • Other “Claw” product
  • Our pre-pivot company
  • Mixed / unsure
EngineUs (correct)Arcade chainOther “Claw” productOur pre-pivot companyMixed / unsure
Claude (Brave index)0/2012/200/200/208/20
OpenAI stack1/80/80/84/83/8
Gemini (Google index)6/205/200/200/209/20

3. How much a number moves when nothing happens

The control brands received no work from us at all, so whatever their numbers did between the two rounds is pure platform noise. That noise is the bar any real effect has to clear — and it is larger than most AEO reporting admits. Anyone showing you a single-run before/after without a control is showing you this table’s contents relabelled as progress.

EngineBrand-directHead categoryLong-tailProblem-phrased
Claude (Brave index)±0pp±3.2pp±4pp±13.3pp
OpenAI stack±0pp±6.3pp±8.3pp
Gemini (Google index)±0pp±8pp±10pp±3.3pp

4. What these engines actually cite

The most-cited domains across the latest round. This is the map of where the work has to happen: not our homepage, but the pages engines already reach for when they answer our category’s questions.

DomainTimes cited
youtube.com289
tryprofound.com133
ayzeo.com114
trakkr.ai98
facebook.com57
clawworldusa.com56
reddit.com55
g2.com52
claw-world.app45
discoveredlabs.com44

What this data cannot tell you

The limits, stated up front.

The two rounds are hours apart, not weeks. The second round ran as soon as both search indexes confirmed they had re-crawled our new site. So the difference between them measures instrument stability, not how long an identity change takes to propagate. We cannot claim a propagation timeline from this data, and we will not.

ChatGPT and Perplexity are sampled by hand, not automated. Automating consumer chat interfaces violates their terms, so the largest-traffic engine is covered by a small manual sample. Conclusions extrapolated to ChatGPT are inference, not measurement.

Five samples per question is enough to catch a step, not a slope. This design detects a brand going from zero to visible; it cannot resolve a one- or two-point drift, and it says so rather than reporting noise as signal.

We broke our own data once, and it is on the record. During one collection run the API returned a rate-limit notice that our script stored as if it were an engine answer, quietly turning 75 records into false negatives. We caught it because a control brand’s number moved in a way reality does not allow, then re-collected the affected records and added a detector. The incident is written up in full in our measurement log — a measurement practice that has never found a bug in itself is a measurement practice nobody is checking.

Operations log

Everything we changed, dated.