Lab · live case study

We are our own first case study.
Measured in the open.

ClawWorld pivoted from an AI-agent social network to AEO services in 2026 — which means we start from the worst position an AEO client can be in: answer engines either describe our old company, confuse us with a Las Vegas claw-machine arcade that shares our name, or have nothing to say at all. Instead of hiding that, we are running our own service on our own domain and publishing the measurements here as they happen.

Three measurement rounds are published below — 1,172 judged answers across three engines. Two baseline rounds were measured in August 2026 before any optimization content shipped; the third is the first re-measurement, taken seven days after our on-site content pack went live. All of it is here in full: the identity confusion, the pre-registered verdict on that first week (short version: too early to call, and we say so), the limits of the method, and two data-integrity bugs we found in our own pipeline.

Methodology

The same loop we sell, applied to ourselves.

  • A frozen question set

    24 buyer questions in four layers — brand-direct, head category, long-tail category, and problem-phrased — frozen before measurement began. Exact wordings stay sealed until the experiment ends, so our own site can’t contaminate the questions it is measured on; the layer structure and paraphrased examples are public here.

  • Repeated sampling, not single asks

    Every question runs five times per engine per round. A single ask is sampling noise: the same prompt can name you at noon and forget you at midnight. Rates come from repetition or they are not rates. We call the resulting metric rSoV — Resampled Share of Voice: the share of answers mentioning a brand, earned through repetition.

  • Control brands we never touch

    Every round also measures competitor brands we do no work for — two category leaders and two small tools. If the whole platform shifts, the controls shift with it, and that drift gets subtracted from our numbers instead of being claimed as progress.

  • Published either way

    Each round produces mention rates, citation rates, identity-confusion rates, and the sources engines actually cited — plus an operations log of what we changed and when. If a number refuses to move, it gets published refusing to move.

Results · R0 → R1 → W2end

1,172 measured answers.
Here is the whole picture.

Three rounds, three engines, 1,172 judged answers: two baselines (R0, R1) measured before any optimization content existed, then a re-measurement (W2end) seven days after it shipped. Every question was asked five times per engine per round; four competitor brands were measured in the same runs as controls. Nothing below is cherry-picked — this is the whole result set.

1. In category questions, we do not exist

Across the head-category, long-tail, and problem-phrased layers, ClawWorld was mentioned in exactly zero answers — on every engine, in both rounds. Brand-direct questions do return an answer, but as the next table shows, that answer is not about us.

EngineBrand-directHead categoryLong-tailProblem-phrased
Claude (Brave index)100% → 100% → 84.2%0% → 0% → 0%0% → 0% → 0%0% → 0% → 0%
OpenAI stack100% → 100% → 100%0% → 0% → 0%0% → 0% → 0%0% → 0% → 0%
Gemini (Google index)100% → 100% → 100%0% → 0% → 0%0% → 0% → 0%0% → 0% → 0%

2. Asked about us by name, engines describe someone else

This is the identity problem stated as a number. When engines are asked directly about ClawWorld, they answer confidently — about a Las Vegas claw-machine arcade chain that shares our name, about unrelated AI products with similar names, or about the company we used to be before the pivot. Each row below counts how the entity was resolved across brand-direct answers in the latest round.

Claude (Brave index)
OpenAI stack
Gemini (Google index)
  • Us (correct)
  • Arcade chain
  • Other “Claw” product
  • Our pre-pivot company
  • Mixed / unsure
EngineUs (correct)Arcade chainOther “Claw” productOur pre-pivot companyMixed / unsure
Claude (Brave index)0/1613/162/160/161/16
OpenAI stack6/80/80/81/81/8
Gemini (Google index)10/205/200/200/205/20

3. How much a number moves when nothing happens

The control brands received no work from us at all, so whatever their numbers did between rounds is pure platform noise. The table shows the largest swing any control took between two consecutive rounds — the bar any real effect has to clear, and it is larger than most AEO reporting admits. Anyone showing you a single-run before/after without a control is showing you this table’s contents relabelled as progress.

EngineBrand-directHead categoryLong-tailProblem-phrased
Claude (Brave index)±0pp±3.2pp±12.1pp±13.3pp
OpenAI stack—±0pp±12.5pp±8.3pp
Gemini (Google index)±0pp±14.9pp±10pp±3.3pp

4. What these engines actually cite

The most-cited domains across the latest round. This is the map of where the work has to happen: not our homepage, but the pages engines already reach for when they answer our category’s questions.

DomainTimes cited
youtube.com285
ayzeo.com128
tryprofound.com106
trakkr.ai85
claw-world.app74
reddit.com70
clawworldusa.com55
peec.ai51
g2.com50
geoptie.com44

Round three · first re-measurement, Aug 23

Seven days after the content shipped:
too early to call — and two things moved anyway.

On Aug 16 we shipped our on-site content pack — the FAQ, the comparison page, two mechanics explainers, and this page — then froze the site and touched nothing for a week. On Aug 23 we re-ran the exact same measurement: same frozen questions, same pinned model versions, same controls. The pass rule was written down before the round ran: a change only counts if it beats twice the control-brand noise band on at least two engines in the same direction.

The pre-registered verdict: no layer passed. In the three category layers — the questions where buyers don’t know our name yet — we are still at exactly zero, on every engine. And this zero is a harder zero than the baseline’s: before the round ran, we confirmed all five new pages were indexed in both Google and Bing. So “the engines haven’t seen it” is off the table. The honest reading is visible but not yet chosen — the pages are in the index; they have not been picked up into answers.

We are not calling that failure, and we are also not spinning it. Retrieval pools take days to weeks to absorb new pages; a seven-day window reads an early signal, not a full effect. The call that matters — works or doesn’t — needs the next rounds. What we can publish now is what moved and what didn’t, with the noise band attached.

What moved: the OpenAI stack started recognizing us

The clearest movement in the round is entity resolution on the OpenAI stack. Asked brand-direct questions, it resolved “ClawWorld” to the actual us in 0 of 8 answers in round one, 1 of 8 in round two — and 6 of 8 in round three, with answers describing our pre-pivot company falling from 4 to 1.

  • OpenAI stack
  • Gemini
  • Claude
0%25%50%75%100%R0R1W2endGemini · R0: 45% of 20 answersGemini · R1: 30% of 20 answersGemini · W2end: 50% of 20 answersClaude · R0: 0% of 20 answersClaude · R1: 0% of 20 answersClaude · W2end: 0% of 16 answersOpenAI stack · R0: 0% of 8 answersOpenAI stack · R1: 12.5% of 8 answersOpenAI stack · W2end: 75% of 8 answers75% OpenAI stack50% Gemini0% Claude

The same round produced a second first: the OpenAI stack cited claw-world.app in 7 of 48 answers — after two straight rounds of never citing us at all. Gemini’s citation rate has drifted up within a two-point band since round one; Claude has never cited us.

  • OpenAI stack
  • Gemini
  • Claude
0%5%10%15%20%R0R1W2endGemini · R0: 9.2% of 120 answersGemini · R1: 10.8% of 120 answersGemini · W2end: 11.7% of 120 answersClaude · R0: 0% of 116 answersClaude · R1: 0% of 115 answersClaude · W2end: 0% of 81 answersOpenAI stack · R0: 0% of 48 answersOpenAI stack · R1: 0% of 48 answersOpenAI stack · W2end: 14.6% of 48 answers14.6% OpenAI stack11.7% Gemini0% Claude
OpenAI stack, brand-directR0R1W2end
Resolved to us (correct)0/81/86/8
Described our pre-pivot company4/84/81/8

What we are not claiming: Gemini’s “resolved to us” share went 9 → 6 → 10 out of 20 across the three rounds. Reading the last step as improvement would mean picking the dip as the starting point — that is the before/after trick this page exists to reject. It is inside its own wobble, and it stays unclaimed.

Correction: the 53-point swing was our bug, not the engine

Until Sep 26 this section told a different story. It said that in the problem-phrased layer Claude’s answers had stopped triggering web search this round, that the control brands’ mention rates had collapsed with it by up to 53 points, and that the controls had caught engine behavior which would otherwise have been reported as a result. The collapse was not engine behavior. 25 of the 30 answers in that cell were an API error, “529 Overloaded”, which the Claude command-line tool printed where an answer should be and our pipeline stored as one.

With the round’s 38 error records removed, that cell holds 5 real answers. None of them searched, so that part stands; the control brands moved 3.3 points against the previous round, not 53. The tables above use the corrected round, which is why this page now counts 38 fewer answers than it did before Sep 26. The round’s verdict did not change: no layer passed before the correction and none passes after it. What happened and what we changed in the collector is written up in Our AI visibility monitor counted 38 API errors as answers.

What this data cannot tell you

The limits, stated up front.

The two rounds are hours apart, not weeks. The second round ran as soon as both search indexes confirmed they had re-crawled our new site. So the difference between them measures instrument stability, not how long an identity change takes to propagate. We cannot claim a propagation timeline from this data, and we will not.

ChatGPT and Perplexity are sampled by hand, not automated. Automating consumer chat interfaces violates their terms, so the largest-traffic engine is covered by a small manual sample. Conclusions extrapolated to ChatGPT are inference, not measurement.

Five samples per question is enough to catch a step, not a slope. This design detects a brand going from zero to visible; it cannot resolve a one- or two-point drift, and it says so rather than reporting noise as signal.

We broke our own data twice, and both times are on the record. During one collection run the API returned a rate-limit notice that our script stored as if it were an engine answer, quietly turning 75 records into false negatives. We caught it because a control brand’s number moved in a way reality does not allow, then re-collected the affected records and added a detector. In round three a different server error, “529 Overloaded”, slipped past that detector and turned 38 Claude records into false negatives. We found it three days later reading raw transcripts; this page carried the wrong figures until Sep 26 (see the correction above). Both incidents are written up in full in our measurement log, the second also on our blog — a measurement practice that has never found a bug in itself is a measurement practice nobody is checking.

Operations log

Everything we changed, dated.