We are our own first case study. Measured in the open.
ClawWorld pivoted from an AI-agent social network to AEO services in 2026 — which means we start from the worst position an AEO client can be in: answer engines either describe our old company, confuse us with a Las Vegas claw-machine arcade that shares our name, or have nothing to say at all. Instead of hiding that, we are running our own service on our own domain and publishing the measurements here as they happen.
Three measurement rounds are published below — 1,172 judged answers across three engines. Two baseline rounds were measured in August 2026 before any optimization content shipped; the third is the first re-measurement, taken seven days after our on-site content pack went live. All of it is here in full: the identity confusion, the pre-registered verdict on that first week (short version: too early to call, and we say so), the limits of the method, and two data-integrity bugs we found in our own pipeline.
Methodology
The same loop we sell, applied to ourselves.
A frozen question set
24 buyer questions in four layers — brand-direct, head category, long-tail category, and problem-phrased — frozen before measurement began. Exact wordings stay sealed until the experiment ends, so our own site can’t contaminate the questions it is measured on; the layer structure and paraphrased examples are public here.
Repeated sampling, not single asks
Every question runs five times per engine per round. A single ask is sampling noise: the same prompt can name you at noon and forget you at midnight. Rates come from repetition or they are not rates. We call the resulting metric rSoV — Resampled Share of Voice: the share of answers mentioning a brand, earned through repetition.
Control brands we never touch
Every round also measures competitor brands we do no work for — two category leaders and two small tools. If the whole platform shifts, the controls shift with it, and that drift gets subtracted from our numbers instead of being claimed as progress.
Published either way
Each round produces mention rates, citation rates, identity-confusion rates, and the sources engines actually cited — plus an operations log of what we changed and when. If a number refuses to move, it gets published refusing to move.
Results · R0 → R1 → W2end
1,172 measured answers. Here is the whole picture.
Three rounds, three engines, 1,172 judged answers: two baselines (R0, R1) measured before any optimization content existed, then a re-measurement (W2end) seven days after it shipped. Every question was asked five times per engine per round; four competitor brands were measured in the same runs as controls. Nothing below is cherry-picked — this is the whole result set.
1. In category questions, we do not exist
Across the head-category, long-tail, and problem-phrased layers, ClawWorld was mentioned in exactly zero answers — on every engine, in both rounds. Brand-direct questions do return an answer, but as the next table shows, that answer is not about us.
Engine
Brand-direct
Head category
Long-tail
Problem-phrased
Claude (Brave index)
100% → 100% → 84.2%
0% → 0% → 0%
0% → 0% → 0%
0% → 0% → 0%
OpenAI stack
100% → 100% → 100%
0% → 0% → 0%
0% → 0% → 0%
0% → 0% → 0%
Gemini (Google index)
100% → 100% → 100%
0% → 0% → 0%
0% → 0% → 0%
0% → 0% → 0%
Mention rate, R0 → R1 → W2end. Brand-direct and head-category figures use the searched-answers subset; the other two use all answers.
2. Asked about us by name, engines describe someone else
This is the identity problem stated as a number. When engines are asked directly about ClawWorld, they answer confidently — about a Las Vegas claw-machine arcade chain that shares our name, about unrelated AI products with similar names, or about the company we used to be before the pivot. Each row below counts how the entity was resolved across brand-direct answers in the latest round.
Claude (Brave index)
OpenAI stack
Gemini (Google index)
Us (correct)
Arcade chain
Other “Claw” product
Our pre-pivot company
Mixed / unsure
Entity resolution across brand-direct answers, latest round. A healthy brand’s bar would be almost entirely teal — on two of the three engines, ours is still mostly other people. The OpenAI stack started flipping to us in round three; that story is told below.
Engine
Us (correct)
Arcade chain
Other “Claw” product
Our pre-pivot company
Mixed / unsure
Claude (Brave index)
0/16
13/16
2/16
0/16
1/16
OpenAI stack
6/8
0/8
0/8
1/8
1/8
Gemini (Google index)
10/20
5/20
0/20
0/20
5/20
Three engines, three different wrong answers — which is itself the finding: identity is resolved per index, not globally.
3. How much a number moves when nothing happens
The control brands received no work from us at all, so whatever their numbers did between rounds is pure platform noise. The table shows the largest swing any control took between two consecutive rounds — the bar any real effect has to clear, and it is larger than most AEO reporting admits. Anyone showing you a single-run before/after without a control is showing you this table’s contents relabelled as progress.
Engine
Brand-direct
Head category
Long-tail
Problem-phrased
Claude (Brave index)
±0pp
±3.2pp
±12.1pp
±13.3pp
OpenAI stack
—
±0pp
±12.5pp
±8.3pp
Gemini (Google index)
±0pp
±14.9pp
±10pp
±3.3pp
4. What these engines actually cite
The most-cited domains across the latest round. This is the map of where the work has to happen: not our homepage, but the pages engines already reach for when they answer our category’s questions.
Domain
Times cited
youtube.com
285
ayzeo.com
128
tryprofound.com
106
trakkr.ai
85
claw-world.app
74
reddit.com
70
clawworldusa.com
55
peec.ai
51
g2.com
50
geoptie.com
44
Round three · first re-measurement, Aug 23
Seven days after the content shipped: too early to call — and two things moved anyway.
On Aug 16 we shipped our on-site content pack — the FAQ, the comparison page, two mechanics explainers, and this page — then froze the site and touched nothing for a week. On Aug 23 we re-ran the exact same measurement: same frozen questions, same pinned model versions, same controls. The pass rule was written down before the round ran: a change only counts if it beats twice the control-brand noise band on at least two engines in the same direction.
The pre-registered verdict: no layer passed. In the three category layers — the questions where buyers don’t know our name yet — we are still at exactly zero, on every engine. And this zero is a harder zero than the baseline’s: before the round ran, we confirmed all five new pages were indexed in both Google and Bing. So “the engines haven’t seen it” is off the table. The honest reading is visible but not yet chosen — the pages are in the index; they have not been picked up into answers.
We are not calling that failure, and we are also not spinning it. Retrieval pools take days to weeks to absorb new pages; a seven-day window reads an early signal, not a full effect. The call that matters — works or doesn’t — needs the next rounds. What we can publish now is what moved and what didn’t, with the noise band attached.
What moved: the OpenAI stack started recognizing us
The clearest movement in the round is entity resolution on the OpenAI stack. Asked brand-direct questions, it resolved “ClawWorld” to the actual us in 0 of 8 answers in round one, 1 of 8 in round two — and 6 of 8 in round three, with answers describing our pre-pivot company falling from 4 to 1.
OpenAI stack
Gemini
Claude
“Resolved to us” share of brand-direct answers, per engine. The teal line is the story; the gray lines are context. Gemini’s wobble is discussed below; Claude has yet to resolve us correctly once.
The same round produced a second first: the OpenAI stack cited claw-world.app in 7 of 48 answers — after two straight rounds of never citing us at all. Gemini’s citation rate has drifted up within a two-point band since round one; Claude has never cited us.
OpenAI stack
Gemini
Claude
Answers citing claw-world.app, all question layers, per engine. Note the axis tops out at 20% — these are early, small movements, drawn to their own scale rather than inflated.
OpenAI stack, brand-direct
R0
R1
W2end
Resolved to us (correct)
0/8
1/8
6/8
Described our pre-pivot company
4/8
4/8
1/8
Both trends are monotonic across three rounds, which is what makes them worth reporting. They are still descriptive: entity resolution has no pre-registered noise band, the sample is 8 answers per round, and neither figure is a verdict. If the next round holds them, they become evidence.
What we are not claiming: Gemini’s “resolved to us” share went 9 → 6 → 10 out of 20 across the three rounds. Reading the last step as improvement would mean picking the dip as the starting point — that is the before/after trick this page exists to reject. It is inside its own wobble, and it stays unclaimed.
Correction: the 53-point swing was our bug, not the engine
Until Sep 26 this section told a different story. It said that in the problem-phrased layer Claude’s answers had stopped triggering web search this round, that the control brands’ mention rates had collapsed with it by up to 53 points, and that the controls had caught engine behavior which would otherwise have been reported as a result. The collapse was not engine behavior. 25 of the 30 answers in that cell were an API error, “529 Overloaded”, which the Claude command-line tool printed where an answer should be and our pipeline stored as one.
With the round’s 38 error records removed, that cell holds 5 real answers. None of them searched, so that part stands; the control brands moved 3.3 points against the previous round, not 53. The tables above use the corrected round, which is why this page now counts 38 fewer answers than it did before Sep 26. The round’s verdict did not change: no layer passed before the correction and none passes after it. What happened and what we changed in the collector is written up in Our AI visibility monitor counted 38 API errors as answers.
What this data cannot tell you
The limits, stated up front.
The two rounds are hours apart, not weeks. The second round ran as soon as both search indexes confirmed they had re-crawled our new site. So the difference between them measures instrument stability, not how long an identity change takes to propagate. We cannot claim a propagation timeline from this data, and we will not.
ChatGPT and Perplexity are sampled by hand, not automated. Automating consumer chat interfaces violates their terms, so the largest-traffic engine is covered by a small manual sample. Conclusions extrapolated to ChatGPT are inference, not measurement.
Five samples per question is enough to catch a step, not a slope. This design detects a brand going from zero to visible; it cannot resolve a one- or two-point drift, and it says so rather than reporting noise as signal.
We broke our own data twice, and both times are on the record. During one collection run the API returned a rate-limit notice that our script stored as if it were an engine answer, quietly turning 75 records into false negatives. We caught it because a control brand’s number moved in a way reality does not allow, then re-collected the affected records and added a detector. In round three a different server error, “529 Overloaded”, slipped past that detector and turned 38 Claude records into false negatives. We found it three days later reading raw transcripts; this page carried the wrong figures until Sep 26 (see the correction above). Both incidents are written up in full in our measurement log, the second also on our blog — a measurement practice that has never found a bug in itself is a measurement practice nobody is checking.
New home page as prerendered static HTML; four retired product pages set to noindex; robots.txt, sitemap, llms.txt and Organization/Service structured data rebuilt; blog titles fixed; company description updated everywhere an engine looks.
Week 1 — crawl fixes and baseline (Aug 15–16)
Bing Webmaster Tools set up and sitemap submitted; 151 URLs pushed via IndexNow; fixed a defect where nonexistent URLs returned an indexable empty page. Both indexes confirmed the new identity, then the baseline rounds above were measured — before any optimization content shipped.
Week 2 — on-site content pack (Aug 17)
Shipped the same day this page went live: a ten-question FAQ with structured data, an honest comparison page with vendor-verified pricing, two mechanics explainers, and this page. Then the site was frozen for the measurement window.
Week 2 — effect window and re-measurement (Aug 18–24)
Content frozen all week; only routine blog posts and two style-level fixes (logged, no indexable text changed). Aug 23: confirmed all five new pages indexed in Google and Bing, then ran round three — same frozen questions, same pinned models, same controls. Verdict published above: no layer passed the pre-registered threshold; the OpenAI-stack identity shift and first citations are on the record as early signals, not results.
Week 3 — cleanup after round three (Aug 26–29)
Aug 26: found that 38 of round three’s Claude answers were API errors (the correction above); fixed the collector and corrected the round in our measurement repository, but did not update this page until Sep 26. Same day: added a note on the pivot, with links to the home page, the comparison page and this page, to the two old blog posts Brave had indexed, and submitted eight URLs to Brave. Aug 28: took down the old product’s pages and regenerated robots.txt, the sitemap and llms.txt; added our Wikidata and Crunchbase entries to our structured data. Aug 29: one contact email across every page; the free-audit button now opens an on-site form.
Week 4 — new pages inside the measurement window (Sep 5–8)
Sep 5: home page redesigned (a stronger statement of who we are, two more FAQ answers). Sep 7: round four measured as a mid-point check; it is not on this page yet, because the hand check of its judged answers has not passed. Sep 7, late evening: the free audit page went live. Sep 8: the pricing and methodology pages went live, and the comparison page’s pricing was restructured. The Sep 5 and Sep 8 changes broke a site freeze we had set for this window; all of them are logged as change sources for the next verdict.
Entity fixes (Sep 14–22)
Sep 14: a Wikidata administrator deleted our item for notability. Sep 22: removed it from our structured data; added a one-line disambiguation (we are not the arcade) to the home page and llms.txt; gave three FAQ answers a short first sentence and a fourth the GEO alias; named Claude’s search crawler in robots.txt; replaced the old product’s sign-up button at the foot of 124 blog posts with a link to our free baseline report.
Round three corrected on this page (Sep 26)
Re-exported this page with the corrected round three (38 error records removed, 1,210 answers down to 1,172), rewrote the paragraph that had presented the 53-point swing as engine behavior, and published the write-up on our blog. Round four is still in hand review; the next round has not been run yet.
Data generated from the measurement repository on 2026-09-26.