GEO
How We Measure Whether AI Recommends a Business, and What It Showed for Mandorla
Most advice about GEO, getting your business recommended by AI assistants, ends with a checklist. Add structured data. Get more reviews. Publish helpful content. All sensible, and almost none of it comes with a before and after, because almost nobody measures.
We wanted the before and after. So in July we built a tracker, pointed it at a client, and started the clock. The chart above is what it recorded.
Why you cannot just ask ChatGPT once
The obvious test is to open an assistant, type "best Italian restaurant in Chiang Mai" and see if your name comes up. The problem is that the answer is not the same twice.
When we asked one assistant for the best restaurants in Chiang Mai five times on 5 September, eleven different restaurants appeared across the five answers. A business that shows up once, or misses once, has learned almost nothing.
So the tracker never asks once. It asks the same question several times and reports a frequency: named in 3 of 3 answers, or in 0 of 5. That is a measurement. A single screenshot is an anecdote.
How the tracker works
Every week it runs the same fixed questions, worded the way a real customer would type them. For Mandorla Sicilian Bistrot, a restaurant we work with, those are "What are the best Sicilian or Italian restaurants in Chiang Mai?" and "Where should I eat Italian food in Chiang Mai?".
Each run follows a few rules, and each rule exists because we hit the problem it prevents.
- Live web search is always on. An assistant answering from memory tells you what was on the internet when it was trained, not what a customer hears today. Where an engine can answer without searching, we check the response for proof that a search happened, and discard it if there is none.
- A failed answer is thrown out, not counted as a miss. Sometimes the assistant says it could not find a reliable list, or offers directory sites instead of businesses. Counting that as "you were not named" would invent bad news. Those runs are dropped, and the number of usable answers is always shown, so 2 of 3 never pretends to be 2 of 5.
- Names must match as whole names. Early on, a venue listed simply as Italia matched the word "Italian" in every single answer and shot to the top of a ranking. Google Maps could not even find it. Now a name only counts when it appears whole, and every competitor is counted with exactly the same rule as the client.
- Every result has a model and a date. Each question gets one row per week in a database: how many usable answers, how many named the business, and the five names mentioned most often. Nothing is estimated.
The same measuring logic sits behind our AI visibility work for clients and the audits we run for prospects, so a client report and an audit can never disagree about the same business.
What it showed for Mandorla
When the tracker started on 16 July, the groundwork was already in place. Structured data describing Mandorla as a restaurant, with its menu, opening hours and reservations link, was live across the site. On 14 July we had published a menu page an assistant can actually read, and a guide to the best Italian food in Chiang Mai that openly names the other good restaurants in town, because a page that reads like a real answer is the kind that gets quoted.
Then we measured, on the first question, every week:
- 16 July: named in 0 of 5 answers.
- 20 July: 0 of 5.
- 27 July: 0 of 1, the only usable answer that week.
- 3 August: 3 of 3. The first time Mandorla was named, twenty days after the guide went live.
- 10 August: 2 of 3.
- 17 August: 3 of 3.
- 24 August: 2 of 2.
- 31 August: 2 of 3.
- 14 September: 4 of 4. There was no run on 7 September.
Never named in July. Named in every weekly measurement since 3 August, inside two months.
The work kept going alongside it. On 16 August the menu went up in English, Thai and Chinese, which matters because assistants often search in English even when they are asked in another language. And a QR review card handed over at the end of service took Mandorla's Google reviews from 139 on 14 July to 205 on 5 September, with the rating holding at 4.6.
What these numbers do not prove
A tracker is only worth something if it is honest about its limits, so here they are.
- The sample size moves. July's zeros are out of five answers. Several later weeks are out of two, three or four, because unusable answers are dropped. Going from never named to usually named is real. It is not a clean five against five.
- The everyday phrasing is still weak. On "Where should I eat Italian food in Chiang Mai?", Mandorla was named in 2 of 5 answers on 3 and 24 August, 1 of 5 on 31 August, and 0 of 5 on 14 September. That question leans on review sites and roundups, and that is where the work goes next.
- We cannot prove we caused it. The answers come from the live web, which changed over the same weeks for reasons that have nothing to do with us. What we can say is that Mandorla went from never appearing to appearing every week during the period we worked on it.
- One engine for the whole series. Every weekly figure comes from the same model, Claude with web search, so the comparison is like for like. On 5 September we also asked Perplexity five times and it named Mandorla in all five, but that was the first time we asked it, so there is nothing to compare it to yet.
Why we measure this way
The most damaging thing an AI visibility report can say is "AI never mentions you" when it does. The owner checks, sees their own name, and never trusts the report again. The opposite mistake, a flattering number built on one lucky answer, is just as bad, only slower to surface.
So every figure we show a client is a stored measurement: this question, this model, this date, named in this many answers out of this many. It is less exciting than a promise. It is also the only kind of GEO result we are willing to put our name to.
If you want to know whether AI assistants recommend your business today, we can run the same measurement on the questions your customers ask, before any work starts. That way there is a real baseline, and anything we claim later has something to be measured against.
Frequently asked questions
What is a GEO tracker?
A tool that asks AI assistants the questions your customers ask, several times each, and records how often your business is named. It turns "does AI recommend us?" from a guess into a number with a date on it.
Why ask the same question more than once?
Because the answer changes. The same question asked five times can return five different lists, so a single answer is an anecdote. We report how often a business was named out of how many usable answers.
Did the tracker prove the GEO work caused Mandorla's result?
No, and we say so. The answers come from the live web, which changes for many reasons. What it shows is that Mandorla went from never named to named in every weekly measurement during the period we worked on it.
Can you track my business?
Yes. We can measure where you stand on the questions your customers ask before any work starts, so there is a real baseline to compare against later.
Find out if AI recommends you
Book a free call and we will measure where you stand on the questions your customers actually ask, before anyone promises you anything.
Book a free call