Case study · AEO GEO · September 2026

AI Visibility Audit: The 4% Problem

A Wrocław renovation company wanted to know what AI assistants tell buyers looking for a contractor. This AI visibility audit asked three of them 10 buyer questions in 4 languages and measured 120 answers. The company was mentioned in 4 of the 96 answers where its name was not in the question.

Everything below comes from one measured run — models, prompts, and the code that ran them are named exactly, so the audit can be repeated and the numbers checked.

ravennapro.pl 120 answers gpt-5.5 · claude-sonnet-5 · gemini-3.7-flash PL · UK · RU · EN
Why measure this

Search is turning into a shortlist

Buyers increasingly ask an AI assistant (ChatGPT, Claude, Gemini) who to hire instead of scanning the top-10 Google results.

The question is not whether you rank but whether you are on the shortlist it returns — and if not, who is.

Method

Three assistants, one comparable test

The measurement is a Python script, run from the command line, that reads a prompt library, sends prompts in bulk to one model, and writes each answer back into that model's own column — three runs, one per model, produce the 120 answers.

Test matrix
DimensionValuesCount
Questions2 branded (does the assistant know this company?) + 8 unbranded (buyer looking for any contractor: best firm, price per m², shortlist with contacts, developer-state finishing, remote owner, language of service, how to choose, timeline)10
LanguagesPolish (primary market), Ukrainian, Russian, English4
Modelsgpt-5.5, claude-sonnet-5, gemini-3.7-flash — each with its own live web search enabled3
Answers10 × 4 × 3 — of which 96 are unbranded, the headline denominator120
Per-provider call shapes

Each provider exposes web search differently, so "with search on" means three different things.

# Anthropic — Messages API, server-side search tool
client.messages.stream(
    model="claude-sonnet-5", max_tokens=16000, system=SYSTEM_PROMPT,
    tools=[{"type": "web_search_20260209", "name": "web_search", "max_uses": 5}],
    messages=[{"role": "user", "content": prompt}])

# OpenAI — Responses API, `instructions` carries the system prompt
client.responses.stream(
    model="gpt-5.5", input=prompt, instructions=SYSTEM_PROMPT,
    tools=[{"type": "web_search"}])

# Google — google-genai, Search grounding
client.models.generate_content_stream(
    model="gemini-3.7-flash", contents=prompt,
    config=types.GenerateContentConfig(
        system_instruction=SYSTEM_PROMPT,
        tools=[types.Tool(google_search=types.GoogleSearch())]))
Headline

Findable by name, invisible among competitors

4%AI visibility among competitors

In the 96 answers where the question did not contain the company name, Ravenna Pro appeared 4 times. All four came from a single model. The other 92 answers handed the buyer a list of competitors.

6 of 8
unbranded questions with zero mentions across all models and languages
24 / 24
branded answers that place the company in the correct city
67%
of branded answers also surface a Warsaw registry address
4%
of branded answers can produce the company's phone number
Unbranded questionPLUKRUEN
Best turnkey renovation company
Cost per m² for a 60 m² flat
Five vetted contractors with contacts
Developer-state finishing, named estates
Owner lives abroad, remote management
Service in Ukrainian
How to choose a contractor
How long a renovation takes
mentioned by gpt-5.5 mentioned by claude-sonnet-5 mentioned by gemini-3.7-flash not mentioned
Diagnosis

Three mechanisms behind the 4%

1 · The site is a business card, not a source

The domain was cited 17 times across 120 answers — but 13 of those sit on the two branded questions, where the model was handed the name and went looking. On six of the eight unbranded questions it was cited zero times.

branded: what is it
7
branded: reliable?
6
developer-state
2
Ukrainian service
2
other six questions
0
Citations of ravennapro.pl, out of 12 possible per question.

2 · The entity resolves to two different companies

Every branded answer places the business in Wrocław — but two thirds of them simultaneously pull a Warsaw registered address and a PKD classification of "architecture / design services" out of the company registries. A quarter drift toward the Italian city of Ravenna, concentrated in Russian and English. The registry record is, in effect, arguing with the marketing site, and the model reports both.

correct city
24
Warsaw address
16
Wrocław street
16
Italy confusion
6
phone number
1
Out of 24 branded answers. The phone number — the one thing a buyer actually needs — was reproduced once.

3 · Shortlists are assembled from directories, not from websites

When a model builds a list of contractors it leans on marketplaces and ranking pages, not on company sites. Oferteo and Fixly are the two most-cited domains in the entire corpus. Ravenna Pro has no profile on either, so it cannot enter the list those pages feed — regardless of the quality of its work.

oferteo.pl
22
fixly.pl
19
wykonczeniawroclaw.pl
18
ravennapro.pl
17
wroclaw.pl
15
kodo.pl
11
Answers citing each domain, out of 120. Oferteo + Fixly = 41 citations.
Competitive set

Who gets recommended instead

Not "competitors" in the abstract — the firms models actually name when asked for advice, and the specific claims they repeat about each.

RUKOS
18
KRASNALBUD
17
MAXYSBUD
15
KODO
14
NEO REMONT
13
DECOROOM
13
GP SYSTEM
10
PPUH ART-RAD
8
Ravenna Pro
4
FirmWhat the models repeat about themSource
RUKOStrading since 2014, "design through handover" model, 5/5 from 45 reviewsown site + Oferteo
PPUH ART-RAD5.0/5 from 12 reviews, 20+ years, itemised quote, schedule, warranty, VAT invoiceOferteo
MAXYSBUDfull scope including developer handover inspection, 2-year warranty, named estatesown site
NEO REMONT2 M PLN contractor liability insurance, post-completion servicethird-party ranking
SIGNUM Interiorsprices from 1390 PLN/m², worked on Olimpia Port, Browary Wrocławskie, Port Popowiceown site
The pattern. None of these claims requires budget or reputation — they are facts published in a legible place: founding year, warranty length, insurance sum, review count, estate names, a starting price. Across the corpus, warranty terms appear in 63 of 120 answers, years in business in 22, insurance in 18, review counts in 17, named estates in 17, prices in 8. That list doubles as the content brief.
Model behaviour

The three assistants are not interchangeable

Visibility in one says very little about visibility in another — which is the practical argument for measuring more than one.

ModelAvg lengthSources / answerOwn-site citationsUnbranded wins
gpt-5.53 8284.312 / 404
claude-sonnet-52 9871.94 / 400
gemini-3.7-flash2 4161.41 / 400

gpt-5.5 searched hardest — 4.3 distinct domains per answer against Gemini's 1.4 — and that alone explains all four unbranded mentions. The company's visibility currently rests on one model happening to dig deeper, which is luck rather than position.

gemini-3.7-flash is the alarming one. One citation in forty, zero unbranded mentions. It runs on live Google Search, so it is the closest available proxy for what Google's own stack knows — and by that proxy the company is close to absent from the ecosystem where most of its buyers decide. It is a proxy, not the AI Overview itself; that surface has no API.

Language differences
LanguageOwn-site citationsBrand mentionsReading
Polish57 / 30home market, densest competition, the only language where models quote prices at all
Ukrainian47 / 30weakest competition — the existing /uk/ version is a real, underused edge
Russian58 / 30most mentions and most Italy confusion; no /ru/ version exists
English36 / 30longest answers, most sources consulted, fewest citations of this site

The single clearest causal finding in the run came from the Ukrainian-language question. Both of its hits happened because a /uk/ version of the site exists; the model opened it, confirmed the page was genuinely in Ukrainian, ranked the company first, and — uniquely in all 120 answers — reproduced the phone number.

What follows

The fixes the data actually supports

Ordered so that nothing gets attributed to the wrong company before the rest is worth doing.

Limits

What this measurement cannot tell you

The point of naming these is that the same 120 cells can be re-run after the fixes land and compared honestly. The baseline to beat: 4 unbranded mentions out of 96, and 17 domain citations out of 120.