Ask the assistants directly, ask the same way every time, and write down what they say. That is the whole method. It takes about an hour to set up and it is the only way to know whether anything you do to your site changes how AI systems describe your business.
Almost nobody does this, which means almost nobody making claims about AI visibility has a baseline to compare against.
Build the prompt set first
You need somewhere between ten and twenty questions, fixed, written down, reused unchanged every time you test. Changing the wording between tests destroys comparability, which is the entire value of the exercise.
Cover four types:
- Category with location. “Who are the best web designers in West Palm Beach?” This is the one that matters most commercially.
- Problem framed. “My small business website is slow and does not rank. Who in South Florida can help?” People increasingly describe a problem rather than name a service.
- Comparative. “What should I look for when hiring an SEO company in Florida?” You are checking whether your framing shows up, not just your name.
- Direct. “What is Palm Projects Media?” This tests whether the assistant can resolve you as an entity at all, which is the foundation for everything else.
Test across assistants, and expect them to disagree
ChatGPT, Google’s AI Overviews, Perplexity, Claude and Copilot draw on different indexes and different retrieval methods. Being cited well in one says little about the others. Perplexity leans heavily on live retrieval and shows its sources. ChatGPT’s behavior varies by whether search is invoked. Google’s overviews lean on its own index.
Run every prompt through every assistant you care about. It is tedious the first time and quick afterward.
What to record
For each prompt and each assistant, note four things: whether you were mentioned at all, roughly where in the answer, which sources were cited, and which competitors appeared.
The competitor list is the most useful column and the one people skip. If the same three businesses appear across every prompt, go and look at their pages. Something about how they are structured is making them retrievable, and it is usually visible.
A spreadsheet is sufficient. The date matters as much as the answer.
Reading the results honestly
Assistants are not deterministic. The same prompt can return different answers minutes apart, and a single result proves nothing. What you are looking for is the pattern across a set of prompts and across repeated tests, not any individual answer.
Be careful with the temptation to over interpret an early win or an early absence. One mention is noise. Consistent mention across ten prompts over three months is a signal.
What to do with an unflattering baseline
If you are absent everywhere, which is the normal starting position for a small business, work backwards from the reasons pages get retrieved: answering questions directly and early, in self contained sections, from a site whose content exists in the raw HTML, belonging to an entity the model can identify.
Then re-run the same prompt set in ninety days. Not sooner. Retrieval indexes update slowly and testing weekly will only show you noise.
What actually moves the result
Once you have a baseline, the interventions worth making are not mysterious, and they are the same ones that improve the site for everyone.
Answer the exact questions in your prompt set, on your own pages, directly and in your own words. If you tested “who are the best web designers in West Palm Beach” and you have no page that addresses how to choose one, there is nothing for a model to retrieve from you on that question.
Make sure your business name, service description and location appear together in unambiguous language rather than being scattered across a site that never quite states them plainly. Models are careful about attributing claims to entities they cannot resolve confidently.
Then be patient in a way that feels uncomfortable. Retrieval indexes update on their own schedule, and a change you make today can take months to be reflected. This is the main reason quarterly testing is the right cadence: anything faster measures variance and tempts you into changing things that were working.
Questions
How often should I re-test?
Quarterly for a stable business. Monthly if you are actively publishing and want a faster read. More often than that measures variance rather than progress.
Can I pay a tool to do this?
Several exist and they save time at scale. For a business with one location and one category, a spreadsheet and an hour a quarter gives you the same information, and doing it manually the first time teaches you more about why the answers look the way they do.
Does being mentioned actually bring business?
It brings a kind of attention that is difficult to attribute, because someone who reads about you in an AI answer and then searches your name arrives as direct or branded traffic. Watch branded search volume alongside your visibility testing. If one rises with the other, that is your answer.
If you want a baseline built for your category and a plan to move it, that is what an AI search engagement covers.