KnowledgeMeasurement
How do you develop buyer prompts for measuring AI visibility?
Buyer prompts are developed from the company's knowledge base: offering, target customers, pricing logic, competitors. This produces 50 questions per service area that reflect what real buyers ask: partly direct provider searches, partly described problems. The client reviews the set, and it then stays stable so that changes remain readable as a trend over weeks.
What is a buyer prompt?
A buyer prompt is the complete question a potential customer asks an AI system when looking for a provider or wanting a problem solved. It differs from a classic keyword in its form: «fiduciary Zurich» becomes something like «Our board wants to outsource the bookkeeping. Which fiduciary firm in the Zurich area specializes in SMEs?»
There are two basic forms. Good prompts sound like customers, not like marketing.
Direct search
Asks for providers in a category.
Problem description
Describes a business situation and asks who can help; nobody asks for a company by name.
Would a customer really type in prompts like these?
Often not word for word, and the measurement still works. Modern AI search decomposes plain requests into more precise variants: the user writes «I am looking for a fiduciary», and the system enriches the request in the background with context (location, history, stored preferences) and generates a series of more specific queries from it. The technical term is fan-out.
Buyer prompts are therefore working hypotheses about those variants, not a bet on an exact wording. They are developed from the company’s knowledge base during onboarding, reviewed by the client, and then sharpened continuously. A prompt that turns out to be unrealistic is replaced, like a wrong keyword in SEO.
What makes up a good prompt set?
Mixture, not mass. For transparency, here is how the set we use to measure Carigiet GEO’s own visibility is built (it runs in Swiss Standard German; the examples below are translated):
| Building block | Share in our own set | Purpose |
|---|---|---|
| Category prompts | around a third | measure baseline visibility in our own provider category |
| Industry prompts | 5 target industries with 5 prompts each | measure whether we appear where defined target customers ask |
| Scenario prompts | the rest | capture concrete trigger moments: a competitor gets recommended, a false statement is discovered, a website relaunch, a cost question |
| Framing mix, across everything | 27 direct searches, 23 problem descriptions | show whether recommendations depend on the question style |
Four examples from this set:
«Which Swiss agency specializes in GEO, i.e. Generative Engine Optimization?»
«We are a fiduciary firm and ChatGPT never recommends us, although competitors appear. Who can help us?»
«We rank first on Google but do not appear in AI answers. Which agency can help?»
«ChatGPT spreads false statements about our company. Who helps us correct that?»
The scope is 50 prompts per service area. More questions do not make the measurement better if they probe the same buying situation twice: 50 well-distributed prompts cover the relevant situations and remain measurable daily.
How do you spot a weak prompt set?
By four patterns:
Wishful thinking instead of buyer language
Prompts full of internal product names and jargon no customer uses. The questions must reflect what real buyers ask, not what would be asked internally.
Category questions only
A set that merely varies «Which providers for X are there?» measures one slice. Buyers frequently describe situations instead of naming categories.
Duplicate prompts
Questions that differ only in wording measure the same buying situation several times and distort the rate. Every prompt needs its own buying situation.
Constant changes
Whoever keeps rebuilding the set starts a new baseline every week and never sees a trend.
A prompt set is a measuring instrument. You calibrate it once, carefully, instead of reinventing it weekly.
How do buyer prompts feed into the measurement?
Every prompt runs daily on ChatGPT, Claude, Perplexity, Gemini, Copilot, and Google AI Overviews, without a login and without memory. Evaluation happens weekly as the recommendation rate: the share of answers in which the company is recommended, compared against the same competitors. The full measurement protocol is described in How do you measure AI visibility reliably?
The prompt set is half the measurement methodology: the cleanest measurement is useless if the wrong questions are asked.