Blog
Building a Prompt Set That Measures What Buyers Ask
The prompt set is the instrument. Every number the Index publishes is a count of what four engines said in reply to 25 fixed questions per category, so the questions decide what the numbers mean, and a badly built set measures something with great precision that no buyer asked. This post goes through how Prompt Set v1.0 was designed, the rules each prompt follows, how a vendor builds a set for its own category, and the four mistakes that turn a set into a mirror.
Six doors into a category
A buyer does not arrive at a software category through one question. Some open with the shortlist question. Some already run a product and want out of it. Some describe their portfolio. Some describe their PMS. Some describe the pain and have no category word for it yet. A set that is all “what’s the best” prompts measures one door and reports it as the building. The Index set spreads 25 prompts per category across six types.
| Type | Count | What it measures |
|---|---|---|
| best | 7 | the open shortlist question, varied by portfolio context |
| alternatives | 5 | who gets named when a buyer starts from an incumbent |
| segment | 4 | whether recommendations change by property type or portfolio size |
| integration | 3 | whether the PMS named changes the answer |
| problem | 4 | who gets named when the buyer describes the pain, not the category |
| comparison | 2 | head-to-head between the two most-asked-about incumbents |
Twenty-five prompts per category in six types: seven open shortlist questions, five alternatives-to questions, four segment questions, three PMS integration questions, four problem descriptions with no product words, and two head-to-head comparisons.
The problem-led prompts are the ones vendors most often leave out of their own checks and the ones we would argue for first. A property manager drowning in after-hours calls does not type “maintenance coordination software”. They type the drowning. If an engine only names you once the buyer has learned your category’s vocabulary, you are invisible at the moment the budget gets created.
How each prompt is written
The prompts are written the way a PM asks, in first person, in one or two sentences, with a real portfolio detail in most of them. Two from the maintenance and vendor network set show the range. Prompt mvn-best-01 reads: “What’s the best maintenance coordination software for a property management company with around 300 doors?” Prompt mvn-best-02 reads: “We manage about 1,200 units across several garden-style apartment communities. Which maintenance software do operators our size use to run work orders and vendors?” Same intent, different door count, different property type, different phrasing of the ask.
Each prompt carries one question, so the answer can be scored as who was named and in what order. Stacked asks produce answers where the second half undoes the first. The seven best prompts vary door count, property type, what the PM already runs and what they are optimizing for, rather than being one sentence with a different number in it.
Vendors appear in prompt text only where a buyer would put them. Alternatives and comparison prompts name real incumbents because a buyer switching from one names it. Best, segment and problem prompts never name a vendor. Integration prompts name the PMS, AppFolio or Buildium or Rent Manager or Yardi, and never a vendor in the category being measured.
No prompt leads. None describes a feature only one vendor has, none mentions a pricing tier, and none names webpossible.com, the Index, or any source it hopes the engine will read. All prompts are US and unpersonalized, with no city unless vendor coverage in a market is the point of the question and no “near me”.
Seed vendors and why their mentions do not count
A prompt that says “alternatives to X” will produce X, because the engine has to name the incumbent to discuss leaving it; the methodology page states this as the reason for the exclusion. If that mention counted, every alternatives prompt would hand its seed vendor a free run, and five of 25 prompts would be a subsidy.
So each prompt lists its seed vendors, meaning the vendor ids in its text. The extractor records the seed vendor’s appearance, stores it with a seeded flag so the exclusion is auditable, and the metrics step does not count it toward that vendor’s mention rate, position score, consistency or rank on that prompt. The vendor still earns credit on the 20 prompts that do not name it.
Versioning, and why text is frozen
Prompts are frozen per edition and versioned as a set. The current set is v1.0, approved 2026-09-02. Changing a prompt’s text does not edit the prompt. It creates a new id and a changelog entry, and the old prompt stays in the file marked retired with the version that retired it. Each prompt carries a permanent id in the form of category, type and number, so a third party can cite one prompt exactly, and the whole set is cited as “AI Shortlist Index Prompt Set v1.0, WebPossible.”
The reason is survivorship. The moment a prompt can be quietly rephrased, the prompts that gave awkward answers get rephrased and the ones that gave flattering answers stay. Six months later the set measures what the editor wanted to see. Freezing the text and logging every change makes that drift visible, and the public prompt library is where anyone can check it.
Building one for your own category
A vendor does not need 25 prompts or four engines to build a defensible set. It needs the shape.
Start with the doors. Write down how buyers arrive: the open shortlist question, the switch from an incumbent, the portfolio description, the PMS they run, the pain with no product word attached. If your sales calls and support tickets are logged, the first sentence a prospect says is usually one of these five, in their words. Forum threads are the other source. In the one run in the September 2026 preview, 38 of the 124 sources ChatGPT cited were on reddit.com. The threads it read are also where buyers write their questions down in public.
Then write 15 to 25 prompts in a PM’s voice, with a door count or a property type in most of them. Tag each one with its type, its segment, the PMS it names if any, and the vendors it names if any. Give each a permanent id. Freeze the text and record the version and the date.
Then run the set the way the methodology runs it: three times per prompt per engine, in fresh sessions, across at least two days, with every raw answer kept. The set is the instrument and the protocol is how you read it.
Four mistakes
Leading prompts are the first. A prompt that mentions the feature only your product has will produce your product, and the number that comes back measures the prompt. If a buyer would not describe that feature before they knew you existed, it does not belong in the text.
Brand in the prompt is the second, and it is the same mistake in plainer form. “Is Acme good for maintenance?” is a question about a brand, and any mention it produces is seeded. If you want to know whether you are on the shortlist, the prompt cannot contain your name. Where a buyer would name a competitor, name the competitor and exclude that mention from the competitor’s count.
Too few prompts is the third. With three prompts and three runs, one flip on one engine moves a mention rate by eleven points. With one prompt, the rate is a coin. Twenty-five is a choice; eight is a floor.
No versioning is the fourth, and it is the quiet one. A set that lives in a shared doc gets edited by whoever ran it last, and the edits go toward the answers people liked. Freeze the text, give each prompt an id, and treat a rewording as a new prompt. That is the difference between a set and a screenshot with extra steps.
The rest of the engine-by-engine work, once the set exists, is laid out in the AI search optimization overview, and the field notes continue on the blog.
Part of the methodology pillar.