Blog

Why We Publish the Prompts Behind the Index

Every prompt the Index sends to an engine is public, with its id, its type, its tags, the vendors it names and the version that added it. So is the capture method for each engine, the run schedule, the extraction rules and the metric definitions. So is the dispute process. A ranking whose inputs are secret asks to be trusted; a ranking whose inputs are published asks to be checked. This post covers why we chose the second, what it costs, and how to reproduce a row with your own API key.

What is public

The prompt library carries all 25 prompts per category, each with a permanent id like mvn-best-01, its type, its segment and PMS tags, its seed vendors and the version that added or retired it. The set is cited as “AI Shortlist Index Prompt Set v1.0, WebPossible.” Changing a prompt’s text creates a new id and a changelog entry, and the old text stays in the file marked retired. The design rules behind the set are written up separately.

The methodology page states the protocol. For ChatGPT that means the app at chatgpt.com: one temporary chat per prompt, signed in to an account used only for the Index, with memory, custom instructions and chat-history search off, the model left at the default and the decision to search left to ChatGPT. Perplexity is the sonar model through its chat completions API. Gemini is the Gemini API with Google Search grounding. Google AI Overviews are captured from US desktop results through a licensed data provider. Three runs per prompt per engine, each in a fresh session, spread across at least two days. Every response written once to a file that is never edited, with the timestamp, the session id, every cited URL, the complete raw capture, and the model identifier where the engine gives one. The ChatGPT app gives none, and reports no token count or price either, so those fields are empty on its runs.

The extraction rules are public too: how a name is matched, how position is counted, how a seeded mention is recorded but excluded, and the hand-labeled accuracy gate the extractor has to pass before metrics publish. And the policy page carries the firewall, the re-check process and the record of disputes.

Reproducibility is the reason

A number that cannot be reproduced is an opinion with decimals. The Index publishes a rank, a mention rate, a position score and a consistency figure per vendor per category, and each one is a claim about what an engine said. The only way for that claim to be more than our word is for someone else to be able to send the same prompts, through the same method, and get an answer distribution in the same neighborhood.

That requires the exact prompt text, because “best maintenance software” and “maintenance coordination software for a company with around 300 doors” are different questions with different answers. It requires the capture parameters, because an API call with web search and a consumer session with memory on are different instruments, a gap the methodology page names and measures with a monthly consumer-UI sample. It requires the run counts and dates, because one run is a draw and three runs on two days is a small sample with a stated shape. And it requires the extraction rules, because whether “Meld” in a sentence counts as Property Meld is a decision that moves a rate.

Publish all four and a vendor, a competitor or a reporter can rerun a row. Withhold any one and the row is a screenshot with a spreadsheet behind it.

Disputes on the record

Any vendor can request one re-run of its category’s prompt set per edition. The re-check follows the same protocol, three runs on separate days on all engines, and is published regardless of outcome, next to the original. Entity corrections, meaning name, aliases, category membership and website, are handled in the edition after verification against a source URL. We respond to disputes with data, on the policy page, where anyone can read both sides.

None of that works with secret prompts. A vendor that believes its row is wrong needs to see the questions that produced it, and a reader deciding whom to believe needs to see the same questions. A dispute over a hidden prompt set resolves to whichever party is louder.

The firewall needs the prompts to be public

The line on every Index and practice page reads: “Paying WebPossible never affects Index placement. The Index measures what AI engines say. Shortlist clients buy the work that earns a position, not the position.”

That sentence is a claim about our own conduct, and a claim about conduct is worth what it can be checked against. If the prompts were private, the easiest way to move a client up would be to write prompts the client happens to answer well, and nobody outside could tell. With the prompts public and frozen per version, that lever is gone. A client’s competitor can run the same 25 prompts through the same API and see the same distribution we see. If our published rate and theirs diverge beyond what three runs on two days should produce, one of us has a problem, and the raw files exist to find out which.

The same logic applies to the practice. Shortlist sells the work that earns a position: the pricing page with a price on it, the integrations page that names the PMS, the sources an engine reads. A client can see, in the prompt library, exactly what question the work is meant to answer. That is a better sales conversation than a promise.

What publishing costs

Gaming is the obvious cost. A vendor can read prompt mvn-best-01 and build a page titled after it, and a marketing team can work down the list. We accept that for two reasons.

The first is that a page which answers a buyer’s question with facts is what should be cited. If reading our prompt set causes a vendor to publish its price, its integration list and its coverage by segment, the buyer wins and the Index has done something useful on the way to measuring it. The second is that the set is built to resist the cheap version of gaming. Twenty-five prompts across six types means a page targeting one prompt moves one twenty-fifth of the rate. The problem-led prompts contain no product words to target. Seeded mentions are excluded, so a vendor cannot ride its own name in an alternatives prompt. And prompt changes go through a versioned changelog, so an edition’s text is fixed before anyone sees the results it produces.

Copying is the other cost. An agency can take the set, run it and publish a rival index. That is fine, and the citation line exists for it. A copied set run through the same protocol should reproduce our numbers within the variance three runs allow. If it does, the Index has a second witness. If it does not, the discrepancy is worth more than the exclusivity would have been.

The real cost is on our side. We cannot quietly rephrase a prompt that gives a strange answer, and we cannot drop an engine because it was unkind to a category. Every such change is a new version, a changelog entry and a public explanation.

How to reproduce a row

Take the 25 prompts for the category from the prompt library, with their ids and seed vendors. For ChatGPT, open a temporary chat at chatgpt.com on an account you keep for this and nothing else, with memory, custom instructions and chat-history search off, and paste the prompt with the model left at the default. Send each prompt three times, each in its own chat with no shared state, and spread the three across at least two days. Store the whole capture each time: the answer text, every link in it and the time you collected it. The app will not give you a model id, a token count or a price, and it will not give you the pages the search opened but did not link.

Then apply the extraction rules. Match each vendor’s canonical name or listed alias in the answer text, case-sensitively for names that are also English words. Record the ordinal of each vendor’s first mention among all vendors named. Record any seeded vendor’s mention with the seeded flag and leave it out of that prompt’s counts. Compute the mention rate as runs in which the vendor was named divided by runs that returned an answer, over the prompts where the vendor is not seeded. Rank by that rate, with ties broken by position score.

Then compare. Our one run of mvn-best-01 on 2026-09-03 named Property Meld first and Latchel second, and cited 124 sources with reddit.com the top domain at 38. That is one draw, and the row on the category page says so on its face. Run the same prompt three times over two days and you have a better estimate than we have published. If yours disagrees with ours once the October edition lands, say so, and we will put both on the record.

The Index itself, with the current edition and every row’s uncertainty shown, is at the AI Shortlist Index, and the field notes continue on the blog.

Part of the AI Shortlist Index.