# How to Measure AI Visibility for Property Ops Software: Notes from the First Run

> Mention rate, position, consistency and the source map, and why one ChatGPT run is not a result. Field notes from the first AI Shortlist Index run.

Canonical: https://webpossible.com/blog/how-to-measure-ai-visibility-for-property-ops-software/
Source: https://webpossible.com/blog/how-to-measure-ai-visibility-for-property-ops-software/
Format: Markdown version for AI agents. The canonical HTML page is at the source URL above.
Last verified: 2026-09-03

---

Blog

# How to Measure AI Visibility for Property Ops Software: Notes from the First Run

By Ryan York · last verified 2026-09-03

Somewhere in the last two years the question a property manager types changed shape. It used to be “maintenance coordination software” into a search box, then a scroll through ten links and three ads. Now it is a full sentence into ChatGPT or Perplexity: “What’s the best maintenance coordination software for a property management company with around 300 doors?” The answer comes back as two or three vendors. Nobody clicked anything to get there, and no analytics tool on the vendor side saw it happen. These are notes on how we measure that moment, written after the first run of the AI Shortlist Index.

## What changed when PMs started asking instead of searching

A search result is a list of pages. An answer is a list of vendors. That difference is the whole reason old measurement stops working. Position tracking tells you where your page sits for a search term and nothing about whether an engine that read that page decided to include you in what a 300-door operator sees. Sessions tell you who arrived; the buyer who took the shortlist straight to a demo request never arrived through you at all.

The second change is that the question carries context now. A PM does not ask for “maintenance software.” She asks for maintenance software for 300 scattered single-family doors on AppFolio, or for a 2,000-unit multifamily portfolio on Yardi with an in-house tech team. The engine answers that specific question, and the shortlist can change with every clause. Measuring one generic query tells you almost nothing about the questions buyers ask.

The third change is that the answer is not stable. Ask twice and you may get two different lists. That is inconvenient for anyone who wants a scoreboard, and it is the reason a single check is worthless as evidence.

## The four numbers we record

The [Index methodology](/ai-shortlist-index/methodology/) defines the metrics in full. The short version is four figures per vendor, per category, per edition.

Mention rate. Of the runs where the vendor could have appeared, the share where it did. Computed per engine and blended across the engines that returned answers. Prompts that put a vendor in the question (“alternatives to X”) do not count toward X’s rate, because that would hand X a mention for free.

Position. When the vendor appeared, where in the answer it came, counting the first vendor in the answer as 1. A vendor that always appears fourth in a five-vendor answer has a different problem from one that appears first half the time and not at all the rest.

Consistency. Across repeated runs of the same prompt on the same engine, how often the runs agreed with each other about whether the vendor appeared. This is the number that keeps everyone honest. A low consistency figure means the engine is flipping, and a flip is not a trend.

The source map. Every URL read or pointed to while the answer was being written, normalized and classified by page type: vendor site, review platform, listicle, forum, trade publication, directory. This is the one that tells a vendor what to do next, because it shows which pages the engine trusted when it built the shortlist.

We also record engine spread, meaning how many of the four engines included a vendor at all. Framing is stored too (recommended, neutral or cautioned, from cue words near the first mention). The four above carry the weight.

## Why one run is not a result

Most do-it-yourself measurement goes wrong at the same spot. Someone on the marketing team asks ChatGPT the question once, screenshots the answer, and takes it to the Monday meeting as “we’re not in there.” Maybe they are not. Or maybe that run went one way and the next one goes the other.

The Index protocol is three runs per prompt per engine, each in a fresh session, spread across at least two days. Twenty-five prompts per category. Four engines. That is 300 responses per category before anyone reads a number, and the consistency metric is published next to the rank so that a rank built on a coin flip is visible as one.

One run gives you an observation. Three runs on separate days give you a measurement with an error bar you can see. The difference matters most for the vendors in the middle of the table, where a single lucky mention moves a rank and a single unlucky omission drops one.

## What the September 2026 preview showed

The preview edition is exactly that, a preview. It ran one engine, one prompt, one time, in one category. On 3 September 2026, through the OpenAI API with web search enabled, which is how the Index captured ChatGPT before it moved to the app, we asked ChatGPT “What’s the best maintenance coordination software for a property management company with around 300 doors?” and it named Property Meld first and Latchel second. The domain ChatGPT cited most often while building that answer was reddit.com.

That is the entire set of verified facts. Two vendors, one order, one most-cited domain. Everything else on the preview pages is either a null (Perplexity, Gemini and Google AI Overviews recorded no runs, so they show as missing rather than as zeros) or a description of what the full edition will hold.

What the single run suggests (suggests, not shows) is still worth writing down. If reddit.com stays the most-cited domain in maintenance coordination across the full run set, then a vendor’s presence in property-manager threads may matter more than its own feature pages for this category. And if the same two vendors hold the top two slots across three runs and four engines, that would be an unusually stable category; our expectation is that they will not, and that the consistency figure will show movement. Both are hypotheses. The October edition will test them.

The preview row data is on the [maintenance and vendor network category page](/ai-shortlist-index/maintenance-vendor-network-software/), with the degraded-engine flags left in place.

## How to do this yourself, badly and then well

Badly is fine to start. Write down the ten questions your buyers ask right before they book a demo, in their words, with a portfolio size and a PMS in the sentence. Run each one on ChatGPT, Perplexity and Gemini. Record who appeared, in what order, and which sources showed up under the answer. Do it again in three days. Compare.

That gets you a rough mention rate and a first look at consistency, and it will already tell you something the rank tracker cannot. Where it falls apart is scale and repeatability. You will not run 75 prompts three times a month by hand. You will start rephrasing prompts that gave bad answers without noticing you are doing it. And you will have no record anyone can audit.

Well means fixing those three things: a frozen prompt set, scheduled runs through documented APIs, and raw responses stored before anyone computes anything. That is the Index, and the [AI visibility measurement guide](/ai-visibility/) walks through each part in detail.

Once you know whether you appeared, the next question is why, and the answer usually sits in the pages the engine chose to read. We wrote up [what an engine pulls from a vendor page](/blog/how-ai-engines-read-a-vendor-page/) separately. More field notes will land on the [blog](/blog/) as editions publish.

Part of [the AI visibility measurement pillar](/ai-visibility/).
