Audit what AI assistants say about you
People ask an assistant before they ask you. What it tells them — which brands it recommends, which figure it quotes for your price or your delivery time, which pages it read that from — is a fact about your market that nobody in the company sees. Trying it by hand gives one answer on one day; the next person gets another.
The AI answer audit Job makes that a measurement. It asks a fixed set of questions of the assistants you choose, keeps every answer word for word with the pages it cited, has a model read each answer against a reference sheet you confirmed, and reports what you then confirmed as wrong — month after month, so the second report is a comparison.
What you fill in
Section titled “What you fill in”Installing the Job creates eight tables and four flows. Four of the tables are yours to fill, by hand or by import:
| Table | What goes in it |
|---|---|
| Brands | Your brand and the ones people compare you with. name is how the brand reads everywhere else; also_known_as holds the nicknames and spellings an answer may use; own_sites the brand’s own domains, comma-separated. Tick Tracked on the brands to count. |
| Questions | The questions, word for word, each with a short question_id. An open question compares brands (“Which phone plan is best for a student?”). A fact question asks one brand’s figure and names the brand and a fact_key, so the answer can be checked. Search evidence records whether buyers really ask it (see below). Tick Active on the ones to ask. |
| Assistants | Who gets asked. A row of kind model names a vetted model string (or leaves it blank for the workspace’s default route), whether it may search the web, a search country, and how many runs per question. A row of kind app is a consumer app whose answers you paste into the Answer log yourself. |
| Reference sheet | What the source says, one row per brand and fact_key: a plain statement with the figure, the source page, and a status. Only rows set to confirmed are used to judge an answer. |
The question text is sent exactly as written, with nothing added — if you want a persona, write the question in that voice and note the persona in its column.
Check that the questions are ones buyers ask
Section titled “Check that the questions are ones buyers ask”A panel of questions you or your agency chose is a judgment. Before you quote a result, give each question evidence that buyers really ask it, and record it in four columns of Questions: Search evidence (searched as asked, searched in other words or no evidence), the Search phrase buyers actually type when it differs from your question, Monthly searches when you know them, and the Evidence source. A blank Search evidence means not checked yet.
Where the evidence comes from, strongest first:
- Your Google Search Console — the real queries that bring people to your site, with impressions and clicks by country. The best evidence there is, and free.
- Your helpdesk or sales inbox — what customers ask in their own words, and the competitors they name.
- Google Keyword Planner and Google Trends — search volume for a phrase, and how it moves.
- Google autocomplete — type the start of what a buyer would ask and see whether Google completes it. It shows that a phrase is searched and the words people use, not how often.
When the evidence says buyers use other words — “upgrades” or “mods” where you wrote “improve” — reword the question the way they search before the next wave.
The four flows
Section titled “The four flows”Set up the board — run once. It creates the AI answer audit dashboard with its charts and the saved view To confirm on Findings. Running it again while the board exists changes nothing.
Ask the questions — on the first of the month, or whenever you press Run. For every active question, every active model assistant and every one of its runs, it sends the question through Ask AI with that row’s model, web search and country, and adds one row to the Answer log: the answer, the model that gave it, the queries it searched and the pages it cited. Every cited page is also a row in Sources, with its site and — when the host is one of a brand’s own_sites — whose site it is.
Two things to know about this flow:
- A failed call is skipped, not fatal. A rate limit or a provider error leaves no row and the run carries on with the next question.
- It only asks what is missing. An answer’s id is
<month>/<question id>/<assistant>/<run>. Run the flow again in the same month and the answers that exist are left alone; only the gaps — a failed call, a question you added, an assistant you switched on — are asked. A wave of 150 questions, three assistants and three runs is 1 350 searched calls and takes hours; if the run stops part-way, run it again and it continues.
When it finishes it starts Read the answers.
Read the answers — for every Answer log row with no Read at, one Ask AI call on the default route reads the answer and reports, as typed items: which listed brands it recommends, which one it puts first, and — for a fact question — whether it gives a concrete figure and which statements contradict the confirmed reference rows for that brand and fact. The reading follows fixed rules: judge only against the reference rows given; a statement the reference does not cover is not a contradiction; fine print worth a trivial amount, and one of two price tiers given as the usual one, are not contradictions; never guess a reference.
It then writes:
- Recommendations — for an open question, one row per tracked brand:
recommendedyes or no,first_choiceyes or no. The rows that say no are the base a share is reported against. - Findings — for a fact question, one row per contradicting statement, with status to confirm. The reference text and the source page are copied from the Reference sheet row the flag names, never from the model; a flag that names no existing row is dropped.
- On the answer itself:
first_choice,figure_given, how many statements wereflagged, andread_at. The answer text is never changed.
A hand-pasted answer from a consumer app is read like any other: give it an answer_id, the question_id and the app’s name as assistant, and the reading fills in the question’s kind, brand and fact from the Questions table. Paste the pages the app cited into its citations cell, one per line, and the reading turns them into Sources rows under the same site and own-site rules as a model answer’s, so the board compares where the apps and the models get their pages. Tick ad_shown if an ad appeared, leave it blank when you do not know. It reads up to 1 000 answers per run and ends with one notification: how many statements wait for confirmation.
You confirm or dismiss. Open the Findings table, view To confirm, and set each statement’s status to confirmed or dismissed. A model flags; a person decides. Nothing counts as a wrong figure until you say so.
Build the report — sets a verdict on every read fact answer: no figure given when the answer gave no figure, wrong figure confirmed when at least one of its findings is confirmed, otherwise no wrong figure found. A dismissed finding counts for nothing. It then snapshots the board into a branded PDF, lists every confirmed finding — what the answer said, what the reference says, the source page — stores the PDF and sends you the link.
What the report counts
Section titled “What the report counts”- Recommendations by brand — every open-question answer that was read, once per tracked brand; the coloured part is the answers that recommend the brand. The bar’s height is the base.
- First choices by brand — the answers that put the brand first. An answer with no clear first choice counts for nobody.
- Fact answers by assistant and verdict — every fact answer, in exactly one of four buckets: the three verdicts, or not judged yet.
- Cited pages by site and by whose site — every page an answer cited, by the site it is on and by the brand whose own site it is; the dash is every third-party site.
- Answers logged by month — how many answers each wave holds.
Every share comes with the number it is a share of. Three of four is not seventy-five percent.
What it does not do
Section titled “What it does not do”- It does not promise a better ranking, a better score or more recommendations. It records what the assistants said and how that moves from month to month.
- An answer obtained through a model’s interface is close to, not the same as, what the consumer app shows the same person; the apps add their own instructions and tools. When that difference matters, paste the app’s own answer into the Answer log.
- Only findings a person marked confirmed count as a wrong figure — in every chart and in the report.
- A fact with no confirmed Reference sheet row is never called wrong, whatever the answer says.
- Every searched answer reads the pages it finds, and every reading is a call. The AI call log shows each one.
Where to go next
Section titled “Where to go next”- Ask AI with web search and keep its sources — the step the asking flow is built on.
- Chart a table and build a dashboard — add a chart of your own to the board.
- Check whether AI search can read your site — the other half of the question: what the assistants can read from you.
- Set up and activate an installed Job — arm the monthly schedule.