AI engines state things about brands with total confidence, and some of those things are wrong. A discontinued product recommended as current, a price that's two years stale, a feature you never built. Answer Accuracy exists to catch these hallucinations in the answers real buyers are seeing, before customers act on them.
The section has three tabs: Claims to review, Being watched and Dismissed. The fact sheet they are checked against is in Brand Settings, under Facts, and a link beside the tabs opens it.
The fact sheet
Accuracy checking is only as good as the truth it checks against, and yours is the fact sheet: a list of short, checkable statements about your brand. What it sells, what it costs, what's current, what's discontinued, who founded it, where it operates.
brandflare drafts the sheet when your brand is set up, and those drafted facts are used from the first check. Your job is to review them:
- Verify the facts that are right. You can verify everything at once.
- Reject the ones that aren't, or correct them.
- Add your own, several at a time, one per line.
Filter the sheet by To review, Verified or Rejected. Facts are grouped by topic, the same topics as your claims, so you can see which subjects the sheet covers. Each fact shows where it came from and how many answers contradict it, which is a quick way to find the fact AI most often gets wrong.
Drafted and verified facts are both checked against answers, so the practical difference is confidence: a claim raised against a verified fact is one you can act on straight away.
Keep the sheet current. When your pricing changes or a product ships or is retired, update the fact sheet first. Otherwise the judge is checking answers against yesterday's truth.
How checking works
During every check, after each engine answers each question, an LLM-as-judge reads the full answer and compares everything it says about your brand with your fact sheet. When a statement contradicts a fact, or asserts something false about your brand outright, it becomes a claim to review, recording:
- Who said it: which engine, in answer to which question, on which date.
- The claim: what the engine said, quoted from the answer.
- Why it's wrong, and what's true: the fact it contradicts.
- A severity: high, medium or low. A slightly outdated founding year is not the same as "this product has been discontinued".
Every claim is double-checked before it's kept, and every claim opens the full answer it came from.
The judge also knows which similarly named organisations are not your brand, so a free zone, the district it sits in and the energy company that shares its name aren't credited with each other's facts.
Only false statements, never omissions
This is the design decision worth understanding. brandflare raises only affirmative false claims: things an engine positively states about your brand that are wrong. It never raises omissions.
If ChatGPT fails to mention your enterprise tier, or leaves you out of a list entirely, that is a visibility problem. It shows up in your Visibility, on Questions & Prompts, and in your score. It is not an accuracy problem, and treating it as one would bury real hallucinations under an endless list of things AI didn't say.
Working through claims
Claims to review opens on your topics: the subjects AI gets wrong about you, such as pricing, products or leadership, worst first. Each topic card shows how many claims wait for review and how serious they are, which engines say them, how many facts on your sheet cover the topic, and whether the topic had more or fewer claims at the last check than at the one before. A topic with no facts is marked, because answers about it can only be checked against what's obviously false. A topic that appears for the first time is marked New until someone opens it.
Topics are set for you, and a new subject gets its own topic when nothing else fits it. Open a topic to review its claims one at a time, worst first, or review every claim in one queue. Coloured pills filter by severity. The card says who says it, in how many of that check's answers, when it was first said, why it's wrong and what's true, and opens the answer it came from. The same claim said in several answers is one decision, not several.
There are five things you can say about it. Three accept that AI has it wrong, and all three keep it counting against your Accuracy, because the engines are still saying it:
- Watch. It leaves the queue and is followed at every check, until the engines stop saying it.
- Fix it. The same, and it becomes a recommendation carrying the claim, what's true, the engines saying it and the sites those answers cited.
- Dismiss. You've seen it and don't need it in the queue. It isn't followed, and dismissing doesn't improve your score.
Two say the flag itself was wrong, and both stop it counting:
- Add to fact sheet. The claim is true after all. It goes on your fact sheet as a verified fact, so later answers are checked against it.
- Not about us. The claim is nothing to do with your brand, which also tells us our judge got it wrong.
Every decision can be undone, from the message that confirms it or from the Dismissed tab. A claim you have decided doesn't come back into the queue at the next check.
Claims to review holds what AI is still saying. A claim that none of your last three checks heard leaves the queue by itself and moves to Dismissed as No longer said. If an engine says it again, it comes straight back.
Being watched holds the claims you chose to follow, each marked as still said or stopped, with when it was last heard. It's where you see whether a fix worked. Dismissed holds the rest, with what each decision did to your score, and a way back to the queue.
For each claim, the useful questions are:
- Is it real? Judges are strong but not infallible, and a stale fact in your sheet can produce a false alarm. Read the answer, check the fact.
- Is it recurring? A claim that appears once may be noise. The same false claim across several checks or engines points at a stale or wrong source the engines keep drawing on.
- Where did it come from? Persistent hallucinations usually trace to an incorrect page the engines retrieve or have memorised. Sources shows which sites the answers cite. Finding and fixing that source is the actual remedy: the playbook is in finding and fixing AI hallucinations, and realistic expectations are in can I correct what AI says about my brand?
Because checks keep running, you don't have to guess whether a fix worked: Being watched tells you when a claim stops appearing.
Editors, admins and owners can review claims and edit the fact sheet; viewers can read both.
Where accuracy shows up elsewhere
- Accuracy issues on your overview and dashboard counts the claims waiting for review, split by severity, and the sidebar shows the same count beside each brand.
- Accuracy is one of the five factors in the Flare score. Claims count against it while AI is still saying them, whether they're open, watched or dismissed. Only the two that say the flag was wrong stop counting.
- Alerts: a new false claim reaches your inbox the same day as the check that found it, with what was said and what is true.
- Export Report lists the false claims that matter most.
What Answer Accuracy is not
It isn't sentiment policing: an engine being lukewarm about your brand is not a false claim. It isn't exhaustive: the judge checks the answers your questions draw out, not every conversation happening on every engine. And it isn't an edit button for AI. brandflare finds the false claims and gives you the receipts; correcting the sources that feed them is work the product helps you target, verify and confirm.