How to Fact-Check AI-Generated Content Before You Publish It
A model can write a wrong number with the exact same confidence as a right one. That's not a bug you can prompt your way around, it's just what generation is. Here's a process for catching it before a reader does.
Most advice about AI writing focuses on making it sound less robotic. Less discussed, and arguably more important, is the fact that AI-generated text can be confidently, fluently, and completely wrong, and there's no visual difference between a sentence that's accurate and one that isn't. A hallucinated statistic reads exactly like a real one. A made up quote is formatted exactly like a sourced one. The writing gives you no signal at all about which sentences to trust, so you have to build that signal yourself, and that means actually checking.
Why this happens, briefly
A language model generates text by predicting the next most plausible word given everything that came before, it's not looking anything up or checking a database of facts as it writes. When you ask it a factual question, it's producing the sequence of words that's statistically most consistent with how that kind of answer usually looks, which is a fundamentally different process than retrieving a verified fact. Most of the time this produces correct information anyway, because most factual statements around most topics do actually follow predictable, well-represented patterns in the model's training data. But nothing about the process includes an internal check of "is this specific number actually true," so when the pattern-matching goes wrong, it fails in the same fluent, confident register as when it goes right.
This is worth sitting with for a second, because it explains why hallucinations feel so surprising when you catch one. The sentence right before the wrong fact was accurate. The sentence right after it goes back to being accurate. There's no stylistic seam, no dip in confidence, nothing in the writing itself that flags the one sentence in the middle as different from its neighbors. A person who doesn't know an answer usually signals it somehow, hedging, hesitating, admitting uncertainty. A model under the same circumstances often just picks the most statistically plausible-sounding answer and states it the same way it states everything else.
Where the risk actually concentrates
Not every sentence carries equal risk, some categories are far more likely to contain an invented detail than others. Worth learning to spot these on sight, so you know where to spend your checking time instead of re-verifying an entire piece line by line.
- Specific numbers and statistics. "62% of businesses report..." is exactly the kind of oddly precise claim that's easy for a model to generate and easy for a reader to accept without questioning. If you didn't feed the number in yourself, verify it before it goes anywhere near a final draft.
- Quotes attributed to real people. Models can produce a quote that sounds exactly like something a person would say, complete with plausible phrasing, and it can be entirely invented. Always trace a quote back to an actual source before using it, never just because it sounds right.
- Citations and sources. A fabricated paper title, journal name, or URL can look completely legitimate at a glance. Click through. If the source doesn't resolve to something real, or the real source doesn't say what's claimed, the citation goes.
- Recent events. Anything past a model's training cutoff is guesswork dressed as memory, and even within the training window, breaking or fast-moving stories are exactly where errors cluster.
- Anything oddly specific about an obscure topic. A model asked about a niche subject will still answer fluently even with thin training data behind it, and confidence doesn't scale down just because the underlying knowledge does.
Phrasing patterns worth extra suspicion
A few specific ways of phrasing a claim show up disproportionately often around fabricated details, worth training your eye to catch on a first read, before you even start the formal checking process.
- Vague attribution. "Studies show," "research indicates," "experts agree," with no actual study, researcher, or source named. This phrasing borrows the credibility of research without pointing at anything you could go check. Treat it as an unverified claim by default, because that's exactly what it is.
- Suspiciously precise numbers with no source. "73.4% of companies" sounds more credible than "most companies" specifically because of the false precision, and that's backwards. A number that specific should come with a source attached. If it doesn't, the precision is a warning sign, not a credibility signal.
- Round numbers presented as exact. The opposite pattern also shows up, suspiciously tidy figures like "10 times faster" or "doubled in five years" dressed up as if they were measured rather than estimated. Ask where the number actually came from.
- Confident claims about niche or highly specific subtopics. The more obscure the claim, the less likely the model had solid training data to draw from, and the tone gives you no indication of that gap. Confidence and accuracy are simply not the same signal, and nothing in the model's phrasing tracks the difference for you.
None of these patterns prove a claim is wrong on their own. They're a signal to go check, the same way a slightly-too-good-to-be-true price is a signal to read the fine print, not proof of a scam by itself.
A workflow that actually holds up
1. Pull out every checkable claim into a list
Read through once and list every number, name, date, quote, and cited source separately from the surrounding prose. This turns a vague "does this feel right" skim into a concrete checklist, which is a much easier and more thorough task, you're not trying to hold the whole piece's credibility in your head at once anymore.
2. Verify each one independently
Search for the number, quote, or event on its own, not as part of the sentence it came in. Models are decent at wrapping a fabricated detail in accurate-sounding surrounding context, so checking the sentence as a whole can feel more convincing than it should. Isolate the fact and check just that.
3. Require two independent sources for anything load-bearing
If a claim is doing real work in your piece, the headline stat, the thing your whole argument leans on, don't settle for one source that happens to confirm it. Two truly independent sources saying the same thing is a real signal. A model's fluent paraphrase of a single source, dressed up to look like confirmation, still only counts as one source underneath the disguise.
4. Click through every citation
Don't just check that a source is cited, check that it exists, that the link resolves, and that it actually says what's being attributed to it. A fabricated citation to a real-sounding journal is one of the more common and more embarrassing failure modes, and it takes thirty seconds to catch.
5. When you can't verify it, cut it
This is the rule that actually matters most: an unverifiable claim is a liability sitting in your draft, not a neutral placeholder you can fix later. If you can't confirm something within a reasonable amount of effort, remove it or rewrite the sentence to not depend on it. A vaguer true sentence beats a specific false one every time.
Fact-checking and rewriting are two different jobs, and it's worth keeping them separate. Once a draft is verified, Unzap.app is built for the second job: paste it in, pick a style, and get a version back that reads like it was actually written by someone, not generated by something.
Try Unzap.appWhat this isn't a substitute for
Fact-checking is not the same job as making writing sound more human, and it's not the same as running a piece through an AI detector either, checking for style tells or a watermark can't tell you anything about whether the content is true. A perfectly natural-sounding, undetectable paragraph can still contain a made up statistic, and a paragraph that reads a little stiff can be completely accurate. Treat verification and voice as two separate passes over a draft, because they catch two completely different kinds of problems.
The actual habit
None of this needs to slow you down much once its a habit rather than an afterthought. Most drafts only have a handful of actually checkable claims buried in them, the rest is framing, transitions, and structure that doesn't carry factual risk at all. Find the five or six sentences doing the actual factual work, verify those properly, and you've covered nearly all of your real exposure. The goal was never to distrust everything a model writes. It's to know exactly which sentences deserve a second look, and to actually give it to them before anyone else reads the piece.
If this is going on a site where search traffic matters, accuracy problems have a second cost beyond just being wrong: thin, unreliable content is exactly the kind of thing search engines are built to demote, regardless of what wrote it. The stakes get even more concrete in a job application, where an invented achievement on a resume is a specific, checkable claim someone might actually ask about in an interview.