When a citation looks convincing
A generated answer can contain a neatly formatted reference and still mislead. NIST identifies confidently presented false content as confabulation; it can also include contradictions or departures from the supplied input. Its generative AI profile recommends checking sources and citations rather than extrapolating reliability from a few successful examples. [1]

What the evidence says
NIST’s generative AI risk profile identifies confabulation among the risks of generative systems: output can be false or misleading while appearing credible. The profile also addresses privacy and information integrity. It is guidance for managing risk, not a finding that every AI answer fails. [1]
Improved on one tested task
Improves workplace performance
Reported an improvement on that task
Invented wording for comparison; not a quotation from a study.
A reference can fail in three different ways
Consider this hypothetical claim: “A study proves that an assistant is accurate in every workplace.” The reference might be invented. It might exist but discuss a different system. Or it might evaluate the right system on a narrow task, while the answer expands the result into a universal claim. Those are three different problems—existence, relevance and scope—and a clickable link only helps with the first step.
Try it with a harmless example
For a reading exercise, ask an assistant to summarize a short paragraph you already understand. Compare the answer against the original: did a qualification disappear, did “may” become “will,” or did it add an unsupported detail? This is an illustrative exercise, not a reliability benchmark. It makes a hidden problem visible: a fluent summary can subtly change the strength of a statement.
The bigger question
NIST’s AI Risk Management Framework is voluntary guidance for managing risks across the AI lifecycle. Its existence does not certify an individual product or settle a particular claim. [2] Our takeaway: judge the connection between claims and evidence, and make room for uncertainty. A useful answer can include a clear boundary around what is unknown.
Go a little deeper
Optional reading · about 1 more minute
A useful task with a visible boundary
Consider a hypothetical assistant sorting meeting notes into decisions, unresolved questions and action items. Its value is organizing material the reader already has, so the output can be compared with the input. If it changes “we might test this” into “we decided to launch,” that is not merely a stylistic edit: it changes the decision. Keeping the original wording accessible helps expose the difference.
Try a small editorial experiment
Original statement: “In a limited trial, the team observed an improvement on one task.” Candidate summary A: “The system improves performance.” Candidate summary B: “The team reported an improvement on the task it tested.” These invented sentences illustrate scope. B retains the boundary of the observation; A broadens it. The lesson is useful for human-written summaries as well as AI-generated ones.
TRY THE SCENARIO
A source says a tool helped on one tested task. Which summary best preserves the evidence?
Choose an answer to see the explanation.
Original sources
Attributed synthesis, not original reporting. Examples labeled hypothetical or illustrative are explanatory. Reviewing a source does not independently validate its findings.
- NIST: Generative AI Profile ↗
Published July 26, 2024; risk-management guidance.
- NIST: AI Risk Management Framework ↗
Framework released January 26, 2023; voluntary guidance.

