P20 · CASE 0013 · POPULATED EVIDENCE RECORD
Can We Trust Answers Generated by AI?
AI-generated answers can support research, drafting and decision preparation, but they should not be trusted by default. Reliance must be claim-specific, task-specific, model/version-specific, time-sensitive and proportional to consequences.
Canonical source basis:
CP1_OBSERVATORY_P20_CASE0013_SOURCE_MAP_AND_EVIDENCE_MATRIX_V0_1_20260731.xlsx
Claims
Claims retain their original research-state labels. A status is not a probability of truth.
| ID | State | Claim | Support refs |
|---|---|---|---|
| C01 | STRONG SUPPORT | AI-generated answers should not be trusted by default. | S01–S09; S12–S18 |
| C02 | STRONG SUPPORT | Reliability is specific to the claim, task, model/version, tools, language and date. | S01–S10; S13–S18 |
| C03 | STRONG SUPPORT | Fluency and confidence are weak indicators of factual correctness. | S01; S02; S12–S18 |
| C04 | PROVISIONAL SUPPORT | Accessible grounding improves reliability only when the context is sufficient and relevant. | S04; S10; S17 |
| C05 | STRONG SUPPORT | The existence of citations does not prove that claims are supported. | S02; S04; S18 |
| C06 | PROVISIONAL SUPPORT | Abstention and calibrated uncertainty can reduce confident errors. | S12; S13; S15–S17 |
| C07 | STRONG SUPPORT | Vendor system cards are useful disclosures but cannot independently validate reliability. | S06–S09 |
| C08 | PROVISIONAL SUPPORT | Multi-step and agentic workflows can amplify small model or tool errors. | S02; S03; S11; S17 |
| C09 | STRONG SUPPORT | Current information requires dated external verification. | S04–S10 |
| C10 | STRONG SUPPORT | High-stakes reliance requires qualified human accountability. | S02–S05; S18 |
| C11 | OPEN | Different models may share the same error or source, so agreement is not independent corroboration. | S10–S17 |
| C12 | PROVISIONAL SUPPORT | Private or displayed reasoning is not evidence unless claims and tools are independently inspectable. | S04; S11; S12 |
| C13 | STRONG SUPPORT | Model improvements reduce some errors but do not eliminate non-zero failure. | S01; S02; S06–S09; S14 |
| C14 | PROVISIONAL SUPPORT | AI is safest as a layered assistant: generate, inspect, verify, correct, decide. | S04; S05; S10; S11; S15–S17 |
| C15 | OPEN | No single universal AI-answer trust score is currently justified. | S01–S18 |
Sources · 18 serialized
| ID | Source | Research status |
|---|---|---|
| S01 | Stanford HAI — AI Index Report 2026 | ACCEPTED EVIDENCE |
| S02 | International AI Safety Report 2026 | ACCEPTED EVIDENCE |
| S03 | International AI Safety Report — Extended Summary | ACCEPTED CONTEXT |
| S04 | NIST AI 600-1 — Generative AI Profile | ACCEPTED METHOD |
| S05 | NIST AI Risk Management Framework | ACCEPTED METHOD |
| S06 | OpenAI GPT-5.6 Preview System Card | ACCEPTED DISCLOSURE |
| S07 | OpenAI GPT-5.5 Instant System Card | ACCEPTED DISCLOSURE |
| S08 | Anthropic — Model System Cards | ACCEPTED DISCLOSURE |
| S09 | Anthropic Claude Opus 4.8 System Card | ACCEPTED DISCLOSURE |
| S10 | Google Research — Sufficient Context in RAG | ACCEPTED EVIDENCE |
| S11 | Google Research — Science One Chain-of-Evidence | ACCEPTED EXPERIMENTAL |
| S12 | Kalai et al. — Why Language Models Hallucinate | ACCEPTED RESEARCH |
| S13 | OpenAI — SimpleQA | ACCEPTED BENCHMARK |
| S14 | SimpleQA Verified | ACCEPTED BENCHMARK |
| S15 | Conformal Linguistic Calibration | ACCEPTED RESEARCH |
| S16 | Behaviorally Calibrated Reinforcement Learning | ACCEPTED RESEARCH |
| S17 | HALT-RAG | ACCEPTED RESEARCH |
| S18 | Stanford — Legal RAG Hallucination Study | ACCEPTED DOMAIN EVIDENCE |
Tests / adversarial controls · 15
| ID | Failure mode / test | State | Required control |
|---|---|---|---|
| A01 | Fluency authority | ACTIVE | Separate style from claim verification. |
| A02 | Citation hallucination | ACTIVE | Open and validate every material citation. |
| A03 | Citation mismatch | ACTIVE | Match each claim to the source. |
| A04 | RAG presence fallacy | ACTIVE | Inspect retrieval relevance and sufficiency. |
| A05 | Benchmark transfer | ACTIVE | Require matched-task evidence. |
| A06 | Model-name authority | ACTIVE | Record version and verify output independently. |
| A07 | Freshness illusion | ACTIVE | Use dated authoritative verification. |
| A08 | Sycophancy | ACTIVE | Reframe prompt and test adverse evidence. |
| A09 | Reasoning theatre | ACTIVE | Verify claims, sources and tools externally. |
| A10 | Tool opacity | ACTIVE | Preserve inputs, outputs and transformations. |
| A11 | Cross-model pseudo-corroboration | ACTIVE | Seek independent evidence, not model votes. |
| A12 | High-stakes compression | ACTIVE | Require qualified human review and primary evidence. |
| A13 | Uncertainty camouflage | ACTIVE | Compare confidence to observed correctness. |
| A14 | Verification laundering | ACTIVE | Define reviewer competence and verification scope. |
| A15 | Version drift | ACTIVE | Version-lock and reassess after material change. |
Serialization boundary
This interface does not complete missing research by inference. The current record is COMPLETE_CLAIMS_SOURCES_ADVERSARIAL. Missing source-level or bilingual object-level fields remain absent until they are read from a canonical Observatory artifact.