Claude Academy
Sign in

Vault / wiki/301/practice/ccao/domain-2-output-evaluation-and-validation.md

updated 2026-07-16

Practice — CCAO-F Domain 2: Output Evaluation and Validation (21%)

19 scenario-based MCQs. Answer key + explanations at the bottom.


Q1

A marketing associate asks why Claude's fabricated statistics are so hard to catch during review. Which explanation is most accurate?

A. Fabricated content is produced by the same plausible-continuation process as accurate content, so it reads just as fluent and confident B. Fabrications cluster around events after the knowledge cutoff, so they hide inside recent-news topics C. Claude marks its uncertain claims with hedging language, but reviewers routinely overlook the hedges D. Fabrications appear only late in long conversations, after the context window has filled with earlier turns

Q2

An operations manager used Claude to draft a competitive analysis that cites specific market-share figures and named industry studies. Before sending it to leadership, which verification approach is best?

A. Ask Claude to reconfirm each figure and keep only the ones it stands behind B. Regenerate the analysis several times in fresh conversations and keep only the figures that appear consistently in every version C. Trace each figure and named study back to an original external source before circulating D. Open a second conversation and have Claude critique the first conversation's claims

Q3

A colleague says a Claude-drafted client brief can skip review because Claude included a citation for every claim. What is the strongest response?

A. Citations make claims easier to check, but Claude can fabricate sources too — every citation still needs verification B. The colleague is right; requesting citations is the accepted way to eliminate hallucination risk C. The brief still needs a review pass for tone, structure, and formatting, but the cited factual claims themselves can be treated as reliable D. The brief is safe to send if regenerating it produces the same citations a second time

Q4

A PM asks Claude to produce a product requirements document the team will revise repeatedly over the next week and eventually share. Which output format should she request?

A. An inline response she copies into a document after each revision B. An artifact, since it is a standalone deliverable the team will iterate on C. A structured table, since requirements feed downstream project tools D. A fresh inline summary regenerated at the start of each working session so the team is always looking at the latest thinking

Q5

An operations analyst asks Claude to extract vendor name, contract value, and renewal date from 30 uploaded contracts so the results can be pasted into a tracking spreadsheet. Which output format best fits?

A. An artifact containing a polished narrative write-up of each contract's key terms, obligations, and renewal provisions B. Inline prose paragraphs walking through the contracts one by one C. An inline executive summary with highlights from notable contracts D. A structured table with one row per contract and one column per field

Q6

During a live meeting, a manager needs Claude to quickly clarify the difference between two pricing terms so she can respond on the spot. Which format best matches the need?

A. An artifact, so the explanation is preserved as a standalone reference B. A structured comparison table she can export to the team's wiki afterward C. A concise inline answer, since this is a quick one-off question D. A Project with the pricing glossary uploaded as a knowledge file

Q7

A recruiter uses Claude to draft candidate personas and notices the engineering personas are consistently young men while the HR personas are consistently women. What does this pattern most likely reflect, and what should she do?

A. A malfunction in Claude's safety filtering; she should pause persona drafting and report the behavior through her organization's support channel before continuing B. Bias absorbed from patterns in training data; she should critically review and revise the defaults rather than accept them C. Hallucination, since the personas describe specific people who do not actually exist D. A capability ceiling of the current model tier; the fix is rerunning the task on a stronger model

Q8

A director asks Claude, "Explain why moving our conference to Austin is the right call," and receives a persuasive analysis. Before treating it as validation of the decision, what should she recognize?

A. The framing steered Claude toward agreement; she should re-ask neutrally or request the case against the move B. Claude's honesty training means it would have pushed back on the framing if the relocation were a mistake C. The persuasiveness of the analysis is itself good evidence that the underlying reasoning is sound D. The analysis can be trusted as long as none of its supporting statistics turn out to be fabricated

Q9

Forty turns into a document-review conversation, Claude begins attributing statements to the uploaded report that are not in it. What is the most likely cause, and the fix?

A. Knowledge cutoff — the report postdates Claude's training data, so she should enable web search for it B. Working-memory degradation — early material competes with newer turns; re-supply or summarize the key sections C. Thin training data — Claude is interpolating, so she should ground it with additional external sources D. Model-tier limitation — the model is not capable enough for document review at this depth, so she should restart the same review on a more capable model

Q10

A consultant uses Claude for four tasks in one afternoon. Which output warrants the most independent verification before she uses it?

A. A summary of the 20-page strategy memo she uploaded at the start of the conversation, which Claude has had in context the whole time B. A rewrite of her own draft email into a warmer, more direct tone C. Competitor revenue figures requested from Claude's memory, with no sources provided D. A brainstormed list of icebreaker activities for a client workshop

Q11

An educator asks Claude about an obscure 1970s local schooling policy and gets a detailed, confident answer full of names and dates. Given how Claude generates text, what should she conclude?

A. The detail suggests the topic was well covered in training data, so the answer is probably sound B. The names and dates are likely accurate, though the interpretive claims deserve a check C. Claude would have said "I don't know" if it lacked the information, so confidence can stand in for accuracy D. Specificity is not evidence of accuracy; on thin topics Claude fluently interpolates plausible details

Q12

A compliance officer asks Claude to summarize current data-privacy regulations for a client memo, providing no documents and using no search. Which risk is most important to address before using the output?

A. Bias — the summary may lean toward one regulatory framework over another B. Context exhaustion — complete regulatory texts are far too long to fit into a single conversation's context window without degrading the summary C. Staleness — Claude's knowledge has a cutoff, so she should supply current texts or use search, then verify D. Format mismatch — regulatory summaries should be requested as structured output

Q13

A team's weekly Claude-drafted metrics summaries occasionally invent numbers. Two fixes are proposed: (1) attach the actual metrics export to each request and instruct Claude to answer only from it; (2) ask Claude to append a confidence score to every figure. Which assessment is correct?

A. Proposal 2 is stronger — confidence scores let reviewers skip checking the figures Claude marks as certain B. Proposal 1 is stronger — grounding constrains Claude to supplied, checkable material; a confidence score is itself generated text C. Both work equally well, since each proposal directly targets the hallucination mechanism D. Neither works — invented numbers at this level can only be fixed by moving to a more capable model

Q14

A manager fact-checks every claim in a Claude-drafted report and finds them all accurate. According to the AI Fluency framework's Discernment competency, what has she not yet evaluated?

A. Nothing further — Discernment is the fact-verification competency, and the fact-check completed it B. Whether the task should have been delegated to Claude in the first place C. Whether the report's AI involvement was properly disclosed to its readers D. The process — did Claude reason sensibly — and the performance of the collaboration itself

Q15

An associate uses Claude both to brainstorm team-offsite ideas and to draft a public press release containing financial figures. Per the Delegation–Diligence loop, how should her verification effort differ?

A. Apply the same rigor to both, since a consistent review process is what matters most B. Scrutinize the brainstorm more heavily, because open-ended creative output hallucinates most C. Verify the press release far more heavily — higher stakes demand narrower delegation and heavier diligence D. Skip heavy verification on both and instead ground each prompt with source documents, which removes the need for human review

Q16

To check a statistic Claude produced, an analyst regenerates the response three times; the same number appears each time, so he marks it verified. What is the flaw in his method?

A. Nothing — consistency across regenerations is a recognized verification technique B. He should have run each regeneration in a separate conversation to keep them independent C. Three runs is too few; five or more are needed before consistency means anything D. Consistency shows what Claude reliably predicts, not what is true — he checked the model against itself

Q17

A customer-success lead has three requests: (1) a quick answer to what "NPS" stands for, (2) a churn-risk table her team will import into their tracker, and (3) a QBR deck outline she will refine across several sessions. Which format assignment is best?

A. (1) inline, (2) structured table, (3) artifact B. (1) artifact, (2) inline, (3) structured table C. (1) structured table, (2) artifact, (3) inline D. (1) inline, (2) artifact, (3) structured table

Q18

After several hallucination incidents, a team decides every Claude-generated product description — thousands per month — must be automatically checked against the product database before publishing. As the associate on the team, what is the appropriate move?

A. Spot-check a weekly sample of descriptions by hand in claude.ai and accept the residual risk B. Scope the requirement and hand it off to developers — validation at this volume is an integration build, not a chat task C. Paste the relevant product database records into each conversation and instruct Claude to verify every one of its own descriptions against them D. Build the validation pipeline herself inside claude.ai, since the associate owns output quality

Q19

A manager runs her first Claude Cowork task: reorganizing a shared drive folder and renaming files. Which practice best reflects the recommended approach to validating agentic work?

A. Run the task on a copy of the folder first and review the result before letting Claude touch the originals B. Rely on the task report, since verifying results is already a built-in step of the Cowork task loop C. Open a fresh conversation and ask Claude whether the reorganization was performed correctly D. Skip detailed review for file operations, since renames and moves are simple to reverse later

Answers

Q1: A. Hallucination is next-token prediction operating without grounding: the model continues text plausibly, so false content carries the same fluency and confidence as true content — fluency ≠ accuracy. (B) is wrong because hallucination isn't confined to post-cutoff topics; thin training data anywhere triggers it. (C) fails because Claude has no built-in "I don't know" reflex that reliably hedges. (D) confuses hallucination with long-context degradation, a separate failure mode.

Q2: C. The exam's fact-checking rule is to verify against sources, not against the model itself. Only tracing figures to original external sources does that. (A) and (D) still ask the model to vouch for itself — a fluent restatement of a fabrication proves nothing. (B) tests consistency, not truth; a repeated fabrication is still a fabrication.

Q3: A. Asking for citations is a useful mitigation because it makes claims checkable, but citations can themselves be confidently fabricated, so verification is still required. (B) overstates the mitigation into a guarantee. (C) wrongly narrows review to style once citations exist. (D) is the consistency-equals-truth fallacy — regeneration checks the model against itself.

Q4: B. A document that will be revised repeatedly and shared is the textbook artifact case: a standalone, iterable deliverable that stays editable and versioned in its own panel. (A) and (D) force manual copying or regeneration and lose iteration continuity — the single trade-off they lose on. (C) mistakes a narrative document for data destined for downstream tools.

Q5: D. The results are consumed downstream (a spreadsheet), so structured output — one row per contract, one column per field — is the right format. (A) misuses an artifact for what is a data-extraction job, not an iterable deliverable. (B) and (C) produce prose that must be manually re-parsed into the tracker, defeating the purpose.

Q6: C. A quick one-off question in the moment is exactly what inline responses are for. (A) adds artifact overhead with no iteration or reuse need. (B) optimizes for a downstream use that doesn't exist yet. (D) is configuration machinery for recurring work, not a live-meeting answer — right tool, wrong timescale.

Q7: B. Systematically skewed defaults are bias inherited from patterns in training data; the Discernment response is to review critically and revise rather than accept what the model produces unprompted. (A) mislabels bias as a safety malfunction. (C) confuses bias (skewed patterns) with hallucination (fabricated facts). (D) is a wrong-layer fix — a stronger tier doesn't remove training-data bias; human review does.

Q8: A. Steerability cuts both ways: a prompt framed as "explain why X is right" steers Claude to argue for X, so agreement is an artifact of the framing, not independent validation. The test is a neutral framing or an explicit request for the opposing case. (B) overtrusts honesty training as a pushback guarantee. (C) is fluency-equals-soundness. (D) misses that the reasoning can be one-sided even when every statistic is real.

Q9: B. Misattributing content from a document that was in context earlier, deep into a long chat, is the working-memory × knowledge collision: early material competes with newer turns. The fix is to re-supply or summarize the key sections. (A) can't explain errors about a document Claude previously handled correctly. (C) is the wrong mechanism — the source was provided, not recalled. (D) is a wrong-layer fix for a context problem.

Q10: C. "Asked to recall, not retrieve" is the highest-risk pattern: specific figures pulled from model memory with no grounding are the classic hallucination setup. (A) and (B) are grounded in material she supplied, so generation is constrained and checkable. (D) is creative output where factual accuracy isn't the success criterion — nothing to verify against.

Q11: D. On thin topics, next-token prediction interpolates plausible-sounding specifics — detail and confidence are properties of the generation process, not signals of accuracy. (A) inverts the inference; obscurity predicts thin data. (B) has it backwards — precise names and dates are exactly what gets fabricated. (C) assumes an "I don't know" reflex the mechanism doesn't have unless explicitly permitted.

Q12: C. The knowledge cutoff means "current" regulations may have changed since training; recent material must be supplied in the prompt or fetched via search/research, then verified — especially in a compliance context. (A) is conceivable but not the dominant risk for recency-sensitive law. (B) invents a problem; nothing was uploaded. (D) is cosmetic next to a correctness risk.

Q13: B. Providing source material and instructing Claude to answer only from it attacks the mechanism — generation without grounding — and makes every figure checkable against the export. A self-reported confidence score is produced by the same plausible-continuation process as the figures themselves, so (A) builds trust on generated text. (C) falsely equates a structural fix with a cosmetic one. (D) is the wrong-layer "bigger model" reflex.

Q14: D. Discernment has three parts: the product (correct, complete, appropriate), the process (did Claude reason sensibly), and the performance (is the collaboration working). Fact-checking covers only the product. (A) collapses Discernment into fact-checking. (B) is Delegation and (C) is Diligence — real competencies, but different Ds.

Q15: C. The Delegation–Diligence loop scales oversight with stakes: what you hand off determines what you must verify and own, and a public release with financial figures is high-stakes. (A) sounds disciplined but wastes scrutiny where errors are cheap and underweights where they're costly. (B) inverts the risk — brainstorms aren't consumed as fact. (D) overtrusts grounding; it reduces hallucination but doesn't discharge accountability for what ships.

Q16: D. Regeneration consistency measures what the model stably predicts, which for a thinly-grounded fact can be a stable fabrication — he never left the model to check a source. (A) endorses the fallacy. (B) and (C) tweak the procedure's mechanics while keeping its core flaw: every variant still verifies the model against itself rather than against an external source.

Q17: A. Quick one-off answer → inline; data consumed by a downstream tool → structured table; a deliverable refined across sessions → artifact. That is the canonical mapping. (B), (C), and (D) each misplace at least one pairing — most tellingly, putting the iterable deck outline anywhere but an artifact or the tracker-bound data anywhere but a structured table.

Q18: B. Automated validation of thousands of outputs per month against a database is an integration build — custom code and API work — and the hallmark associate move is to scope it and hand it to developers. (A) substitutes sampling for the stated requirement of checking every description. (C) doesn't scale and verifies the model against itself with the database as a prop for volume the chat surface can't handle. (D) attempts developer work on the wrong surface.

Q19: A. Cowork best practice is to start with reversible tasks — work on copies before granting write access to originals — and to review before you rely. (B) is the tempting runner-up: the task loop does include a verify step, but Claude's self-verification doesn't replace human review, since the human stays accountable. (C) asks the model to vouch for its own work. (D) waves off review on a reversibility assumption the recommended practice explicitly doesn't make for originals.