Claude Certified Architect — Professional (CCAR-P) Complete Study Guide
Professional-level • Architecture-focused • Scenario-driven • 2026 edition
This guide is designed for the Claude Certified Architect — Professional (CCAR-P) exam. It is deliberately more architecture-heavy than a developer guide: the exam expects you to connect business outcomes to system design, integration choices, operational controls, evaluation, governance, and lifecycle ownership.
Study approach: For every scenario, identify the binding business constraint first, then the architectural decision it forces, then the operational evidence that would prove the design works.
Version note: The public CCAR-P blueprint is version 1.0, effective July 2026. Model names, product surfaces, MCP implementation details, and Claude tooling can evolve faster than certification blueprints. Prioritize the durable architecture principles and use current Anthropic documentation for implementation details.
Exam at a Glance
| Item | CCAR-P |
|---|---|
| Certification | Claude Certified Architect — Professional |
| Exam code | CCAR-P |
| Level | Professional |
| Questions | 63 |
| Time limit | 120 minutes |
| Format | Multiple-choice and multiple-response |
| Passing score | 720 scaled score on a 100–1,000 scale |
| Delivery | Pearson VUE, online proctored or test center |
| Exam fee | $175 USD |
| Credential validity | 12 months |
| Formal prerequisite | None |
| Typical audience | Mid- to senior-level solution architects, AI/ML engineers, technical leads, senior software engineers |
The professional blueprint is intended for people who can own a Claude solution across its full lifecycle: discovery, design, delivery, governance, operations, and iteration.
The Seven Domains
| Domain | Weight | Core Architect Question |
|---|---|---|
| 1. Solution Design & Architecture | 17% | What is the simplest end-to-end design that meets business, reliability, and risk requirements? |
| 2. Claude Models, Prompting & Context Engineering | 13% | How should model capability, prompts, context, reuse, and cost be shaped for this workload? |
| 3. Integration | 19% | How should Claude connect securely and observably to data, tools, RAG, and enterprise systems? |
| 4. Evaluation, Testing & Optimization | 16% | How will we prove quality, diagnose failures, and improve without breaking another dimension? |
| 5. Governance, Safety & Risk Management | 14% | What risks exist, which controls mitigate them, and what evidence proves those controls operate? |
| 6. Stakeholder Communication & Lifecycle Management | 14% | How do we convert stakeholder needs into decisions, SLAs, documentation, ownership, and feedback loops? |
| 7. Developer Productivity & Operational Enablement | 7% | How do we equip teams to build, debug, operate, and govern the solution consistently? |
Weighting insight: Domains 1, 3, and 4 account for 52% of the exam. Add Domains 5 and 6 and you cover the architecture, operations, governance, and organizational judgment that defines the Professional level.
CCAR-P in One Mental Model
A professional architect should be able to move through this chain:
Business outcome
|
v
Requirements + constraints
|
v
Architecture pattern
|
v
Model + context strategy
|
v
Integration + data boundaries
|
v
Controls + governance
|
v
Evaluation + observability
|
v
Deployment + ownership
|
v
Feedback + iteration
The exam repeatedly tests whether you can choose the structural fix, not merely a mitigation.
Examples:
- Too many tools → remove or scope unnecessary tools, not just add more confirmations.
- Stale RAG → fix freshness/indexing or call the source of truth, not just use a stronger model.
- Repeated static prompt cost → cache the stable prefix, not simply truncate useful context.
- High-impact action → enforce authorization/approval before execution, not audit only afterward.
- Vague "seamless" requirement → translate it into measurable latency, handoff, and failure-state requirements.
Domain 1 — Solution Design & Architecture
Weight: 17%
This domain tests whether you can translate a business problem into an end-to-end Claude solution and choose the right architecture pattern.
The blueprint expects skill in:
- translating business problems into AI solution requirements;
- designing input → processing → output → feedback architectures;
- selecting workflow, agentic, or augmented-LLM patterns;
- designing multi-agent orchestration when justified;
- decomposing complex problems;
- aligning technical design with business value, cost, and service-level targets.
1. Start With the Business Outcome
Do not begin with:
- "Which Claude model?"
- "Should we use agents?"
- "Can we add RAG?"
- "Should this be multi-agent?"
Begin with:
- What outcome must improve?
- Who owns the outcome?
- What is the baseline today?
- What is the measurable target?
- What errors are acceptable?
- What errors are unacceptable?
- Which actions are reversible?
- What must remain deterministic?
- What must stay human-owned?
- What evidence will prove success?
Example
Weak requirement:
"Use Claude to automate claims."
Architect-quality decomposition:
| Question | Example answer |
|---|---|
| Desired business outcome | Reduce manual claim triage time |
| Baseline | 18 minutes average handling time |
| Target | Under 8 minutes |
| Automation scope | Extract, classify, summarize |
| Deterministic rules | Coverage eligibility, monetary thresholds |
| Human ownership | Final denial decision |
| Failure tolerance | No unsupported denial recommendation |
| Data constraint | Sensitive personal data |
| SLA | 95% of triage under 20 seconds |
| Audit requirement | Full trace of input, retrieval, output, reviewer decision |
Only after this decomposition should architecture begin.
2. Decompose Work by Nature
A useful decomposition separates work into three buckets.
A. Work suited to Claude
Examples:
- classification with ambiguity;
- summarization;
- drafting;
- reasoning across unstructured text;
- planning;
- natural-language interpretation;
- choosing among allowed tools.
B. Work better handled by deterministic systems
Examples:
- arithmetic;
- fixed thresholds;
- access control;
- transactional state;
- exact policy rules;
- identity validation;
- rate limiting;
- idempotency;
- schema validation.
C. Work that should remain human-controlled
Examples:
- irreversible legal decisions;
- high-impact exceptions;
- sensitive approval;
- ambiguous cases with major consequences;
- ethical judgment where policy requires a person.
Architect rule: Do not assign a deterministic business invariant to a probabilistic component simply because the model can usually reproduce it.
3. Workflow vs Agent
Workflow
The application owns control flow.
Use when:
- steps can be enumerated;
- checkpoints matter;
- observability must be simple;
- variation should be minimized;
- latency needs a predictable ceiling.
Input
|
v
Classify
|
v
Retrieve
|
v
Generate
|
v
Validate
|
v
Output
Agent
The model chooses what happens next.
Use when:
- the next step cannot be fully predetermined;
- tool choice depends on discoveries;
- number of steps varies;
- investigation/planning is open-ended.
Goal
|
v
Claude decides action
|
v
Tool / observation
|
v
Claude decides next action
|
+------> repeat
|
v
Done
Exam instinct: If the process can be written as a stable flowchart before runtime, prefer a workflow unless agentic flexibility creates measurable value.
4. Single Augmented Call vs Workflow vs Agent
Think of architecture as a spectrum.
| Shape | Use when | Main risk |
|---|---|---|
| Single augmented call | One request can reliably solve the task with context/tools | Overengineering if you add more stages |
| Workflow | Steps are known and can be explicitly controlled | Excess calls/latency if decomposed too far |
| Agent | Next step depends on discoveries | Cost, latency, observability, control |
Five elimination questions
- Predictability — Can steps be listed before execution?
- Error cost — How costly is a wrong action?
- Observability — Can operations reconstruct the path?
- Latency — Is worst-case runtime bounded?
- Cost — How many calls and how much growing context may accumulate?
5. Workflow Patterns
Prompt chaining
A -> B -> C
Use when each step depends on the previous result.
Routing
Input
|
v
Classifier
/ | \
A B C
Use when known request types require different handling.
Parallelization
/-> A -\
Input -+-> B --+-> Merge
\-> C -/
Use only when tasks are independent.
Evaluator-optimizer
Generate
|
v
Evaluate
|
v
Feedback
|
v
Improve
Use when quality can be judged against clear criteria and one pass is insufficient.
6. Architecture by Failure Mode
Do not pick patterns because they sound advanced.
Ask:
"Where can this design fail, and where can I intercept that failure?"
Workflow advantage
You can insert deterministic controls:
Generate
|
Validate schema
|
Check policy
|
Human approval
|
Execute
Agent challenge
The trajectory itself is dynamic.
Therefore you need:
- turn limits;
- tool permissions;
- trace IDs;
- state checkpoints;
- action validation;
- budget limits;
- approval gates;
- post-run reconciliation.
7. Multi-Agent Architecture
Use multiple agents when there is a concrete reason:
- context isolation;
- specialist instructions;
- independent work;
- parallelism;
- organizational/tool boundaries.
Manager-worker
/-> Research agent
/
User -> Manager ----> Security agent
\
\-> Cost agent
|
v
Synthesis
Responsibilities
Coordinator
- holds overall goal;
- decomposes;
- delegates;
- tracks work;
- reconciles completion;
- synthesizes.
Worker
- receives bounded task;
- uses narrow context/tools;
- returns structured result.
8. Multi-Agent Failure Modes
Missing worker result
The coordinator receives 9 of 10 results and summarizes confidently.
Fix:
- count dispatched work;
- count completed work;
- reconcile before synthesis;
- surface incomplete coverage.
Conflicting workers
Fix:
- explicit tie-break rule;
- adjudicator;
- human escalation;
- source-quality hierarchy.
Coordinator loses state
Fix:
- checkpoint orchestration state;
- durable job record;
- stable task IDs.
Fragmented traces
Fix:
- propagate one end-to-end correlation/trace ID.
Exam rule: Silent partial completion is more dangerous than explicit failure.
9. Human Checkpoints
Put human review at meaningful boundaries.
Good:
Agent proposes deployment plan
|
v
Human reviews complete plan
|
v
Approved execution
Poor:
Human approves every trivial read operation
Human review is expensive and can become rubber-stamping.
Use it where:
- consequence is high;
- action is irreversible;
- uncertainty is high;
- regulation requires it;
- exception handling matters.
10. Reference Architecture: Tool-Using Agent
Use when:
- task is dynamic;
- external systems are needed;
- actions depend on prior observations.
Controls:
- tool allowlist;
- per-tool authorization;
- maximum turns;
- request budget;
- audit trail;
- idempotency;
- human gate for writes.
Failure mode:
Unlimited autonomy with broad tools and no stop/approval boundaries.
11. Reference Architecture: RAG Over Stable Knowledge
Use when:
- corpus is larger than context;
- knowledge is relatively stable;
- answers need grounding/citations.
Documents
|
Ingest
|
Chunk
|
Index
|
Query -> Retrieve -> Rerank -> Claude -> Answer
Failure mode:
Using the index as if it were the current transactional system of record.
For live balances, inventory, job status, entitlement, or account state, use a tool/API to the authoritative system.
12. Reference Architecture: Document Processing
Example:
Document
|
Extract
|
Normalize
|
Claude interpretation
|
Confidence / validation
|
+-> clean -> continue
|
+-> exception -> human
Failure mode:
No exception path, so low-confidence extraction flows through as though it were verified.
13. Reference Architecture: Classify and Route
Request
|
Classifier
|
+-> FAQ/RAG
+-> Transactional tool
+-> Human
Failure mode:
Sending everything into an open-ended agent when a router plus known handlers would be cheaper, more observable, and easier to govern.
14. Architecture Feasibility
A good feasibility verdict is one of:
- feasible as scoped;
- feasible only with explicit constraints;
- not feasible.
Constraints might include:
- document size;
- data freshness;
- latency;
- review requirement;
- availability of source-of-truth API;
- regional deployment;
- model capability;
- cost ceiling.
Professional rule: "Feasible if X remains true" is stronger than an unconditional yes built on an unstated assumption.
15. Business Value Alignment
Tie design to measurable value.
Possible value pillars:
- efficiency;
- productivity;
- transformation/new capability;
- solution cost;
- quality;
- service levels;
- risk reduction.
Business case shape
Measured baseline
|
Projected improvement
|
Operational run cost
|
Human-review cost
|
Risk / compliance cost
|
Net value + payback
Common errors
- baseline guessed instead of measured;
- projection assumes full automation but design still uses reviewers;
- cost model uses only average input size;
- model/API cost excludes retrieval, storage, support, and people.
16. Service-Level Architecture
Define more than latency.
Possible SLOs:
- task accuracy;
- p95 latency;
- availability;
- retrieval freshness;
- tool success rate;
- escalation rate;
- cost/request;
- hallucination rate;
- audit completeness.
Architecture should explain how each is measured.
Domain 1 Common Traps
- Starting with a model instead of a business outcome.
- Agent where deterministic workflow is enough.
- Multi-agent by default.
- Probabilistic enforcement of deterministic policy.
- Human approval added after architecture is complete.
- No reconciliation of delegated work.
- No worst-case latency/cost budget.
- RAG used for live state.
- Feasibility stated without constraints.
- Business value stated without baseline.
Domain 1 Scenario Drills
Scenario 1
A legal intake system always performs extraction, policy lookup, clause comparison, and draft generation.
Best answer: Workflow.
Scenario 2
Claude must investigate an unfamiliar production incident and choose logs/tools dynamically.
Best answer: Agent, with bounded tools, limits, and traceability.
Scenario 3
A coordinator delegates 100 document reviews and receives only 97 results.
Best answer: Block final synthesis until coverage is reconciled or missing items are explicitly handled.
Scenario 4
An agent is asked to enforce a fixed refund threshold written only in its prompt.
Best answer: Move the threshold to deterministic business logic.
Scenario 5
Stakeholder says "make onboarding seamless."
Best answer: Translate "seamless" into latency, handoff, error-state, and visibility requirements before architecture.
Domain 1 Rapid Review
- Business outcome first.
- Decompose model / deterministic system / human.
- Known path → workflow.
- Unknown dynamic path → agent.
- Multi-agent only for justified specialization/isolation.
- Reconcile worker completion.
- Human gates before consequential actions.
- RAG for knowledge; tools for live state.
- Feasibility must include constraints.
- Value claims need a measured baseline.
Domain 2 — Claude Models, Prompting & Context Engineering
Weight: 13%
The blueprint expects you to:
- choose models based on capability, speed, and cost;
- design system prompts, templates, and guardrails;
- apply appropriate prompting techniques;
- manage context/token growth;
- use caching, modular prompts, and reusable packaged instructions where appropriate.
1. Model Choice Is an Architecture Decision
Do not default all stages to one model.
Different stages may need different capability.
Example:
Cheap classifier
|
v
Balanced general model
|
v
Strong model only for difficult exception
Choose with:
- task difficulty;
- error consequence;
- latency target;
- throughput;
- cost;
- tool-use reliability;
- context size;
- eval results.
Professional rule: Model selection must be justified by measured task performance, not brand hierarchy.
2. Start Balanced, Move on Evidence
A practical approach:
- choose a reasonable baseline;
- run representative evals;
- identify failing categories;
- move stronger only where required;
- move cheaper only where evals prove it is safe.
Avoid:
- strongest model everywhere;
- cheapest model everywhere;
- model routing with no per-route evaluation.
3. Treat Model Changes as Releases
A model change can affect:
- tool selection;
- output shape;
- latency;
- cost;
- refusals;
- reasoning;
- verbosity;
- retrieval use;
- token consumption.
Before rollout:
- run regression suite;
- segment by input class;
- compare latency/cost;
- define rollback threshold before test;
- canary/shadow test;
- monitor post-release.
4. Four Context Concepts
Keep these distinct.
Context window
The active information Claude can attend to for one request.
Retrieval
External knowledge fetched at request time.
Application state
Authoritative state held outside Claude.
Memory / summary
Continuity your application stores and passes back.
Trap: Conversation context is not a database and retrieval is not current transactional state.
5. Context Strategy Options
Load up front
Use when:
- context is bounded;
- stable;
- reused;
- cacheable.
Risk:
- token cost;
- attention dilution;
- growth over time.
Carry recent context
Use for:
- conversation;
- agent loop;
- staged workflow.
Risk:
- important old details disappear;
- prefix mutates and cache reuse drops.
Fetch just in time
Use when:
- corpus is large;
- only a subset is relevant;
- data changes independently;
- citations matter.
Risk:
- retrieval misses relevant material.
Compact / summarize
Use for:
- long sessions;
- reducing repeated history.
Risk:
- summaries can lose IDs, values, constraints, and prior decisions.
6. Context Budget
Do not size only to the hard model limit.
Budget for:
System instructions
+ tools
+ conversation
+ retrieved context
+ tool results
+ output headroom
Large windows are ceilings, not targets.
7. Stable vs Dynamic Prompt Structure
Good cache-friendly order:
Stable system instructions
Stable tool definitions
Stable reference context
Dynamic retrieved context
Dynamic conversation
Current user request
Changing early content breaks reuse of later cached prefix sections.
8. Prompt Caching
Use when large stable prefixes recur.
Good candidates:
- policy manual;
- stable tool definitions;
- long standard instructions;
- codebase context reused over many calls;
- repeated reference documents.
Caching trades freshness against reuse.
Do not cache stale data that must always reflect current state.
9. Prompt Templates
A good reusable prompt defines:
- role/scope;
- task;
- inputs;
- constraints;
- output contract;
- error/unknown behavior;
- examples where useful.
Bad template:
"Analyze the case."
Better:
"Using only the supplied case and policy excerpts, return risk level, evidence, unresolved questions, and recommended next action. If policy evidence is missing, state
insufficient_evidencerather than guessing."
10. Guardrails in Prompts vs Structure
Prompt instruction:
"Never omit the evidence field."
Stronger structural approach:
- schema requires
evidence; - downstream validation rejects missing evidence.
Architect rule: Use prompts for model behavior; use structure/code for guarantees.
11. Zero-Shot, Few-Shot, Explicit Reasoning
Zero-shot
Use when the task is clear enough from instructions.
Few-shot
Use when:
- label boundaries are subtle;
- output shape is unusual;
- examples communicate behavior more efficiently than prose.
Explicit reasoning/decomposition
Use when the problem benefits from intermediate analysis.
Do not require verbose reasoning when:
- simple classification;
- extraction;
- deterministic lookup.
Every added technique costs tokens and latency.
12. Bias in Prompt Construction
Bias can enter through:
- leading wording;
- skewed few-shot examples;
- uneven retrieval corpus;
- assumptions embedded in the task definition.
Evaluate across groups/categories, not just aggregate quality.
13. Skills and Reusable Procedures
Reusable packaged instructions are useful when:
- the same procedure should be shared;
- it needs versioning;
- it should load on demand;
- rollback/governance matters.
Architect question:
Should this be loose prompt text copied between teams, or a versioned reusable capability?
At scale, versioning and ownership matter.
14. Context Failure Modes
Context overload
Symptoms:
- forgotten constraints;
- tool confusion;
- cost rise;
- latency increase.
Stale summary
Symptoms:
- earlier decision or identifier lost.
Wrong retrieval
Symptoms:
- fluent answer grounded in irrelevant chunk.
Tool result flooding
Symptoms:
- giant backend payload dominates context.
Mitigations:
- targeted retrieval;
- compact results;
- selective history;
- subagents;
- tool search;
- summary checkpoints.
Domain 2 Common Traps
- Strongest model without eval evidence.
- Model swap with no regression gate.
- Confusing retrieval with memory/state.
- Caching dynamic live state.
- Variable content placed before stable cache prefix.
- Long prompts replacing structural validation.
- Few-shot examples treated as universal truth.
- Context window treated as target rather than ceiling.
- Summary used without preserving critical state.
Domain 2 Scenario Drills
Scenario 1
A 60k-token policy prefix is identical on most requests.
Best answer: Put the stable material first and use prompt caching.
Scenario 2
A classifier performs well on a balanced model but latency/cost are high.
Best answer: Evaluate a cheaper/faster tier specifically on the classifier workload.
Scenario 3
A model update improves average score but significantly hurts one high-risk input category.
Best answer: Block or partially route the migration based on segmented evals.
Scenario 4
A long agent conversation loses an important customer ID during compaction.
Best answer: Preserve load-bearing state outside lossy summary text.
Domain 2 Rapid Review
- Model selection is eval-driven.
- Model changes are releases.
- Context window ≠ retrieval ≠ state ≠ memory.
- Stable prefix first.
- Cache repeated stable context.
- Use the lightest prompting technique that works.
- Structural rules beat prose-only guarantees.
- Long context can degrade quality before the hard limit.
- Preserve critical state outside summaries.
Domain 3 — Integration
Weight: 19% — the largest domain
The blueprint expects you to:
- audit tool/agent configuration for capability bloat;
- analyze authentication and authorization gaps;
- reason about accuracy/latency trade-offs;
- design observability at scale;
- design RAG chunking/indexing;
- match retrieval strategy to corpus/query shape;
- choose among direct APIs, MCP, CLI, and agent-to-agent patterns;
- use progressive discovery rather than monolithic context/tool surfaces.
1. Integration Has Three Layers
Entry point
Who or what interacts with the system?
Examples:
- web app;
- Slack;
- IDE;
- service API;
- contact-center desktop.
Build-time interface
What do engineers integrate against?
Examples:
- Claude SDK;
- raw API;
- MCP;
- Agent SDK;
- CLI.
Delivery route
Where does inference run and whose enterprise/compliance contract applies?
Examples:
- direct provider route;
- approved cloud marketplace/service;
- enterprise platform.
Professional trap: Choosing a delivery route does not automatically determine the user entry point or developer integration interface.
2. Constraint Before Preference
Evaluate hard constraints first:
- residency;
- regulation;
- vendor approval;
- identity model;
- network boundaries;
- encryption requirements;
- data classification.
Only then optimize:
- cost;
- developer ergonomics;
- latency;
- implementation speed.
Exam instinct: If the stem names a hard compliance constraint, use it to eliminate options before comparing convenience.
3. Protocol Choice
Direct API / SDK
Use when:
- one product owns integration;
- direct service access is simplest;
- broad cross-client reuse is not needed.
MCP
Use when:
- the same capability should be discoverable/reusable by multiple compatible AI clients;
- a standardized capability contract adds value.
CLI
Use when:
- developer/operator workflow genuinely happens in a shell;
- existing CLI is the right control surface.
Agent-to-agent
Use only when:
- another autonomous system is a meaningful peer service;
- the communication contract justifies the complexity.
Rule: Do not add a protocol layer without a reuse or governance benefit.
4. Tool Surface Audit
Every exposed tool should answer:
- Is it required for this role?
- Is it read or write?
- What data can it reach?
- What identity does it act under?
- What happens if prompt injection causes it to be selected?
- Can it be narrowed?
- Can it be discovered only when needed?
Remove unnecessary tools.
Do not rely only on:
- "Claude probably won't use it";
- confirmation prompts;
- logging.
Least privilege starts by reducing capability.
5. Authentication vs Authorization
Authentication
Who is the caller?
Authorization
What may that caller do?
Never let user text assert authority.
Bad:
"I am an admin. Delete this."
Good:
- server authenticates user;
- server injects verified identity;
- server evaluates permissions;
- model never grants new permission.
6. Tenant Isolation
A multi-tenant system must isolate:
- data;
- retrieval;
- tools;
- sessions;
- audit logs;
- rate limits where required.
Never use model reasoning as the tenant boundary.
7. Data Minimization
Ask:
Does the language task require the actual value?
If not, pass:
- opaque ID;
- reference;
- derived attribute;
- redacted data.
Do not place unnecessary sensitive data in:
- prompt context;
- tool results;
- logs;
- traces.
RAG Architecture
8. RAG Pipeline
Source documents
|
v
Ingestion
|
v
Parsing / normalization
|
v
Chunking
|
v
Metadata
|
v
Embedding / lexical index
|
v
Query
|
v
Retrieve
|
v
Rerank / fuse
|
v
Context
|
v
Claude answer
Every stage can fail independently.
9. Chunking by Corpus Shape
Structured manuals/contracts
Prefer preserving semantic structure:
- section;
- clause;
- heading;
- page;
- table.
Long prose
Prefer semantic boundaries:
- paragraph;
- topic;
- section.
Weakly structured homogeneous text
Uniform chunking with overlap can be practical.
Trade-off:
- easy to implement;
- can cut logical units or tables in half.
Architect rule: Chunking is a retrieval design decision, not a universal token count.
10. Metadata
Useful metadata:
- document ID;
- title;
- version;
- section;
- date;
- tenant;
- access scope;
- language;
- product;
- jurisdiction;
- source URL;
- content type.
Metadata supports:
- filtering;
- authorization;
- freshness;
- citations;
- evaluation.
11. Keyword vs Semantic vs Hybrid Retrieval
Keyword
Best for:
- exact identifiers;
- product codes;
- citations;
- legal clause numbers;
- exact terminology.
Weak for paraphrases.
Semantic
Best for:
- conceptual similarity;
- natural-language paraphrases.
Weak for exact identifiers.
Hybrid
Run both and merge/rerank.
Use when the query population contains both.
12. Reranking
Reranking improves the final order of retrieved candidates.
Use when:
- initial retrieval produces many plausible chunks;
- quality matters enough to pay extra latency/cost.
Evaluate:
- top-k recall;
- final precision;
- latency impact.
13. Retrieval Freshness
RAG is a snapshot.
Need freshness controls:
- re-index schedule;
- event-driven ingestion;
- versioning;
- stale document removal;
- update watermark;
- freshness metrics.
For live transactional data:
Call the system of record instead of relying on an index.
14. Retrieval Evaluation
Measure separately:
Retrieval quality
- Did relevant evidence enter top-k?
- Was required document retrieved?
- Was stale material returned?
Generation quality
- Did Claude use the evidence correctly?
- Did it cite unsupported claims?
- Did it abstain when evidence was absent?
Do not blame the model for missing context that retrieval never supplied.
15. Progressive Discovery
Avoid preloading:
- hundreds of tools;
- entire knowledge corpus;
- every schema;
- every connector.
Instead:
Discover domain
|
Load relevant tools
|
Retrieve relevant context
|
Execute
Benefits:
- fewer tokens;
- better selection;
- smaller attack surface;
- easier governance.
16. Observability Architecture
Trace across:
User request
|
Model call
|
Retrieval
|
Tool call
|
Queue / service
|
Downstream API
|
Response
Capture:
- trace ID;
- model;
- prompt version;
- token counts;
- latency per span;
- retrieval IDs;
- tool names/arguments;
- authorization decision;
- errors/retries;
- final result;
- human review.
17. Reliability Controls at the Right Layer
Retry
Near the failing call.
Circuit breaker
At dependency boundary.
Queue
For durable asynchronous work.
Fallback
In orchestration.
Idempotency
At side-effecting operation.
Human escalation
At consequence/uncertainty boundary.
Trap: Putting every reliability control in the agent prompt.
18. Multi-Tenant API Keys
Avoid one shared credential if you need:
- per-tenant limits;
- per-tenant attribution;
- blast-radius control;
- separate revocation;
- auditability.
Use scoped identity/credentials where architecture requires it.
19. Integration Failure Taxonomy
| Symptom | Investigate first |
|---|---|
| Correct answer before document refresh, wrong after | indexing/retrieval |
| Cross-tenant record visible | authorization/filtering |
| Tool chosen incorrectly | tool surface/description |
| Tool times out | dependency trace/retry |
| Agent gets expensive | tool/context bloat |
| Exact ID queries fail | keyword/index strategy |
| Paraphrase queries fail | semantic retrieval |
| Responses stale | freshness/source-of-truth choice |
Domain 3 Common Traps
- Compliance checked after convenience.
- MCP used with only one client and no reuse benefit.
- RAG used for live state.
- Authentication confused with authorization.
- User text trusted as identity.
- Shared tool catalog with unnecessary capabilities.
- Observability deferred until after launch.
- Giant preloaded context/tool surface.
- One retrieval strategy used for every corpus/query.
- Retry/circuit breaker/fallback placed at wrong layers.
Domain 3 Scenario Drills
Scenario 1
Only support agents need refund tools, but every employee agent has them.
Best answer: Remove/scoped the refund tools from roles that do not require them.
Scenario 2
An index is refreshed nightly, but account balances change every minute.
Best answer: Call the account system of record for current balance.
Scenario 3
Users search both exact case IDs and natural-language descriptions.
Best answer: Hybrid lexical + semantic retrieval is a strong candidate.
Scenario 4
A system integrates one private internal API and no other client will reuse the tool surface.
Best answer: Direct SDK/API may be simpler than adding MCP.
Scenario 5
Security review asks how a refund decision can be reconstructed.
Best answer: End-to-end trace including identity, authorization, model/tool path, and result.
Domain 3 Rapid Review
- Integration is the heaviest domain.
- Hard constraints eliminate options first.
- Entry point, build interface, and delivery route are different.
- Remove unnecessary tools.
- AuthN identifies; AuthZ permits.
- RAG is not transactional truth.
- Chunk by corpus shape.
- Keyword exact, semantic paraphrase, hybrid mixed.
- Progressive discovery beats monolithic exposure.
- Trace the whole path.
Domain 4 — Evaluation, Testing & Optimization
Weight: 16%
The blueprint expects you to:
- define metrics for quality, latency, cost, safety, and security;
- design representative evaluation datasets;
- combine automated, model-based, and human methods;
- run A/B tests and iterative improvements;
- diagnose prompt, retrieval, hallucination, orchestration, and model-fit problems;
- optimize token/latency/cost;
- monitor production behavior.
1. Define Success Before Building
Weak:
"Responses should be accurate."
Strong:
"At least 97% of sampled policy answers must be fully supported by retrieved evidence, with zero unsupported high-severity compliance claims."
A measurable requirement includes:
- behavior;
- threshold;
- evaluation method;
- population;
- consequence.
2. Evaluation Dataset Composition
Include:
- representative normal cases;
- edge cases;
- adversarial cases;
- malformed inputs;
- rare high-risk cases;
- known prior failures;
- multilingual cases if relevant;
- multi-turn transcripts if product is conversational.
Trap: A clean happy-path dataset can produce impressive scores that do not predict production.
3. Three Grading Levels
Code-based
Use for deterministic checks:
- JSON validity;
- enum;
- numeric threshold;
- required fields;
- exact policy IDs;
- forbidden action;
- latency/cost.
Model-based judge
Use for:
- faithfulness;
- instruction following;
- quality;
- tone;
- semantic equivalence;
- nuanced policy adherence.
Human review
Use for:
- calibrating judge;
- novel high-stakes behavior;
- ambiguous policy;
- fairness review;
- expert-domain judgment.
Rule: Use the cheapest reliable grader. Stakes determine how much you care; ambiguity determines whether deterministic grading is possible.
4. Calibrate Model Judges
A judge can be wrong.
Calibrate against:
- human-labeled benchmark;
- diverse cases;
- borderline cases.
Where possible:
- use a different model/configuration from the one being evaluated;
- track judge disagreement;
- revisit rubric.
5. Evaluation Before Production Code
Benefits:
- forces clear success definition;
- exposes assumptions early;
- creates regression gate;
- makes model/prompt changes measurable.
If you cannot describe how to test a claimed capability, the design claim is weak.
6. Failure Taxonomy
Prompt failure
Symptom:
- instruction ambiguous;
- output varies in avoidable ways.
Fix:
- clarify prompt/template;
- add structure/examples.
Hallucination / grounding failure
Symptom:
- answer contains unsupported claim.
Fix:
- retrieval/tool grounding;
- citation requirement;
- verification;
- abstention behavior.
Model mismatch
Symptom:
- task exceeds selected tier capability.
Fix:
- model selection after eval.
Retrieval failure
Symptom:
- correct source not supplied.
Fix:
- chunk/index/query/reranking/freshness.
Orchestration failure
Symptom:
- dropped subtask;
- loop lost state;
- wrong route.
Fix:
- contracts, reconciliation, trace, state.
Tool failure
Symptom:
- incorrect/failed external action.
Fix:
- tool implementation, auth, schema, retries.
7. Model Drift vs Data Drift vs Version Change
Model drift
Behavior shifts on stable inputs.
Data drift
Input distribution changes.
Version change
A model/prompt/index/tool version changed.
Different causes require different actions.
Do not prescribe a model change when the real variable is the document corpus.
8. A/B Testing
A credible experiment needs:
- hypothesis before running;
- randomized/sticky assignment;
- primary metric chosen in advance;
- secondary guardrail metrics;
- adequate sample size;
- predetermined acceptance rule.
Example:
"Prompt B will improve task success by at least 3 points without increasing p95 latency more than 10% or cost more than 5%."
9. Practical vs Statistical Significance
A statistically significant improvement can be useless.
Example:
- +0.2% accuracy;
- +80% cost;
- +50% latency.
Architect decision considers:
- business impact;
- operational cost;
- risk;
- confidence.
10. Shadow Testing
Use when live exposure is unacceptable.
Live traffic
|
+-> Current system -> user
|
+-> Candidate system -> offline evaluation
Benefits:
- realistic inputs;
- no user exposure.
Limit:
- cannot measure true user behavioral outcomes.
Good for regulated/high-risk pre-release validation.
11. Segment Results
Do not rely on one average.
Segment by:
- use case;
- customer group;
- risk tier;
- language;
- input length;
- document type;
- retrieval source;
- model route;
- tool;
- failure class.
Averages hide concentrated failure.
12. Tail Latency
Monitor:
- p50;
- p95;
- p99.
SLO breaches occur in the tail, not the median.
For agents, tail latency can be driven by:
- extra turns;
- retries;
- long context;
- slow tools;
- fallback path.
13. Cost Distribution
Average tokens can hide heavy-tail workloads.
Track per request:
- input tokens;
- output tokens;
- cache hit/write;
- number of calls;
- tool costs;
- retries;
- model route.
High-cost outliers often correlate with low-quality loops.
14. Optimization Order
A useful order:
- remove unnecessary calls;
- reduce irrelevant context;
- reduce tool output;
- cache stable prefixes;
- route simple tasks to cheaper model;
- shorten output;
- parallelize safe independent work;
- optimize retrieval;
- tune infrastructure.
Do not optimize price per token while ignoring redundant model calls.
15. Production Monitoring
Track:
- task success;
- safety violations;
- retrieval quality;
- tool error rate;
- fallback rate;
- human escalation;
- p95 latency;
- spend;
- cache hit rate;
- model mix;
- input distribution;
- drift.
Monitoring becomes useful only when tied to:
- thresholds;
- owners;
- actions.
Domain 4 Common Traps
- "Accurate" with no measurable threshold.
- Happy-path-only eval set.
- Human review used when code could grade exactly.
- Judge model uncalibrated.
- A/B primary metric chosen after seeing results.
- Median latency used as SLA proof.
- Average cost hides tail.
- Blaming model when retrieval changed.
- Model upgrade with no regression suite.
- Metrics collected with no decision/action attached.
Domain 4 Scenario Drills
Scenario 1
Answers become wrong immediately after a document refresh; model version did not change.
Best answer: Investigate ingestion/index/retrieval first.
Scenario 2
A prohibited tool call is either present or absent.
Best answer: Deterministic code-based evaluation.
Scenario 3
A candidate architecture cannot be exposed to users yet due to regulatory risk.
Best answer: Shadow test on copied live traffic.
Scenario 4
New prompt improves average quality but doubles p99 latency.
Best answer: Evaluate against pre-agreed secondary SLA guardrail, not quality alone.
Domain 4 Rapid Review
- Define success before building.
- Representative + edge + adversarial evals.
- Code-grade objective behavior.
- Judge subjective behavior; calibrate it.
- Diagnose failure class before fix.
- Segment results.
- A/B test with predeclared metrics.
- Shadow test when live exposure is risky.
- Optimize tails and per-request cost.
- Monitoring needs triggers and owners.
Domain 5 — Governance, Safety & Risk Management
Weight: 14%
The blueprint expects you to:
- implement guardrails and safety controls;
- identify LLM risks and failure modes;
- apply human validation strategically;
- map regulations to architecture;
- address bias, fairness, transparency, and responsible AI.
1. Safety Is an Architecture Input
Do not treat governance as paperwork added at the end.
Before design:
- classify data;
- identify regulated processes;
- define prohibited actions;
- define review obligations;
- define explainability needs;
- define audit evidence.
Architecture choices change based on these requirements.
2. Four Safety Layers
Trained model behavior
Broad safety baseline.
Does not know:
- your customer authorization rules;
- your internal policy;
- your retention requirement;
- your line-of-business restrictions.
System/application instructions
Guide role and behavior.
Not sufficient as hard enforcement.
Screening
Detects risky input/output.
Can be deterministic or model-based.
Action authorization
Determines whether a caller may perform a real-world action.
Should be deterministic and auditable.
3. Direct Prompt Injection
Malicious user input tries to override policy.
Mitigation layers:
- clear trusted instruction hierarchy;
- input controls;
- narrow tools;
- deterministic action checks;
- output validation.
4. Indirect Prompt Injection
Untrusted instructions arrive through:
- webpage;
- document;
- email;
- retrieval result;
- tool output.
This is especially important for agentic/RAG systems.
Rule: Retrieved text is data, not authority.
5. Data Exposure Risk
Risk exists even if Claude behaves correctly.
Data can leak into:
- prompt;
- tool result;
- trace;
- log;
- analytics;
- human review system.
Minimize fields before the model boundary.
6. Input, Output, and Action Controls
Input screening
Should request reach Claude?
Output screening
Should answer reach user?
Action authorization
May this side effect run?
These solve different problems.
Trap: Output filtering cannot undo an unauthorized action that already executed.
7. Deterministic vs Model-Based Safety Checks
Deterministic
Use for:
- allowlist;
- blocklist;
- schema;
- identity;
- scope;
- fixed threshold;
- known prohibited operation.
Model-based
Use for:
- nuanced harmful intent;
- policy interpretation;
- semantic risk classification.
Chain controls where necessary.
8. Fail Open vs Fail Closed
If a screening/authorization dependency fails:
Fail open
Request continues.
Use only where bypass is acceptable.
Fail closed
Request/action stops.
Use when uncontrolled execution could cause unacceptable harm.
Professional question: What is the safer failure direction for this control?
9. Supply-Chain Risk
Reusable:
- skills;
- packages;
- MCP servers;
- scripts;
- connectors;
can introduce malicious or excessive behavior.
Review for:
- network calls;
- filesystem access;
- shell execution;
- credential access;
- runtime downloads.
Use:
- trusted source policy;
- least privilege;
- sandboxing;
- version pinning;
- approval process.
10. Fairness
Bias can enter through:
- source corpus;
- prompt framing;
- examples;
- retrieval;
- routing;
- historical labels.
Evaluate outcomes by subgroup.
An aggregate metric can hide unequal harm.
11. Transparency
Different audiences need different explanations.
Affected user
Needs:
- understandable reason;
- relevant inputs;
- next step/appeal where appropriate.
Regulator/auditor
Needs:
- reproducible evidence;
- comparable-case consistency;
- controls;
- trace.
Engineering team
Needs:
- full technical trace;
- versions;
- tool path;
- data path.
12. Human Review by Stakes
Two key factors:
- reversibility;
- consequence.
Then consider:
- uncertainty/confidence;
- regulation;
- novelty.
Pre-action approval
Use for irreversible/high-impact actions.
Post-action audit
Use for reversible low-cost behavior.
Sampling
Use for ongoing quality monitoring.
13. Avoid Rubber-Stamping
Human oversight fails when:
- too many items;
- no useful context;
- approval UI hides evidence;
- reviewer cannot understand why item was flagged.
A reviewer needs:
- input;
- model output;
- evidence;
- risk reason;
- proposed action.
14. Regulation to Control Mapping
Do not answer compliance scenarios with:
"Use GDPR-compliant infrastructure."
Instead map obligation to:
Requirement
|
Control
|
Owner
|
Evidence
|
Review cadence
Example:
| Requirement | Control | Evidence |
|---|---|---|
| Data minimization | Redact fields before prompt | Payload sample / dataflow test |
| Access control | Server-side AuthZ | Access logs / policy test |
| Retention | TTL/delete job | Storage policy + deletion evidence |
| Auditability | Trace all tool writes | Queryable audit record |
15. Risk Register
A professional risk register includes:
- risk;
- likelihood;
- impact;
- preventive control;
- detective control;
- corrective action;
- owner;
- evidence;
- residual risk;
- review cadence.
Residual risk must be explicitly accepted by an accountable owner.
16. Responsible AI Questions
Ask:
- Who benefits?
- Who may be harmed?
- Can user challenge/appeal?
- Are outcomes unequal?
- Is automation proportionate?
- Is the system transparent enough?
- Are people aware when AI is involved where required?
- Is there a safe fallback?
Domain 5 Common Traps
- Policy statement mistaken for control.
- Prompt instruction used as authorization.
- Output filter used to protect an already-executed action.
- Fail-open default on critical safety service.
- No owner/evidence for compliance control.
- Human review applied to every trivial step.
- Aggregate fairness metric only.
- Reusable tool package trusted without audit.
- Sensitive data logged unnecessarily.
- Residual risk never assigned.
Domain 5 Scenario Drills
Scenario 1
An action authorization service is unavailable during a production write.
Best answer: For high-impact writes, fail closed unless policy explicitly allows otherwise.
Scenario 2
A PDF retrieved from an external repository tells Claude to send credentials to a URL.
Best answer: Treat document instructions as untrusted; tool/network permissions must prevent unauthorized action.
Scenario 3
Every output requires human approval, creating a 12-hour queue.
Best answer: Redesign review based on stakes/risk rather than universal approval.
Scenario 4
A compliance design states "we comply with retention requirements" but has no technical deletion mechanism.
Best answer: Requirement has not been converted into an operating control/evidence artifact.
Domain 5 Rapid Review
- Safety is architecture.
- Prompts are guidance, not hard enforcement.
- Input/output/action controls differ.
- Indirect injection comes through retrieved/tool content.
- Deterministic AuthZ.
- Choose fail-open/closed deliberately.
- Human review by consequence and reversibility.
- Compliance = requirement → control → owner → evidence.
- Fairness requires subgroup evaluation.
- Residual risk needs an owner.
Domain 6 — Stakeholder Communication & Lifecycle Management
Weight: 14%
The blueprint expects you to:
- conduct structured discovery;
- communicate architecture decisions and trade-offs;
- manage feedback and SLAs;
- document implementation guidance;
- support discovery → design → handoff → monitoring → iteration.
1. Discovery Is Structured Elicitation
Do not start architecture during the first five minutes of discovery.
Capture:
- goal;
- users;
- workflow today;
- pain;
- volume;
- latency;
- quality threshold;
- prohibited behavior;
- integration boundaries;
- data classification;
- cost constraints;
- compliance obligations;
- failure handling;
- owner;
- assumptions.
2. Convert Preferences Into Constraints
Stakeholder says:
"Fast."
Ask:
- p50 or p95?
- acceptable wait?
- streaming allowed?
- which step dominates?
- what happens on timeout?
Stakeholder says:
"Seamless."
Translate into:
- handoff latency;
- error state;
- hidden technical complexity;
- user reauthentication;
- human escalation path.
Rule: Adjectives are discovery prompts, not requirements.
3. Four Requirement Buckets
What must it do?
Functional outcomes.
What must it never do?
Forbidden actions.
What must it cost?
Latency, throughput, budget.
What must it prove?
Audit, compliance, evidence.
This framework surfaces hidden constraints early.
4. Requirements Traceability
Create rows linking:
Stakeholder statement
|
Underlying requirement
|
Architectural decision
|
Validation/evidence
Example:
| Stakeholder statement | Requirement | Design decision | Evidence |
|---|---|---|---|
| "No customer data leaves region" | Residency | Approved regional route | Architecture + provider config evidence |
| "Answers must be trustworthy" | Grounded answer threshold | RAG + abstention | Eval results |
| "Agents can refund" | Controlled write | AuthZ + approval + idempotency | Audit trace |
5. Architecture Trade-Off Communication
A strong trade-off statement contains:
- benefit;
- concession;
- reversal cost.
Example:
"Using a shared MCP service standardizes tools across three clients, but adds another service boundary and authorization layer. Reversing later would require each client to implement the APIs directly."
In regulated environments, add:
- compliance impact.
6. Architecture Decision Record
An ADR should include:
- decision;
- date;
- context;
- options considered;
- rejected alternatives;
- decision criteria;
- trade-offs;
- consequences;
- owner;
- review trigger.
The "why" matters as much as the diagram.
7. Communicating to Executives
Focus on:
- business outcome;
- risk;
- cost;
- timeline;
- major dependency;
- decision required.
Avoid opening with:
- vector embeddings;
- token windows;
- JSON schema.
8. Communicating to Engineers
Focus on:
- interfaces;
- data flow;
- contracts;
- failure handling;
- permissions;
- observability;
- deployment;
- test strategy.
9. Feedback Loop
Observability alone is not a feedback loop.
A loop is:
Signal
|
Trigger
|
Owner
|
Decision
|
Action
|
Review
Example:
| Signal | Trigger | Owner | Action |
|---|---|---|---|
| Grounding score | < 96% for 2 days | AI platform | Investigate retrieval |
| p95 latency | > 8s | SRE | Trace slow spans |
| Human escalation | > 20% | Product | Review task scope |
| Cost/request | > budget | Architect | Review model/context mix |
10. Scheduled Reviews
Not every control fires on an alert.
Schedule:
- output audits;
- access review;
- vendor compliance review;
- retrieval freshness review;
- model upgrade review;
- risk register review.
11. Lifecycle Phases
Discovery
- goals;
- baseline;
- requirements;
- constraints;
- ownership.
Design
- architecture;
- integrations;
- controls;
- eval plan;
- cost model.
Handoff / implementation guidance
- interfaces;
- ADRs;
- tests;
- rollout;
- operational ownership.
Monitoring
- metrics;
- alerts;
- risk;
- feedback.
Iteration
- eval updates;
- optimization;
- model/prompt changes;
- architecture refactoring.
Exam trap: An action may be good practice but belong to a later phase than the stem asks about.
12. Handoff Readiness
Before handoff:
- design decisions documented;
- interfaces stable enough;
- controls owned;
- eval criteria defined;
- failure modes identified;
- logging/monitoring specified;
- rollout/rollback agreed;
- open assumptions assigned.
13. Outcome Reporting
Executives need more than system metrics.
Report:
- use case;
- baseline;
- target;
- measured result;
- run cost;
- reviewer effort;
- risk/incident summary;
- scalability/next opportunity.
Capture baseline before implementation.
14. SLA / SLO Alignment
Possible commitments:
- availability;
- latency;
- quality;
- freshness;
- support response;
- maximum escalation time;
- recovery objective.
Architect should know:
- who owns each;
- how measured;
- what happens when breached.
15. Expectation Management
Explicitly state:
- what AI can do;
- what it cannot guarantee;
- where human review remains;
- known failure modes;
- acceptable uncertainty;
- cost/latency trade-offs.
Avoid overselling "automation."
Domain 6 Common Traps
- Architecture sketched before discovery is complete.
- Stakeholder adjectives treated as requirements.
- Trade-off presented without downside/reversal cost.
- Diagram without decision rationale.
- Metrics without triggers/owners.
- Baseline captured after deployment.
- Phase-order confusion.
- Compliance requirement discovered late.
- Human review cost omitted from business case.
Domain 6 Scenario Drills
Scenario 1
Sponsor asks for "instant" answers.
Best answer: Define measurable latency requirement and acceptable degraded/fallback behavior.
Scenario 2
Architect chooses MCP and presents only its technical benefits.
Best answer: Also communicate operational/compliance cost and reversal implications.
Scenario 3
A dashboard tracks quality but nobody is responsible for acting when it degrades.
Best answer: Add explicit trigger, owner, action, and review loop.
Scenario 4
The project wants to claim 50% productivity improvement but did not measure pre-project baseline.
Best answer: The outcome claim is weak; baseline must be measured consistently.
Domain 6 Rapid Review
- Discovery before design.
- Translate adjectives into measurable constraints.
- Must do / must not / must cost / must prove.
- Communicate benefit + concession + reversal cost.
- ADR explains why.
- Metrics need trigger + owner + action.
- Know lifecycle phase ordering.
- Baseline before intervention.
- Document for engineers, auditors, and future architects.
Domain 7 — Developer Productivity & Operational Enablement
Weight: 7%
The blueprint expects you to:
- configure Claude tools/environments for teams;
- improve developer workflows using AI-assisted tools;
- support debugging and operational issue resolution.
1. Shape vs Govern
Shape how the agent works
Examples:
- project instructions;
- skills;
- subagents;
- MCP servers;
- reusable prompts.
Govern what it may do
Examples:
- permissions;
- hooks;
- approval flows;
- sandbox;
- identity/authorization.
Rule: A guarantee belongs in enforcement, not in wording.
2. Team Baseline
A team environment should standardize:
- project instructions;
- approved tools;
- permission posture;
- model guidance;
- cost guidance;
- review expectations;
- testing workflow;
- support path.
Avoid each developer inventing a separate setup.
3. Claude Code Configuration Surfaces
CLAUDE.md
Persistent project guidance.
Settings
Behavior/configuration.
Permissions
Allow/ask/deny.
Hooks
Deterministic lifecycle actions.
Skills
Reusable procedures.
Subagents
Specialized isolated contexts.
MCP
External capabilities.
Know which surface matches the requirement.
4. Distribution Scope
Possible levels:
- organization;
- team/group;
- repository;
- explicit versioned programmatic dependency.
Trade-offs:
- reach;
- version control;
- rollback;
- governance.
Repository-scoped guidance is useful when behavior should version with code.
5. Model and Spend Guardrails
Team enablement should define:
- default model;
- approved models;
- when stronger reasoning is justified;
- cost caps;
- usage limits;
- observability.
Unmanaged model choice can create surprise spend.
6. Adoption Strategy
Good adoption:
- identify real workflow;
- pilot with champion;
- capture friction;
- package best practices;
- expand by team;
- measure usage and outcome.
Bad adoption:
- buy licenses;
- announce availability;
- assume value follows.
7. Avoid Basic-Chat Plateau
Value comes from integration with work:
- repository exploration;
- implementation;
- tests;
- review;
- debugging;
- incident investigation;
- repeatable skills.
Not just ad hoc chat.
8. Human Understanding
Generated code still needs:
- correctness review;
- security review;
- maintainability;
- explainability by owner.
A developer should understand what they ship.
9. Verification Checklist
Before AI-assisted code ships:
- tests pass;
- lint/static analysis;
- security checks;
- dependency review;
- behavior understood;
- edge cases covered;
- generated changes reviewed.
Automate deterministic checks.
10. Operational Troubleshooting
Quality declines gradually
Investigate:
- model/prompt version;
- index drift;
- data distribution.
Latency jumps
Investigate:
- context growth;
- cache hit loss;
- tool/dependency latency;
- extra agent turns.
Tool fails intermittently
Investigate:
- credential;
- rate limit;
- timeout;
- dependency;
- retry behavior.
Spend rises without traffic growth
Investigate:
- model mix;
- context size;
- cache hit rate;
- retries;
- agent-loop turns.
11. Runbooks
A runbook should map:
Symptom
|
Likely causes
|
Checks
|
Immediate action
|
Escalation owner
|
Recovery validation
Good architecture leaves behind operational leverage, not just a design diagram.
12. Support Ownership
Define:
- who owns model behavior;
- who owns retrieval;
- who owns tools;
- who owns security policy;
- who owns platform runtime;
- who owns business process.
Without ownership, every incident escalates to the architect.
Domain 7 Common Traps
- Team instructions used as security controls.
- Every developer configures tools differently.
- AI code ships without understanding/testing.
- Adoption measured only by seat count.
- No cost guidance.
- No runbooks.
- Architect remains permanent first-line support.
- Debugging starts by changing prompt without tracing symptom.
Domain 7 Scenario Drills
Scenario 1
Every developer has different project instructions and tool permissions.
Best answer: Establish a shared team/repository baseline.
Scenario 2
Spend rose with no increase in request count.
Best answer: Compare model route, context size, caching, retries, and loop behavior against baseline.
Scenario 3
AI-generated code passes unit tests but developer cannot explain what it does.
Best answer: Human understanding/review is still required before shipping.
Domain 7 Rapid Review
- Shape vs govern.
- Team baseline.
- CLAUDE.md guidance; permissions/hooks enforcement.
- Standardize tools and review.
- Measure adoption by workflow outcome, not seats.
- AI output still requires engineering discipline.
- Troubleshoot from evidence.
- Leave runbooks and ownership behind.
Cross-Domain Architecture Decision Matrix
| Scenario clue | Think first |
|---|---|
| vague business goal | discovery + measurable requirement |
| deterministic business rule | code/workflow |
| unknown runtime path | agent |
| many independent specialists | multi-agent only if isolation helps |
| incomplete worker results | reconciliation |
| static repeated prefix | prompt caching |
| live transactional state | system-of-record tool |
| exact IDs/codes | lexical retrieval |
| paraphrased knowledge search | semantic retrieval |
| mixed exact + semantic | hybrid retrieval |
| unnecessary tools | remove/scope |
| current user claims admin role in prompt | ignore claim; server identity/AuthZ |
| high-impact write | pre-action authorization/approval |
| backend tool transient failure | bounded retry at call boundary |
| dependency unhealthy | circuit breaker |
| durable async work | queue |
| candidate model change | regression eval |
| user cannot be exposed to candidate | shadow test |
| metric average looks fine | segment/tail analysis |
| compliance requirement | map to control + owner + evidence |
| "seamless" | translate to measurable constraint |
| dashboard metric no owner | build feedback loop |
| team tool setup varies | shared baseline |
| repeated incident | runbook |
Architectural Anti-Patterns
1. Agent Everywhere
Problem: Dynamic autonomy where control flow is known.
Why it fails: More cost, less observability, wider safety surface.
Better: Workflow until dynamic planning is actually required.
2. RAG as Database
Problem: Indexed snapshot used for live state.
Why it fails: Confidently stale answers.
Better: Source-of-truth API/tool.
3. Prompt as Security Boundary
Problem: "Never perform X" is the only protection.
Why it fails: Prompt injection and model variability.
Better: Deterministic authorization/permission/hook.
4. Monolithic Tool Catalog
Problem: Hundreds of tools always visible.
Why it fails: Context bloat, wrong selection, larger attack surface.
Better: Scope and progressive discovery.
5. Aggregate-Only Evals
Problem: Average score hides high-risk subgroup failure.
Better: Segment by category/risk/user/input.
6. Monitoring Without Operations
Problem: Data is collected but nobody acts.
Better: Trigger + owner + action + review.
7. Governance by Document
Problem: Policy exists but no control/evidence.
Better: Requirement → operating control → proof.
8. Unmeasured Business Case
Problem: Productivity claim with no baseline.
Better: Baseline before rollout, same metric after.
9. Human Review Everywhere
Problem: Consent fatigue and bottleneck.
Better: Route by risk/consequence.
10. Architect as Permanent Operator
Problem: Team cannot support system without architect.
Better: Shared tooling, observability, runbooks, ownership.
CCAR-P Exam Decision Framework
When reading a scenario, use this order.
Step 1 — Find the binding constraint
Look for:
- regulation;
- latency;
- cost;
- data freshness;
- side effect;
- authorization;
- input volume;
- quality threshold;
- team capability.
Step 2 — Identify the layer
Is this primarily:
- architecture;
- model/context;
- integration;
- evaluation;
- governance;
- stakeholder/lifecycle;
- enablement?
Step 3 — Find what changed
Ask:
- Did the model change?
- Did documents change?
- Did prompt change?
- Did traffic change?
- Did permissions change?
- Did tool fail?
This is especially important in debugging/eval questions.
Step 4 — Prefer structural fix
Remove cause before adding mitigation.
Examples:
- remove unnecessary tool instead of just log it;
- fix retrieval instead of strengthen prompt;
- durable queue instead of more retries;
- deterministic AuthZ instead of warning text.
Step 5 — Check lifecycle phase
Is the option correct now, or only later?
Step 6 — Check evidence
How would the architect prove the control or outcome works?
Practice Questions
These are original study questions based on the public CCAR-P blueprint. They are not real exam items.
Question 1
A customer service agent has read-only lookup tools plus an unused tool that can issue refunds. Only a separate specialist role ever needs refunds.
What is the strongest architectural action?
A. Add logging whenever the refund tool is selected B. Add a confirmation prompt C. Remove the refund tool from the customer service agent D. Use a stronger model
Answer: C
The structural fix is reducing capability and privilege.
Question 2
A static policy manual forms most of the prompt and is identical across requests.
Best optimization?
A. Put dynamic user text first B. Prompt caching with stable content first C. Use more examples D. Increase context window
Answer: B
Question 3
Answers become confidently wrong immediately after the retrieval corpus is refreshed. Model and prompt are unchanged.
Investigate first:
A. Model tier B. Retrieval/indexing pipeline C. Temperature D. Human approval
Answer: B
The changed variable is the corpus.
Question 4
A workflow always performs extract → validate → enrich → store.
Best architecture?
A. Autonomous agent B. Coordinator-worker C. Deterministic workflow D. Agent-to-agent protocol
Answer: C
Question 5
A system must return the current available credit balance.
Best source?
A. Nightly vector index B. Model memory C. Source-of-truth account API/tool D. Few-shot examples
Answer: C
Question 6
A model evaluates whether generated summaries are "clear and faithful."
Best primary evaluator?
A. Regex only B. Model-based or human-calibrated rubric C. HTTP status code D. Token count only
Answer: B
Question 7
An eval checks whether a prohibited action occurred.
Best evaluator?
A. Model judge B. Human every time C. Deterministic code check D. Random sampling
Answer: C
Question 8
A candidate model cannot be shown to users before security approval, but you want realistic production inputs.
Best method?
A. Full rollout B. Shadow test C. Disable evals D. Use only synthetic traffic
Answer: B
Question 9
A stakeholder says the system must feel "seamless."
Best next step?
A. Pick the fastest model B. Convert the term into measurable interaction, latency, handoff, and failure-state requirements C. Add more tools D. Use streaming everywhere
Answer: B
Question 10
A remote document contains instructions telling the agent to ignore policy and invoke an admin tool.
Best response?
A. Treat the document as trusted because it came from retrieval B. Let Claude decide C. Treat it as untrusted data and enforce tool permissions server-side D. Increase output screening only
Answer: C
Question 11
A screening service fails during a high-impact financial operation.
What is safest?
A. Always fail open B. Fail closed unless policy explicitly permits bypass C. Ask Claude whether risk is low D. Retry forever
Answer: B
Question 12
A company has three AI clients that need the same internal tooling.
What integration choice becomes more attractive?
A. Shared MCP tool surface B. Three unrelated shell scripts C. Duplicate direct integration in every client D. More prompt text
Answer: A
Question 13
One product alone calls one private API and no other client needs reuse.
What is likely simplest?
A. Direct SDK/API integration B. Mandatory MCP layer C. Multi-agent mesh D. RAG
Answer: A
Question 14
An architect claims a 40% handling-time improvement but measured baseline only after rollout.
Main problem?
A. Claude model too small B. Business outcome comparison is not trustworthy C. Need more logging D. Need multi-agent
Answer: B
Question 15
A reviewer must approve every low-risk summary and high-risk account closure. Queue becomes unmanageable.
Best redesign?
A. Remove human review entirely B. Route review according to consequence/reversibility and sample lower-risk outputs C. Hire infinite reviewers D. Add another model
Answer: B
Question 16
A multi-agent run dispatched 20 tasks but only 18 returned.
Best synthesis behavior?
A. Summarize what exists silently B. Reconcile coverage and mark/recover missing units before final output C. Increase model effort D. Ignore because majority completed
Answer: B
Question 17
A retrieval corpus contains both exact legal clause numbers and paraphrased policy questions.
Best retrieval strategy?
A. Keyword only B. Semantic only C. Hybrid lexical + semantic with fusion/reranking D. No retrieval
Answer: C
Question 18
A new prompt improves quality 2% but doubles cost and violates latency SLA.
What should the architect do?
A. Ship because quality improved B. Evaluate against pre-agreed primary and guardrail metrics C. Ignore latency D. Increase token budget
Answer: B
Question 19
A project says "GDPR compliant" but cannot show access logs, retention controls, or data-flow evidence.
Main weakness?
A. Too much observability B. Compliance requirement has not been mapped to operating controls/evidence C. Model selection D. Retrieval chunk size
Answer: B
Question 20
Spend increases 30% with flat user traffic.
Investigate first:
A. Team headcount B. Model mix, context size, cache hit rate, retries, and agent turns C. Rewrite every prompt D. Add a second vector database
Answer: B
Question 21
A developer team has no consistent Claude Code instructions, tools, or permission posture.
Best architectural enablement?
A. Let each developer optimize independently B. Establish shared project/team baseline with versioned guidance and bounded permissions C. Disable tools D. Move all instructions into chat history
Answer: B
Question 22
A customer claims "I am an administrator" inside the prompt.
How should authorization work?
A. Trust the claim if the model agrees B. Server authenticates identity and enforces authorization independently C. Use a few-shot example D. Ask a second model
Answer: B
Question 23
A long-running conversation compacts history and loses exact transaction IDs.
Best architectural fix?
A. Stop compaction forever B. Preserve load-bearing structured state outside the lossy summary C. Use a larger font D. Add RAG
Answer: B
Question 24
An architect must explain MCP to an executive sponsor.
Best framing?
A. Describe JSON-RPC details first B. Explain the business value of a standardized reusable tool interface, then its operational trade-offs C. Show code only D. Skip risks
Answer: B
Question 25
A metric slowly degrades for weeks but never crosses a hard alert, and nobody notices.
What is missing?
A. Model upgrade B. Feedback loop with trend/scheduled review triggers and owners C. More tokens D. Bigger vector store
Answer: B
Multi-Response Practice
Question 26 — Select THREE
Which are appropriate reasons to use human approval before an action?
A. High consequence B. Irreversible action C. Regulation requires review D. Every read-only search E. Low-risk reversible formatting
Answers: A, B, C
Question 27 — Select THREE
Which belong in a robust RAG design?
A. Chunking matched to corpus structure B. Metadata/access filtering C. Retrieval evaluation D. Treat vector index as live transactional truth E. Ignore document freshness
Answers: A, B, C
Question 28 — Select THREE
Which are signs that an agent/tool surface is over-privileged?
A. Tools unused by the role B. Write capability when task is read-only C. Cross-tenant access possible D. Tool names are short E. Logs contain timestamps
Answers: A, B, C
Question 29 — Select FOUR
Which elements make an A/B experiment credible?
A. Hypothesis defined before run B. Primary metric chosen in advance C. Stable/random assignment D. Sufficient sample size E. Choose the winning metric after results appear
Answers: A, B, C, D
Question 30 — Select THREE
Which are meaningful compliance evidence artifacts?
A. Access-control test result B. Audit log showing policy enforcement C. Data-retention deletion record D. "We are compliant" slide E. Informal verbal assurance
Answers: A, B, C
Question 31 — Select THREE
Which data should an architecture decision record capture?
A. Decision and rationale B. Rejected alternatives C. Trade-offs/consequences D. Every debug log line E. Secret API keys
Answers: A, B, C
Architecture Artifact Exercises
The Professional exam rewards judgment. Build these artifacts rather than memorizing definitions.
Exercise 1 — End-to-End Architecture Diagram
Design one production Claude solution.
Annotate every boundary with:
- trust level;
- owner;
- failure mode;
- fallback;
- observable signal;
- data classification.
Exercise 2 — Model Routing Table
| Task class | Model tier | Quality threshold | Latency budget | Cost ceiling | Fallback |
|---|---|---|---|---|---|
| Classification | Efficient | 98% | 1s | low | stronger model |
| Complex analysis | Strong | rubric ≥ target | 10s | medium | human |
| High-risk recommendation | Strong + validation | strict | 15s | high | human |
Use your own workloads.
Exercise 3 — RAG Threat Model
Document:
- corpus source;
- ingestion trust;
- access filtering;
- injection risk;
- stale-data risk;
- chunking;
- index freshness;
- retrieval evaluation;
- citation behavior;
- fallback.
Exercise 4 — Evaluation Matrix
| Failure class | Test cases | Metric | Threshold | Grader | Owner |
|---|---|---|---|---|---|
| Unsupported claim | 200 | grounded accuracy | ≥ 98% | model + human calibration | AI team |
| Schema invalid | all | parse success | 100% | code | app team |
| Tool misuse | adversarial | prohibited call | 0 | code | security |
Exercise 5 — Risk Register
For 10 risks, capture:
- likelihood;
- impact;
- preventive control;
- detective control;
- owner;
- evidence;
- residual risk;
- review date.
Exercise 6 — Architecture Decision Record
Write an ADR for:
Direct API vs MCP for internal tools.
Include:
- context;
- constraints;
- options;
- recommendation;
- trade-offs;
- reversal cost;
- security;
- operational ownership.
Exercise 7 — Executive vs Technical Presentation
Explain the same architecture twice.
Executive version — 5 minutes
- business outcome;
- cost;
- risk;
- timeline;
- decision needed.
Technical version
- components;
- interfaces;
- auth;
- data flow;
- failure handling;
- evals;
- observability.
Six-Week Study Plan
Week 1 — Solution Design & Architecture
Study:
- requirements decomposition;
- workflow vs agent;
- multi-agent;
- feasibility;
- service levels;
- business value.
Build:
- architecture diagram;
- one ADR.
Week 2 — Integration
Study:
- APIs/MCP;
- auth/AuthZ;
- tool scoping;
- RAG;
- hybrid retrieval;
- freshness;
- observability.
Build:
- RAG threat model;
- tool audit.
Week 3 — Evaluation & Optimization
Study:
- metrics;
- datasets;
- deterministic vs judge vs human;
- failure taxonomy;
- A/B;
- shadow testing;
- tail latency/cost.
Build:
- evaluation matrix;
- regression suite.
Week 4 — Governance & Risk
Study:
- prompt injection;
- safety layers;
- controls;
- fail-open/closed;
- human review;
- compliance;
- fairness;
- risk register.
Build:
- risk register;
- compliance evidence map.
Week 5 — Models, Prompting & Context
Study:
- model routing;
- context strategies;
- caching;
- templates;
- reuse;
- compaction.
Build:
- model routing table;
- cache-friendly prompt architecture.
Week 6 — Stakeholders + Developer Enablement
Study:
- discovery;
- ADRs;
- SLAs;
- lifecycle phases;
- feedback loops;
- team Claude Code baseline;
- runbooks.
Build:
- executive presentation;
- technical handoff;
- operational runbook.
Final two days:
- practice questions;
- domain rapid reviews;
- revisit Integration, Domain 1, and Evaluation first.
Night-Before CCAR-P Review
Start with the business outcome, not the model.
Deterministic business rules belong in deterministic code.
Known path → workflow. Dynamic path → agent.
Multi-agent only when specialization/context isolation earns the overhead.
Reconcile delegated work before synthesis.
Model changes are releases and need eval gates.
Context window ≠ retrieval ≠ state ≠ memory.
Stable repeated prefix → prompt caching.
Hard constraints eliminate integration options before cost/ergonomics.
AuthN says who. AuthZ says what they may do.
RAG for knowledge. System-of-record tools for live state.
Keyword exact, semantic paraphrase, hybrid mixed.
Progressive discovery beats monolithic tool/context exposure.
Define evaluation before production.
Use code graders where correctness is objective.
Segment results; averages hide failure.
Shadow test when live exposure is unacceptable.
Input screening, output screening, and action authorization are different controls.
Human review belongs where consequence and irreversibility justify it.
Compliance means control + owner + evidence, not policy prose.
Stakeholder adjectives are not requirements.
Metrics need triggers, owners, and actions.
Baseline must be measured before claiming improvement.
Team enablement needs shared configuration, verification, and runbooks.
60-Second Readiness Checklist
Domain 1
- [ ] Can I translate business goals into requirements?
- [ ] Can I distinguish workflow, agent, and multi-agent?
- [ ] Can I identify where deterministic rules belong?
- [ ] Can I explain feasibility constraints?
- [ ] Can I connect design to measurable business value?
Domain 2
- [ ] Can I justify model selection with evals?
- [ ] Can I explain context window vs retrieval vs state vs memory?
- [ ] Can I design a cache-friendly prompt?
- [ ] Can I decide zero-shot vs few-shot vs more reasoning?
- [ ] Can I manage context growth?
Domain 3
- [ ] Can I choose direct API vs MCP?
- [ ] Can I audit tool privilege?
- [ ] Can I separate AuthN/AuthZ/tenant isolation?
- [ ] Can I design RAG chunking and retrieval by corpus/query shape?
- [ ] Can I identify live data that should bypass RAG?
- [ ] Can I design end-to-end traces?
Domain 4
- [ ] Can I define measurable metrics?
- [ ] Can I build representative eval datasets?
- [ ] Can I choose code vs model vs human grader?
- [ ] Can I diagnose prompt/retrieval/model/orchestration failures?
- [ ] Can I run credible A/B or shadow tests?
- [ ] Can I reason about p95/p99 and cost tails?
Domain 5
- [ ] Can I identify direct/indirect injection?
- [ ] Can I place input/output/action controls correctly?
- [ ] Can I choose fail-open vs fail-closed?
- [ ] Can I route human review by stakes?
- [ ] Can I map regulation to control/owner/evidence?
- [ ] Can I explain residual risk?
Domain 6
- [ ] Can I run structured discovery?
- [ ] Can I translate vague words into measurable constraints?
- [ ] Can I communicate trade-offs and reversal cost?
- [ ] Can I write an ADR?
- [ ] Can I design feedback loops?
- [ ] Can I distinguish lifecycle phases?
Domain 7
- [ ] Can I distinguish guidance from enforcement?
- [ ] Can I standardize Claude tooling for a team?
- [ ] Can I define verification gates?
- [ ] Can I troubleshoot quality, latency, tool, and cost symptoms?
- [ ] Can I leave behind runbooks and ownership?
Current-Product Details to Verify Near Exam Day
These can evolve quickly:
- model lineup;
- context-window sizes;
- prompt-caching parameters;
- thinking/effort controls;
- tool search/deferred loading behavior;
- Claude Code settings/hook surface;
- Agent SDK APIs;
- MCP lifecycle details;
- enterprise delivery options;
- pricing;
- rate limits.
Use the current official documentation for those details, but keep the architecture principles in this guide as the stable study framework.
Official / Primary References
Certification
- Anthropic certification program: https://claude.com/blog/four-role-based-claude-certifications
- Anthropic Partner Academy certifications: https://anthropic-partners.skilljar.com/page/partner-certifications
- Anthropic certification FAQ: https://anthropic-partners.skilljar.com/page/faq-certifications
Claude Platform
- Documentation home: https://platform.claude.com/docs/
- Models: https://platform.claude.com/docs/en/about-claude/models/overview
- Model selection: https://platform.claude.com/docs/en/about-claude/models/choosing-a-model
- Prompt caching: https://platform.claude.com/docs/en/build-with-claude/prompt-caching
- Tool use: https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview
- Tool search: https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool
- Streaming: https://platform.claude.com/docs/en/build-with-claude/streaming
- Message Batches: https://platform.claude.com/docs/en/build-with-claude/batch-processing
- Rate limits: https://platform.claude.com/docs/en/api/rate-limits
- API errors: https://platform.claude.com/docs/en/api/errors
- Prompt-injection mitigation: https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/mitigate-jailbreaks
Claude Code / Agent SDK
- Claude Code docs: https://code.claude.com/docs/
- Settings: https://code.claude.com/docs/en/settings
- Hooks: https://code.claude.com/docs/en/hooks
- Subagents: https://code.claude.com/docs/en/sub-agents
- Agent SDK: https://code.claude.com/docs/en/agent-sdk/overview
MCP
- Model Context Protocol: https://modelcontextprotocol.io/
Final Exam Strategy
Prioritize by weight and by how much professional judgment each domain requires:
- Integration — 19%
- Solution Design & Architecture — 17%
- Evaluation, Testing & Optimization — 16%
- Governance, Safety & Risk Management — 14%
- Stakeholder Communication & Lifecycle Management — 14%
- Claude Models, Prompting & Context Engineering — 13%
- Developer Productivity & Operational Enablement — 7%
The recurring Professional-level principle is:
The architect's job is not to make Claude capable of doing everything. It is to assign each responsibility to the component best suited to own it, constrain the dangerous paths, prove the system meets its business outcome, and leave behind evidence and operational ownership.
If two answer choices both sound reasonable, prefer the one that:
- addresses the constraint named in the scenario;
- fixes the cause rather than only reducing the symptom;
- keeps deterministic guarantees outside probabilistic behavior;
- makes the system more observable and auditable;
- can be defended against a measurable requirement.

