COMPLETE PRINTABLE EDITION · 2026
Senior GenAI
Interview Handbook
A 16-chapter preparation book for Senior GenAI, AI Platform, GenAI Solutions Architect, and Forward Deployed roles on AWS and Google Cloud.
PREPARED FOR PURNENDU DASCHAPTER 01 · FOUNDATION
The Preparation Strategy
27 min read · 14 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- Turn a senior GenAI job description — Bedrock, Vertex AI, RAG, agents, evals, LLMOps — into a ranked preparation backlog instead of studying the whole AI field.
- Decode recurring GenAI JD language into the hiring signal behind it, and map every signal to honest evidence, a gap-closing artifact, or an explicit learning plan.
- Prioritize preparation work on an effort-versus-signal grid, and defend why eval harnesses and cost models beat certificate cramming.
- Assemble a five-artifact GenAI proof stack — eval harness, RAG benchmark, guarded agent demo, inference-cost analysis, and a two-cloud reference build.
- Construct senior answers that connect requirements, design choices, failure modes, token economics, operations, and outcomes.
- Run a weekly preparation cadence with timed rehearsal across design, coding, troubleshooting, behavioral, and written formats.
- Calibrate every claim and example metric so no practice result is ever presented as personal production experience.
1. Preparation is evidence selection
This book is deliberately not a survey of artificial intelligence. It is a conversion plan: turn existing production-engineering competence into evidence for Senior GenAI, AI Platform, GenAI Solutions Architect, and Forward Deployed interviews on AWS and GCP. The highest-return work is therefore not memorizing more model names. It is choosing which claims you want an interviewer to believe and constructing honest, inspectable support for each claim.
Use a three-layer model. A signal is what the role needs to infer — for example, "can decide whether a RAG change is safe to ship" or "can keep an agent's blast radius bounded." Evidence is what makes that inference reasonable — an eval harness with a CI gate, a retrieval benchmark, an incident analysis, a cost model, or an accurately described past decision. A delivery is how the evidence appears in the interview: a two-minute story, a whiteboard design, a working repository, or a written memo. Weak preparation jumps from a topic name straight to a rehearsed explanation. Strong preparation joins all three layers and lets feedback flow backwards.
flowchart LR
JD["Live GenAI job description"] --> SIG["Required hiring signal"]
SIG --> EV["Proof artifact (honest, inspectable)"]
EV --> DEL["Interview delivery (story, design, repo, memo)"]
DEL --> FB["Feedback and exposed gaps"]
FB --> SIG
The five-part readiness test
For every priority topic, ask whether you can pass all five tests below. A topic is not interview-ready until each is plausible. This kills the common illusion that reading Bedrock or Vertex documentation equals being able to reason in a live system-design conversation.
- Define — state the concept precisely, including its unit of measure (recall@k, p95 TTFT, cost per successful task).
- Design — place it in an architecture under stated quality, latency, cost, and security constraints.
- Implement — build a small honest version: a harness, a benchmark, a guarded tool call.
- Debug — inject a deliberate failure (bad chunking, poisoned tool output, judge drift) and diagnose it out loud.
- Defend — argue one consequential trade-off and name the reversal condition that would change your mind.
2. Decode the job description into a signal matrix
GenAI job descriptions in 2026 are keyword-dense but highly decodable. "Hands-on with Amazon Bedrock," "productionize LLM applications on Vertex AI," "design agentic workflows," and "establish evaluation frameworks" are not synonyms — each phrase encodes a different question the interview loop must resolve. Extract the verbs and objects from two or three live postings, then build the matrix below before you study anything. Re-check postings at application time; GenAI requirements drift quarter to quarter.
| JD phrase (verbatim pattern) | Question the interviewer is resolving | Best evidence form | Weak substitute | Deep dive |
|---|---|---|---|---|
| "Hands-on with Amazon Bedrock / Vertex AI" | Can they make model-selection, throughput, guardrail, and cost decisions on a managed platform, not just call an API? | Two-cloud reference build with a decision memo on model choice, pricing mode, and quotas | Console screenshots or a certification badge alone | Ch. 08–09 |
| "Production RAG / grounded generation" | Can they diagnose retrieval quality and latency separately from generation quality? | Golden question set, hybrid-versus-dense benchmark with nDCG/recall and p95 latency, chunking ablation | A framework quickstart that answers five demo questions | Ch. 04 |
| "Design agentic workflows / multi-step automation" | Can they bound autonomy: idempotent tools, approval boundaries, recovery from partial failure? | Agent demo with an explicit state machine, guardrails, budget caps, and a replayed failure transcript | A happy-path chatbot that calls one tool | Ch. 05 |
| "Establish evaluation frameworks / LLM quality" | Can they decide whether a change is safe to ship? | Versioned eval set, LLM-judge calibrated against human labels, slice analysis, CI gate that has actually blocked a change | One aggregate "accuracy" number from a leaderboard | Ch. 06 |
| "LLMOps / observability / reliability" | Can they operate model-backed systems like production software? | Traces with token-level spans, drift and safety alarms, rollback plan, runbook, SLOs | An architecture diagram with no failure path | Ch. 11 |
| "Optimize inference cost / scale efficiently" | Do they reason in dollars per successful task, not tokens per request? | Cost model comparing on-demand tokens, provisioned throughput, caching, and model routing, with a stated denominator | "We switched to a cheaper model" | Ch. 02 |
| "Fine-tuning / PEFT / model adaptation" | Do they know when adaptation beats prompting and RAG — and when it does not? | Adaptation decision memo plus a small LoRA run with a before/after eval delta and cost accounting | Name-dropping LoRA and QLoRA | Ch. 03 |
| "Customer-facing / forward deployed / solutions" | Can they turn customer ambiguity into a safe, phased rollout that shows value early? | Discovery memo, assumptions log, pilot plan with adoption metrics and stop conditions | A premature product pitch | Ch. 12 |
Label the evidence honestly
Mark each cell experienced, built for practice, understood but not operated, or gap. Those labels are valuable in the room. A credible statement such as "I have not operated Bedrock provisioned throughput at enterprise scale; here is the cost model and load test I built to learn the trade-offs" is stronger than an inflated production claim — and it survives follow-up probes, which inflated claims never do. Never present a tutorial metric, a synthetic benchmark, a team result, or a hypothetical design as your personal production outcome.
3. The 16-chapter curriculum as a preparation map
The rest of this handbook is organized so that the signal matrix above has a place to send you. Do not read it linearly like a textbook. Anchor on the chapters your matrix flags as high-frequency and weak-evidence; treat the others as reference material with a small breadth budget. The mindmap below is the territory; your matrix is the route.
mindmap
root(("GenAI prep map"))
Foundations
02 LLM internals and serving
03 Prompting and adaptation
Applications
04 Retrieval and production RAG
05 Agentic systems
06 Evaluation and quality
Platform
07 Enterprise integration
10 Data and cloud platform
11 LLMOps and security
Clouds
08 AWS Bedrock and SageMaker
09 GCP Vertex AI and Gemini
Synthesis
12 System design and FDE
13 Leadership and coding
Execution
14 Role to topic map
15 Twelve week sequence
16 References and readiness
A useful traversal for most senior GenAI loops follows the spine below: decode signals first, refresh only the foundations those signals need, build artifacts in the application chapters, map them onto both clouds, then rehearse synthesis. Chapter 15 turns this spine into a full 12-week execution sequence, so this chapter deliberately stays at the strategy layer.
4. Prioritize by effort versus hiring signal
Preparation time is the scarce resource, so manage it as a portfolio. Score each candidate activity on two axes: how much build effort it demands, and how much hiring signal it generates for your target roles. The quadrant chart below positions the artifacts this chapter recommends against common low-signal temptations. Positions are judgments to be re-derived from your own matrix, not universal constants.
quadrantChart
title Effort versus hiring signal
x-axis Low build effort --> High build effort
y-axis Low hiring signal --> High hiring signal
quadrant-1 Flagship bets
quadrant-2 Quick wins
quadrant-3 Parking lot
quadrant-4 Money pits
Eval harness with CI gate: [0.55, 0.92]
RAG benchmark report: [0.45, 0.85]
Guarded agent demo: [0.68, 0.8]
Inference cost model: [0.28, 0.75]
Two cloud reference build: [0.82, 0.72]
Story ledger rewrite: [0.2, 0.62]
Leaderboard trivia tracking: [0.15, 0.12]
Pretraining an LLM from scratch: [0.92, 0.18]
Certificate cramming: [0.55, 0.28]
To turn the picture into a schedule, use a deliberately simple priority product. Its purpose is to expose opportunity cost, not to be precise: if a role repeatedly mentions Bedrock, RAG quality, and evaluation, another evening on pretraining mathematics is a poor trade unless the posting explicitly asks for it.
priority(topic) = role_frequency × evidence_gap × interview_urgency
# Example only: each factor scored 1 (low) to 3 (high)
evals_and_rag_quality = 3 × 2 × 3 = 18
agent_guardrails = 2 × 3 × 2 = 12
pretraining_theory = 1 × 2 × 1 = 2
pie title Default preparation weights in percent
"Retrieval and production RAG" : 15
"Evaluation and quality" : 15
"Cloud stacks AWS and GCP" : 15
"System design and FDE" : 15
"Agentic systems" : 10
"Serving and token economics" : 10
"LLMOps and security" : 10
"Leadership and coding" : 10
Deepen
Pick one or two recurring, high-value skills where you already have foundation — for most 2026 GenAI loops, evaluation and retrieval quality are the prime candidates because they anchor every ship decision.
Repair
Pick the risk that can sink a loop outright: vague token-economics reasoning, no agent failure story, weak SQL, or no answer to "how did you know it was safe to ship?" Practise it in small, timed units.
Maintain
Keep existing strengths fluent through spaced recall. Do not rebuild familiar Python, FastAPI, or core cloud networking from zero; those are covered as refreshers in Chapters 07 and 10.
Defer
Record fascinating but low-signal topics — new model gossip, exotic architectures — in a dated parking lot with a revisit trigger. Deferral is a strategy decision, not a value judgment.
5. The GenAI proof stack: five artifacts, one scenario
A proof stack is a small set of artifacts that covers many signals without becoming a portfolio museum. The efficient move is to build all five artifacts around one shared scenario — for example, an enterprise document assistant with retrieval, tool use, and strict tenancy rules. The shared scenario lets you reason deeply in any interview instead of maintaining unrelated demos, and every artifact feeds the same evidence dossier.
flowchart TD
SCN["Shared scenario (enterprise doc assistant)"] --> EVH["Eval harness with CI gate"]
SCN --> RAGB["RAG benchmark (hybrid vs dense)"]
SCN --> AGD["Agent demo with guardrails"]
SCN --> COST["Inference cost analysis"]
SCN --> REF["Two-cloud reference build"]
EVH --> DOSS["Evidence dossier + design memos"]
RAGB --> DOSS
AGD --> DOSS
COST --> DOSS
REF --> DOSS
| Artifact | Signals it covers | Core metrics it must carry | Built in |
|---|---|---|---|
| Eval harness — versioned dataset, LLM-judge calibrated against your own human labels (see Zheng et al., LLM-as-a-Judge), CI gate | Evaluation judgment, ship/no-ship discipline, LLMOps | Task success rate with slices, judge–human agreement, gate pass/fail history | Ch. 06 |
| RAG benchmark — golden set, chunking and hybrid-retrieval ablations over the classic retrieval-augmented pattern (Lewis et al.) | Retrieval depth, grounded-generation quality | Recall@k, nDCG@10, faithfulness, p95 retrieval latency | Ch. 04 |
| Agent demo with guardrails — explicit state machine, idempotent tools, budget caps, approval boundary, replayable failure transcript (patterns in Anthropic's building-effective-agents guidance) | Agent reliability, bounded autonomy, safety | Task completion rate, tool-error recovery rate, guardrail-violation count | Ch. 05 |
| Inference-cost analysis — tokens × price versus provisioned throughput versus caching and routing, expressed per successful task | Token economics, FinOps credibility | Cost per successful task, cache hit rate, break-even utilization for provisioned capacity | Ch. 02 |
| Cloud reference build — the same scenario deployed thin on both clouds with IAM, logging, and a rollback path | Platform ownership, AWS and GCP fluency | Deployment reproducibility, p95 end-to-end latency, monthly cost estimate | Ch. 08–09 |
AWS
- Amazon Bedrockmanaged FM inference, guardrails, agents
- Bedrock Knowledge Basesmanaged RAG over your corpus
- OpenSearch Serverlessvector plus keyword hybrid retrieval
- Lambda + API Gatewayserving glue and tool endpoints
- CloudWatch + Cost Explorertraces, alarms, token spend
Google Cloud
- Vertex AIGemini models, Model Garden, tuning
- Vertex AI Search / RAG Enginemanaged grounding and retrieval
- Vertex AI Vector SearchANN retrieval at scale
- Cloud Runserving glue and tool endpoints
- Cloud Monitoring + BigQuerytraces, alarms, spend analysis
Keep a story ledger beside the artifact stack
Artifacts prove capability; the story ledger proves experience. For every real story, record: situation and stakes; exact personal responsibility; collaborators; the alternatives considered; the action; the observable result; what you would change; and which claims need qualification. Replace confidential names and values only when necessary, and say that values are rounded or anonymized. If you lack a metric, state what was observed and what you would measure now — never reverse-engineer an impressive number.
6. Answer at senior scope
Senior answers reveal a decision process. Begin by locating the goal, users, risk, and constraints. Establish a simple baseline. Decompose the system and make boundaries explicit. Compare alternatives against criteria. Cover failure, security, rollout, observability, and ownership. Close with validation and what would change your mind. The same sequence works in GenAI system design, troubleshooting, project deep dives, and customer discovery — Chapter 12 drills it against full design prompts.
flowchart LR
G["Goal, users, risk"] --> C["Constraints (quality, latency, cost, privacy)"]
C --> B["Smallest credible baseline"]
B --> T["Alternatives and trade-offs"]
T --> F["Failure modes and security"]
F --> R["Rollout and ownership"]
R --> M["Measurement and reversal conditions"]
Separate facts, assumptions, and decisions
Say "the requirement states," "I am assuming," and "I would choose" rather than blending all three. Ask a small number of high-value clarifying questions, then proceed with declared assumptions. In a 45-minute design session, exhaustive discovery is impossible; the signal is whether your assumptions are consequential and whether the design can absorb being wrong about them.
Make influence observable
Staff-level influence is a mechanism, not a title: you wrote an options memo, instrumented a disputed bottleneck, ran the eval that settled a model-choice argument, created an adoption path, mentored an owner, or changed a standard. Explain the resistance and how feedback altered the plan. "I convinced everyone" is weaker than a transparent decision process that let reasonable people converge. Chapter 13 covers the delivery mechanics.
7. The preparation operating cadence
Preparation needs short feedback loops, and GenAI preparation needs them more than most because the vendor surface moves monthly while the underlying judgment moves slowly. Run a weekly cycle that always ends in deliverables — an artifact increment, a recorded mock, a rewritten story — never in "hours studied."
- Monday — diagnoseRe-scan two live postings; re-rank the backlog with the priority product; pick the week's one artifact increment and one repair item.
- Tuesday–Thursday — buildDeep-work blocks on the proof stack: extend the eval harness, run one benchmark ablation, harden one agent tool, or update the cost model with current published pricing.
- Friday — rehearseOne timed mock in a rotating format: design, coding, troubleshooting, behavioral, or a written memo. Record it.
- Weekend — review and updateScore the recording against a rubric; extract one or two observable corrections; update the signal matrix, story ledger, and parking lot.
| Format | What to practise | Evidence to capture | Useful review question |
|---|---|---|---|
| GenAI system design | Ambiguity, scale, retrieval and model choice, guardrails, rollout | Final diagram and decision log | Did constraints actually drive the architecture? |
| Technical deep dive | One real or practice system, including a failure | Recording, claim ledger, missing-detail list | Did I distinguish my work from the team's? |
| Coding | Correctness, tests, streaming and async patterns, complexity | Timed solution and post-review patch | Did I validate edge cases before optimizing? |
| Troubleshooting | Hypotheses, instrumentation, isolating retrieval vs generation vs infra | Incident timeline and next experiment | Did I change one variable at a time? |
| Behavioral | Ownership, conflict, failure, influence mechanism | Two-minute and six-minute versions | Was the action mechanism concrete? |
| Written memo | Concise asynchronous decision-making | One page with a requested decision | Can a reader act without a meeting? |
Use a stop-doing list
- Stop adding notes once a topic passes the five-part readiness test; move to the next gap.
- Stop polishing demo UX while the artifact lacks a baseline and a metric definition.
- Stop tracking weekly model-release gossip unless a target role requires vendor-frontier depth; log it in the parking lot instead.
- Stop rehearsing only successful stories — failure analysis is where operational judgment becomes visible.
8. Make every outcome measurable
Senior GenAI interviews repeatedly probe evidence across quality, latency, cost, safety, reliability, and customer impact. Treat these as a balanced evidence set rather than quoting whichever number looks best. Each dimension has a defensible form and a weak substitute, and interviewers are well calibrated to the difference. The AWS Well-Architected Generative AI Lens uses essentially this decomposition for production reviews, which makes it a useful shared vocabulary in AWS-facing loops.
| Dimension | Defensible evidence | Weak substitute |
|---|---|---|
| Quality | Versioned task set, metric definition, slices, baseline, judge calibration | One impressive response |
| Latency | p50/p95/p99 TTFT and end-to-end at stated concurrency and boundary | An unqualified average |
| Cost | Provider plus platform cost per successful task, with cache and routing effects | Token count alone |
| Safety | Guardrail-violation rate on an adversarial set, escalation path, audit trail | "We enabled the guardrail feature" |
| Reliability | Availability or successful-task SLI over a declared window, with rollback evidence | "It was stable" |
| Customer impact | Adoption, time saved, error reduction, or business result with provenance | Feature shipped |
If exact production values are confidential, use approved ranges or explain the measurement method and direction. Never invent precision. The senior signal is the causal chain from decision to measurable outcome — and an honest account of confounders.
Interview playbook
Use the CONTEXT answer frame when a question is broad:
- C — Customer and consequence: who needs what, and what happens if the system is wrong, unsafe, or late?
- O — Objectives and constraints: define quality, latency, cost per task, privacy, scale, and timeline.
- N — Narrow baseline: the smallest system that tests the value proposition — often managed RAG before custom anything.
- T — Trade-offs and topology: draw boundaries, compare alternatives (managed versus custom, prompt versus tune, single model versus router), justify the choice.
- E — Exceptions: dependency failure, bad retrieval, prompt injection, abusive input, partial agent completion.
- X — eXecution: sequencing, ownership, migration, approvals, and communication.
- T — Tests and telemetry: eval gates, SLOs, guardrail monitoring, feedback loops, and reversal conditions.
Common traps
- Model-first: naming a model, vector database, or agent framework before defining the problem and constraints.
- Unbounded "we": making team output sound like personal implementation. State your role explicitly.
- Metric theater: quoting an eval score without dataset, denominator, baseline, judge calibration, or business meaning.
- Perfect hindsight: omitting uncertainty, disagreement, or what changed during delivery.
- Demo scope: ignoring authorization, tenant isolation, retries, evaluation, rollout, or on-call ownership.
- Vendor recital: listing Bedrock or Vertex features instead of criteria, measurement, and failure behavior.
For experience questions, use Situation → Stakes → My responsibility → Options → Action → Result → Reflection. If a result is qualitative, keep it qualitative. If discussing a portfolio exercise, label it as such before presenting any numbers.
Question bank
Practise aloud. Each answer should use evidence appropriate to its claim and should survive the probes without invented detail.
Q1Why does your preparation prioritize evaluation and retrieval over foundation-model training?
Strong answer outline
- Anchor to recurring responsibilities in live senior GenAI postings: ship decisions, RAG quality, platform operations — not pretraining.
- Explain comparative advantage: production engineering background converts fastest into eval and retrieval evidence.
- Describe the opportunity cost and a small breadth budget for serving internals (Chapter 02) so model reasoning stays credible.
Follow-up probes
- What job-description change would make you rebalance?
- What concrete artifact will the prioritized time produce?
Pass if the answer connects role frequency, honest evidence gaps, and a named deliverable; fail if it dismisses model fundamentals categorically.
Q2Walk me through how you would decode a GenAI job description into a study plan.
Strong answer outline
- Extract verbs and objects — "operate," "evaluate," "advise," "Bedrock," "agentic" — from two or three live postings.
- Map each phrase to the hiring signal behind it and label current evidence: experienced, practiced, understood, or gap.
- Rank by role frequency × evidence gap × urgency; commit to weekly deliverables, not hours.
Follow-up probes
- Show me one real phrase you decoded and what it changed.
- What did you explicitly defer, and what is the revisit trigger?
Pass if a real opportunity-cost decision is named; fail if the plan is "cover everything important."
Q3What counts as proof of production-RAG competence?
Strong answer outline
- Name four layers: code proof, decision proof (ADR), measurement proof (golden set, ablations), operations proof (traces, rollback).
- Separate retrieval metrics (recall@k, nDCG) from generation metrics (faithfulness, task success) and system metrics (p95, cost per task).
- Include a failure injection — bad chunking or index drift — and the degradation path.
Follow-up probes
- Which artifact would you show first and why?
- What can a small portfolio project not prove?
Pass if limitations are explicit and the artifact is inspectable; fail if a screenshot of a good answer is treated as production evidence.
Q4You have more depth on one cloud than the other. The role is on the weaker cloud. How do you prepare and how do you answer?
Strong answer outline
- State that platform judgment transfers: managed FM inference, vector retrieval, IAM boundaries, and observability exist on both clouds.
- Show the mapping concretely — Bedrock ↔ Vertex AI, Knowledge Bases ↔ Vertex AI Search, OpenSearch ↔ Vector Search — and where the mapping breaks (pricing modes, quota models, guardrail features).
- Present the two-cloud reference build as the gap-closing artifact and name what you still have not operated.
Follow-up probes
- Which service pair has the most misleading equivalence?
- What would you validate in your first week on the weaker cloud?
Pass if transfer is argued mechanism-by-mechanism with honest gaps; fail if the answer is "clouds are basically the same."
Q5How do you present a benchmark without overstating it?
Strong answer outline
- Label the environment: production, anonymized production, or synthetic practice.
- State corpus, queries, judgment provenance, metric, baseline, hardware, and run conditions.
- Show trade-offs, variance across runs, limitations, and the next test needed before a real decision.
Follow-up probes
- Is the gain statistically or just numerically visible?
- Could there be leakage between your golden set and your tuning loop?
Pass if another engineer could interpret and challenge the result; fail if only a favorable percentage survives retelling.
Q6Which single proof artifact would you build first for these roles, and why?
Strong answer outline
- Argue for the eval harness: it is the artifact every other artifact depends on for credibility — benchmarks, agent claims, and cost trade-offs all need a quality denominator.
- Describe its parts: versioned dataset, calibrated judge, slices, CI gate, and a change it actually blocked.
- Show how it compounds: the RAG benchmark and agent demo reuse the same harness.
Follow-up probes
- How do you know your LLM judge is trustworthy?
- What would make you build the cost model first instead?
Pass if the choice is justified by dependency structure, not fashion; fail if the answer is "the most impressive demo."
Q7You have never run Bedrock Agents or Vertex AI Agent Builder in production. The JD asks for agentic experience. How do you answer?
Strong answer outline
- Say so directly, then separate transferable agent engineering — state machines, idempotent tools, budgets, approval boundaries — from vendor-specific operations.
- Present the guarded agent demo: injected tool failure, replayed transcript, guardrail-violation count.
- Name the operating claims you cannot make (quota behavior at scale, real user abuse patterns) and how you would validate them.
Follow-up probes
- Which agent failure mode worries you most in production?
- What would you test in week one with real tenant data?
Pass if honesty is paired with relevant depth and a learning plan; fail if a demo is relabeled as production ownership.
Q8How do you prepare for an ambiguous GenAI system-design prompt?
Strong answer outline
- Identify user, value, scale, quality bar, latency budget, cost ceiling, privacy, and failure consequence.
- Ask only the questions that materially change the design, then declare assumptions and proceed.
- Start with a managed baseline, evolve it at pressure points, and reserve time for guardrails, rollout, and measurement.
Follow-up probes
- Which assumption in your last mock was most dangerous?
- How would your design change at ten times the traffic — or one tenth the budget?
Pass if ambiguity becomes explicit decisions; fail if questioning consumes the session or everything is silently assumed.
Q9How do you choose metrics for a GenAI proof artifact?
Strong answer outline
- Trace user value to component metrics: retrieval quality, generation faithfulness, guardrail violations, latency percentiles, cost per successful task.
- Always pair a primary metric with a counter-metric — quality with latency, cost with success rate.
- Define slices and acceptance thresholds before running the favored variant.
Follow-up probes
- Which of your metrics can be gamed, and how?
- What is the denominator of "cost per successful task" in your artifact?
Pass if metrics have definitions and decision consequences; fail if a dashboard is mistaken for a quality model.
Q10Show me how you reason about token economics for a workload.
Strong answer outline
- Start from traffic shape: requests per day, input/output token distribution, cacheable prefix share, latency tolerance.
- Compare on-demand per-token pricing, batch, and provisioned throughput on both clouds; compute the utilization break-even, citing current published pricing rather than memory.
- Add levers in order of leverage: prompt and context reduction, caching, model routing, then capacity commitments — each with its quality counter-metric.
Follow-up probes
- When does provisioned throughput lose to on-demand despite high volume?
- How does an aggressive router protect quality?
Pass if the unit of analysis is dollars per successful task with stated assumptions; fail if the answer is a cheaper-model reflex. Depth lives in Chapter 02.
Q11Models and vendor features change monthly. How do you keep preparation current without chasing news?
Strong answer outline
- Split knowledge into durable (retrieval math, eval design, failure modes, cost reasoning) and volatile (model names, prices, feature flags).
- Invest build time in durable artifacts; refresh volatile facts from primary docs in a short weekly slot, just before interviews.
- Keep a parking lot with revisit triggers so novelty does not hijack the backlog.
Follow-up probes
- Name one durable principle that survived the last two model generations.
- What volatile fact did you re-check this week?
Pass if the durable/volatile split is explicit and primary sources are named; fail if currency means reading headlines.
Q12What makes a GenAI project story senior rather than merely trendy?
Strong answer outline
- Frame the consequential decision and its uncertainty — model choice, managed versus custom, ship gate — not the novelty of the stack.
- Clarify personal ownership, alternatives, stakeholder influence, and operational follow-through after launch.
- State result and reflection, including a decision you would now revise.
Follow-up probes
- Who disagreed with the decision and why?
- Who operated the system after launch, and what paged them?
Pass if the story reveals judgment and leverage; fail if seniority is implied by using this year's tools.
Q13What should you do when a mock interview goes badly?
Strong answer outline
- Separate knowledge gaps from delivery, structure, and time-management failures — they need different fixes.
- Choose one observable behavior correction and one technical correction; write both down.
- Re-answer the same prompt after a delay and compare against the rubric.
Follow-up probes
- What evidence would show the correction worked?
- When do you seek external critique instead of self-review?
Pass if feedback becomes a testable change; fail if the response is simply more reading or more unrelated mocks.
Q14Give your two-minute positioning for a Senior GenAI or Forward Deployed role.
Strong answer outline
- State the production problem class you solve — for example, taking LLM applications from demo to operated, evaluated, cost-bounded systems — without inflating title or scope.
- Select two or three substantiated strengths matched to this role's signals, each backed by an artifact or a real story.
- Name the ownership you are seeking and bridge to one evidence-rich example the interviewer can pull on.
Follow-up probes
- Why this role rather than a research or foundation-model role?
- Which of your claims should we investigate first?
Pass if every sentence can lead to concrete evidence and fits in two minutes; fail if it is a biography, a tool list, or an unsupported superlative.
Proof artifact: the GenAI evidence dossier
Create a version-controlled dossier that maps one target role to inspectable proof across the five-artifact stack. This artifact tests preparation discipline; it is not a claim about Purnendu's past outcomes, and its README must say so.
Build steps
- Save a dated copy or structured summary of one live senior GenAI job description. Extract no more than ten decision-relevant signals using the matrix format from Section 2.
- Create the matrix with columns for signal, importance, evidence label, artifact or story, limitation, and next action. Commit it so the history shows re-ranking over time.
- Stand up the shared scenario thin: a small document corpus, an eval set of 100–200 questions with labels, and a baseline RAG pipeline on one cloud's managed path (Bedrock Knowledge Bases or Vertex AI Search).
- Add the five artifacts incrementally: eval harness with a CI gate; hybrid-versus-dense RAG benchmark; agent demo with guardrails and a replayable failure; inference-cost model using current published Bedrock and Vertex pricing; and the second-cloud mirror of the reference build.
- Write six real-experience story cards. Record exact personal scope and mark every number as exact, rounded/anonymized, or unavailable.
- Record a 45-minute design mock and a two-minute positioning statement. Score both with the same rubric, then repeat one week later.
- Publish a README covering reproducibility, secrets and privacy boundaries, current pricing-check dates, and what the practice system does not prove.
Metrics to capture
- Coverage: percentage of top signals with at least one inspectable proof; report gaps separately rather than hiding them in an average.
- Readiness: count of topics passing all five tests — define, design, implement, debug, defend.
- Artifact quality: each artifact carries its core metrics from the Section 5 table, with baselines and counter-metrics.
- Delivery: rubric scores for framing, assumptions, trade-offs, failure handling, measurement, and clarity across recorded mocks.
- Claim hygiene: number of unqualified "we" statements, unsupported metrics, or practice results phrased as experience. Target zero.
- Cadence: planned versus completed weekly deliverables, not passive study hours.
Deliberate failure injection
Three injections, one lesson each. First, remove the baseline and dataset description from the RAG benchmark report and ask a reviewer to interpret the claimed gain; it should become non-actionable, demonstrating why measurement provenance is engineering quality. Second, disable one guardrail in the agent demo, replay the failure transcript, and document the blast radius and the recovery path. Third, delete the denominator from the cost model — present "monthly spend" without "per successful task" — and note how the number stops supporting any decision. Restore each and record the difference.
What to present
Present the one-page signal matrix first, then one proof chain end to end: role need → decision → artifact → metric → injected failure → lesson. Use example numbers only if clearly labeled as hypothetical or portfolio measurements. Keep the complete repository available, but do not force an interviewer through every file.
Chapter review
Interview preparation for senior GenAI roles is a constrained engineering program. Decode live JD language into hiring signals, allocate effort by evidence gap on the effort-versus-signal grid, build a five-artifact proof stack around one shared scenario mapped to both Bedrock and Vertex AI, rehearse senior decision-making on a weekly cadence, and keep every claim honestly labeled. The goal is not to sound universally expert. It is to make relevant, honest competence easy to verify.
Glossary
- Hiring signal
- The conclusion an interviewer must draw about capability, judgment, or collaboration — the decoded meaning behind a JD phrase.
- Signal matrix
- A table mapping JD phrases to signals, best evidence forms, honest evidence labels, and next actions.
- Proof stack
- Five GenAI artifacts on one shared scenario: eval harness, RAG benchmark, guarded agent demo, inference-cost analysis, two-cloud reference build.
- Evidence label
- An explicit distinction among production experience, practice implementation, conceptual understanding, and a gap.
- Reversal condition
- New evidence or a threshold that would cause a technical decision — managed RAG, provisioned throughput, model choice — to be revisited.
- Story ledger
- A claim-calibrated inventory of real experiences, responsibilities, results, and lessons.
- Token economics
- Reasoning about GenAI cost in dollars per successful task across pricing modes, caching, and routing — not raw token counts.
- Counter-metric
- A measure that exposes harmful optimization of a primary metric, such as latency beside retrieval quality or quality beside cost.
Mastery checklist
- I can derive a ranked backlog from a current senior GenAI job description.
- I can decode Bedrock/Vertex/RAG/agent/eval JD phrases into the signal behind each.
- I can label each claim as experience, portfolio proof, conceptual knowledge, or gap.
- I can place any preparation activity on the effort-versus-signal grid and defend the placement.
- I have (or have scheduled) all five proof-stack artifacts on one shared scenario.
- My cost artifact states a denominator and a break-even, checked against current published pricing.
- I can explain one system through goal, constraints, alternatives, failure, rollout, and measurement.
- I have six stories with exact personal scope and no invented metrics.
- I have repeated a mock prompt after applying rubric-based feedback.
- I can name what I deliberately deferred and the trigger for revisiting it.
Primary sources
Links checked . Vendor pricing and features are volatile; re-verify at application time.
- Amazon Bedrock — official documentation
- Amazon Bedrock — pricing modes (on-demand, batch, provisioned throughput)
- AWS Well-Architected Framework — Generative AI Lens
- Vertex AI — official documentation
- Generative AI on Vertex AI — official documentation
- Vertex AI — pricing
- Anthropic — Building effective agents (engineering guidance)
- Lewis et al. — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv:2005.11401)
- Zheng et al. — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv:2306.05685)
CHAPTER 02 · PRIORITY 0
LLM Internals, Serving & Inference Optimization
32 min read · 15 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- Explain a decoder-only transformer at whiteboard depth — tokenization, embeddings, attention, MLP, residual stream — and connect each part to a serving cost.
- Derive KV-cache memory from first principles and use it to size batch, context, and hardware before touching a benchmark.
- Explain why time-to-first-token is compute-bound and time-per-output-token is memory-bandwidth-bound, and what each implies for optimization.
- Reason about continuous batching, PagedAttention, quantization, and speculative decoding as trade-offs with measurable failure modes, not as feature names.
- Compare vLLM, TGI, TensorRT-LLM, and SGLang, and defend a serving-engine choice for a stated workload.
- Build a defensible cost-per-million-tokens model and make a build-versus-buy call across Bedrock, Vertex AI, SageMaker, and self-hosted GPU serving.
- Benchmark TTFT, TPOT, and throughput honestly — under realistic load, with percentiles, and without vendor-benchmark traps.
1. Decoder-only anatomy, priced per component
Interviewers for platform and architect roles rarely want the training math. They want to know whether you can map each architectural component to a runtime cost, because that mapping is what makes serving decisions rational. A modern chat model is a decoder-only transformer: it converts text into tokens, tokens into vectors, pushes those vectors through a stack of identical blocks, and emits a probability distribution over the next token.
Tokenization is byte-pair encoding (BPE) or a byte-level variant: a learned merge table that greedily compresses frequent byte sequences into single vocabulary entries (vocabularies today run 32K–256K). Everything downstream — context limits, latency, and your invoice — is denominated in tokens, and tokenizers are not interchangeable: code, numbers, and non-English text can tokenize 1.5–4× longer than English prose, which silently changes both cost models and cross-engine benchmark comparisons. Embeddings map each token ID to a d_model-dimensional vector; position is injected not by adding position vectors but by RoPE (rotary position embeddings), which rotates query/key vectors by an angle proportional to position — a detail that matters later because it is the hook for context extension.
Each block applies pre-normalization (RMSNorm), self-attention, a residual add, another norm, then a gated MLP (SwiGLU, expanding to roughly 3–4× d_model), and a second residual add. Two facts earn senior credit: about two-thirds of parameters live in the MLPs, not attention — so weight memory and decode bandwidth are mostly an MLP bill; and the residual stream means each block edits a shared representation rather than replacing it, which is why layers can be quantized or even skipped with graceful rather than catastrophic degradation. Attention computes, for every position, a similarity score against every prior position — an n × n score matrix, which is the O(n²) term everyone cites. Crucially, that quadratic cost is paid during prompt processing; once past keys/values are cached, each new token attends to n cached entries — linear per step.
MHA → MQA → GQA → MLA: shrinking the cache, not the compute
The attention variants exist for one dominant reason: the KV cache (next section) scales with the number of key/value heads. Multi-query attention (MQA) shares one K/V head across all query heads; grouped-query attention (GQA) shares K/V among groups — Llama-3-70B uses 8 KV heads against 64 query heads, an 8× cache reduction; multi-head latent attention (MLA, introduced by DeepSeek-V2) caches a low-rank latent compression of K/V, cutting cache size by an order of magnitude at the price of extra projection compute.
| Variant | KV heads | Cache vs MHA | Trade-off |
|---|---|---|---|
| MHA | = query heads | 1× | Maximum quality headroom, maximum cache. |
| MQA | 1 | ~1/64× | Cheapest cache; measurable quality loss on some tasks. |
| GQA | groups (e.g., 8) | ~1/8× | Near-MHA quality; the current open-weights default. |
| MLA | latent vector | ~1/10–1/30× | Big cache savings; more matmuls per token, more kernel complexity. |
2. KV cache and the two-phase request
Every autoregressive request has two phases with opposite performance characters. Prefill processes all prompt tokens in parallel — big matrix multiplies, high arithmetic intensity, compute-bound — and its duration is essentially your time-to-first-token (TTFT). Decode generates one token per forward pass; each step must stream the model weights and the growing KV cache through the GPU's memory system to do comparatively little math, so it is memory-bandwidth-bound, and it sets time-per-output-token (TPOT).
sequenceDiagram
participant C as "Client"
participant R as "API gateway"
participant S as "Engine scheduler"
participant G as "GPU worker"
C->>R: submit prompt
R->>S: enqueue request
S->>G: admit and allocate KV blocks
G->>G: prefill all prompt tokens in one pass
G-->>C: first token streamed
loop one token per step until stop
G->>G: read weights plus KV cache
G-->>C: stream next token
end
G->>S: release KV blocks
The KV cache stores every layer's keys and values for every token so decode never recomputes them. Memorize the formula and one worked example:
kv_bytes_per_token = 2 × n_layers × n_kv_heads × head_dim × dtype_bytes
Llama-3-70B (80 layers, 8 KV heads, head_dim 128, FP16):
2 × 80 × 8 × 128 × 2 = 327,680 B ≈ 320 KB per token
→ 8K-token session ≈ 2.6 GB → 32K ≈ 10.5 GB → 128K ≈ 42 GB per sequence
Set that against hardware: FP16 weights for a 70B model are ~140 GB, so an 8×H100 node (640 GB HBM) has roughly 400+ GB left for KV after weights and activation workspace — call it ~1.2M cached tokens, or about 300 concurrent 4K-context sessions, but only nine concurrent 128K sessions. This single calculation explains why long context is an economics problem before it is a quality problem, and why GQA/MLA and KV quantization exist at all.
The roofline argument is worth reciting. At batch 1, decoding one token reads every weight byte to perform ~2 FLOPs per parameter — arithmetic intensity near 1 FLOP/byte, while an H100 needs roughly 295 FLOPs per byte moved (~989 dense BF16 TFLOPS over ~3.35 TB/s HBM3) to saturate. Hence the batch-1 ceiling: an 8B model in FP16 (~16 GB) on one H100 cannot exceed ~3,350/16 ≈ 200 tokens/s no matter how clever the kernels, and measured numbers sit below that. Batching multiplies useful FLOPs per weight byte read, which is why every serving optimization ultimately serves one goal: keep many sequences in flight per weight pass. Prefill math is equally quotable: a 2,000-token prompt on that 8B model needs ~2 × 8e9 × 2000 = 32 TFLOPs, so at ~50% utilization TTFT has a floor around 65 ms — before queueing, which dominates in practice.
3. Continuous batching and PagedAttention
Naive serving batches requests statically: collect N prompts, run them together, return when the longest finishes. Two pathologies follow — the batch runs at the speed of its slowest member while finished slots idle, and arrivals wait for the next batch to form. Orca introduced iteration-level scheduling — continuous batching — where the scheduler recomposes the batch at every decode step: finished sequences leave immediately, queued requests join mid-flight. This alone yields multi-fold throughput gains over static batching and is table stakes in every modern engine.
flowchart TD
Q["Incoming request queue"] --> A["Admission control checks free KV blocks"]
A -->|"blocks available"| B["Join running batch at next step"]
A -->|"pool exhausted"| P["Queue, or preempt a sequence by swap or recompute"]
B --> F["Single forward pass for whole batch"]
F --> D{"Sequence hit stop token or max length?"}
D -->|"yes"| E["Stream final token and free KV pages"]
D -->|"no"| B
E --> A
Continuous batching made KV allocation the new bottleneck. Pre-PagedAttention engines reserved contiguous KV memory for each request's maximum possible length; measured waste from internal fragmentation and over-reservation ran 60–80% of KV memory. PagedAttention (the vLLM paper) applies virtual-memory thinking: KV lives in fixed-size blocks (e.g., 16 tokens), each sequence holds a page table, and blocks are allocated on demand — waste drops to under ~4%, which converts directly into batch size and therefore throughput. Paging also enables prefix sharing with copy-on-write: a thousand requests carrying the same 2K-token system prompt can reference one physical copy of its KV blocks. SGLang generalized this into RadixAttention, an automatic prefix-cache tree, which is why it shines on agentic workloads that re-send conversation prefixes on every step (Chapter 05 depends on this).
- Chunked prefill — split large prompt prefills into slices interleaved with decode steps, so one 100K-token upload does not spike everyone else's TPOT.
- Preemption — when the KV pool exhausts, engines evict a sequence (recompute later, or swap to CPU). Watch this metric: preemptions are the canary for KV pressure.
- Prefill–decode disaggregation — run prefill and decode on separate GPU pools and ship KV between them; removes phase interference at the cost of a KV transfer path. An emerging default for large deployments.
- Goodput, not throughput — the number that matters is tokens/s delivered while meeting TTFT/TPOT SLOs, not peak tokens/s at unbounded latency.
4. The quantization ladder
Quantization is the highest-leverage cost knob after batching, because decode speed is proportional to bytes moved. The ladder runs from BF16 (training-native baseline) down to 4-bit weights, and the senior move is knowing which rung serves which regime: weight-only quantization accelerates the memory-bound low-batch regime; weight-and-activation formats (FP8/INT8) exploit faster tensor cores and win in the compute-bound high-batch regime, where weight-only 4-bit can actually lose to FP8 because of per-step dequantization overhead.
| Rung | What is quantized | Hardware | Typical quality cost | Use when |
|---|---|---|---|---|
| BF16 / FP16 | nothing (baseline) | all | reference | Quality baselining, evals, low-risk default. |
| FP8 (E4M3) | weights + activations | Hopper/Ada and newer | usually <1% on aggregate evals | High-throughput production on H100-class GPUs; ~2× compute and half the weight bytes. |
| INT8 (LLM.int8, SmoothQuant) | weights ± activations | Ampere and newer | small, outlier-sensitive | Pre-Hopper fleets; activation outliers need special handling. |
| 4-bit weight-only (GPTQ, AWQ) | weights only | all | ~1% aggregate, but task-skewed | Fit big models on small GPUs; fastest decode at low batch. |
| NF4 (QLoRA) | weights, for fine-tuning | all | designed for training memory | PEFT on constrained GPUs — a Chapter 03 topic, not a serving format. |
| FP8 KV cache | the cache itself | Hopper-class | usually negligible | Double token capacity per GPU; pairs with long context. |
The quality caveat is where candidates get filtered. Aggregate benchmarks (MMLU-style) routinely show <1% degradation for well-executed 4-bit quantization of large models, while narrow capabilities — math, code generation, low-resource languages, strict instruction-following at long context — degrade first and disproportionately. GPTQ minimizes layer-wise reconstruction error using second-order information; AWQ instead protects the ~1% of weight channels with the largest activation magnitudes. Both are calibration-dependent: quantize with generic web-text calibration data and evaluate on your legal-summarization traffic, and you may see regressions the model card never showed. The rule: the quantization decision is an evaluation decision — your own task harness (Chapter 06) gates every rung change.
flowchart TD
S["Start at BF16 with a task eval harness"] --> Q1{"Bottleneck today?"}
Q1 -->|"memory capacity or low-batch latency"| W4["Weight-only 4-bit AWQ or GPTQ"]
Q1 -->|"throughput at high batch on Hopper"| F8["FP8 weights and activations"]
Q1 -->|"KV capacity limits concurrency"| KV["FP8 KV cache first"]
W4 --> E["Re-run task evals plus latency benchmark"]
F8 --> E
KV --> E
E --> G{"Quality within agreed budget on your slices?"}
G -->|"yes"| SHIP["Ship, monitor drift and complaints"]
G -->|"no"| ROLL["Step back up one rung or change calibration set"]
5. Faster decode: speculation, MoE, and long context
Three techniques dominate the “make decode cheaper” conversation, and each has a regime where it backfires. Speculative decoding (Leviathan et al., Chen et al.) uses a cheap drafter to propose k tokens, which the target model verifies in a single parallel pass; rejection sampling guarantees the output distribution is exactly the target model's. Because verification prices like a tiny prefill, you convert several bandwidth-bound steps into one compute-heavier step — a 1.5–3× TPOT win when acceptance rates are high (drafter matches the target's style; predictable text like code) and batch is small. At high batch the GPU is already compute-saturated, so speculation adds work and can reduce throughput; engines increasingly disable it dynamically under load. EAGLE and Medusa-style approaches replace the separate drafter with lightweight heads on the target model itself, removing the two-model operational burden.
flowchart LR
DR["Drafter proposes k tokens"] --> V["Target verifies all k in one parallel pass"]
V --> AC{"How much of the draft matched?"}
AC -->|"all k accepted"| N["Emit k plus one bonus token"]
AC -->|"partial"| PR["Emit accepted prefix plus one corrected token"]
N --> DR
PR --> DR
Mixture-of-experts models decouple parameters from per-token compute: Mixtral 8×7B holds ~47B parameters but routes each token through ~13B. The serving implications are the interview substance: all experts must be resident, so memory is priced at total parameters while compute is priced at active parameters; at batch 1 you read only the routed experts (bandwidth win), but as batch grows nearly every expert is activated by someone, so weight-read amortization converges back toward dense behavior; and multi-GPU MoE adds expert-parallel all-to-all communication plus load-balancing pathologies when routing concentrates on hot experts. MoE is a hardware-utilization bet, not a free lunch.
Long context is bounded by two costs you now own: KV memory growing linearly (Section 2's formula) and prefill compute growing quadratically-ish in practice. Models extend beyond trained context by rescaling RoPE — position interpolation, NTK-aware scaling, and YaRN — which stretches rotary frequencies so unseen positions land in familiar ranges; quality at extended lengths must be verified with needle-and-haystack and task-level evals, since "supports 128K" often means "does not crash at 128K." Operationally, pair long context with FP8 KV, chunked prefill, and prefix caching — and remember that the cheapest long-context token is the one you retrieved instead (Chapter 04's argument).
6. Serving engines: vLLM, TGI, TensorRT-LLM, SGLang
All four mainstream engines now implement continuous batching, paged KV, quantized formats, and OpenAI-compatible APIs, so the differentiation is operational: kernel pedigree, model-coverage velocity, prefix-caching sophistication, and how much build engineering they demand.
| Engine | Core strength | Cost you accept | Choose when |
|---|---|---|---|
| vLLM | Community default; fastest new-model support; PagedAttention origin; huge feature surface (LoRA serving, spec decode, P/D disaggregation). | Config surface is large; peak perf sometimes trails compiled engines. | Default choice, model churn expected, k8s self-hosting. |
| TGI | Hugging Face ecosystem integration, hardened Rust server, simple operational story. | Smaller optimization community than vLLM today. | HF-centric stack, straightforward deployments. |
| TensorRT-LLM | NVIDIA-tuned compiled kernels; frequently the raw tokens/s ceiling on NVIDIA GPUs; pairs with Triton/NIM. | Engine builds per model/GPU/config; slower model onboarding; NVIDIA lock-in. | Stable model list, extreme throughput targets, NVIDIA fleet. |
| SGLang | RadixAttention automatic prefix caching; strong structured-output and multi-call programs; excellent agent/high-QPS results. | Younger ecosystem; fewer enterprise integrations. | Agentic traffic with heavy shared prefixes; JSON-constrained decoding at scale. |
Interview framing that lands: the engine choice is reversible (they share API shapes); the benchmark methodology is what protects the decision. Comparing engines on tokens/s with different tokenizers, different default sampling, or unmatched quantization is the classic self-deception — normalize the workload first (Section 9).
7. GPU economics and build-versus-buy
Cost per million tokens is the unit that makes inference decisions comparable across managed APIs and self-hosting. The formula is trivial — GPU-hour cost divided by sustained tokens per hour — and every input is a place candidates go wrong: they use peak benchmark throughput instead of SLO-compliant goodput, and they assume 100% utilization when real diurnal traffic delivers 20–40%.
| GPU (approx. specs) | HBM | Bandwidth | Serving role |
|---|---|---|---|
| L4 | 24 GB | ~0.3 TB/s | Small models (≤8B quantized), embeddings, bursty low-QPS endpoints. |
| L40S | 48 GB | ~0.86 TB/s | Mid-size models, cost-efficient inference without HBM pricing. |
| A100 80GB | 80 GB | ~2.0 TB/s | Prior-gen workhorse; no FP8. |
| H100 80GB | 80 GB | ~3.35 TB/s | FP8 tensor cores; the current serving baseline (specs). |
| H200 | 141 GB | ~4.8 TB/s | KV-heavy and long-context serving; fewer GPUs per 70B replica. |
Managed API, on-demand
Pay per token, zero capacity risk, frontier-model access. Right up to the volume where a provisioned or self-hosted floor beats the token price — and always right for spiky, low-volume, or frontier-quality workloads.
Provisioned throughput
Reserve model capacity (Bedrock provisioned throughput, Vertex AI Provisioned Throughput) for predictable latency and volume discounts on steady traffic — commitment risk replaces queueing risk.
Self-host open weights
Wins on price only with sustained high utilization, tolerance for open-weight quality, and a team that can own GPU capacity, upgrades, and incident response. Also the only option under strict data-residency or air-gap constraints.
Hybrid routing
The common senior answer: managed frontier models for hard/low-volume traffic, self-hosted small models for high-volume narrow tasks, with an eval-gated router (Chapters 05–06) deciding.
8. Serving on AWS and GCP
Both clouds offer the same three altitudes — fully managed per-token APIs, managed endpoints you configure, and raw GPUs you orchestrate — and interviewers for cloud-aligned roles expect you to pick an altitude from workload characteristics, then name concrete services fluently.
AWS
- Bedrock on-demandpay-per-token managed FM inference; also batch mode at a discount
- Bedrock Provisioned Throughputreserved model units for steady, latency-sensitive volume
- Bedrock Custom Model Importserve your fine-tuned open weights behind the Bedrock API
- SageMaker real-time endpoints + LMI containersmanaged autoscaling endpoints running vLLM/TensorRT-LLM backends
- EKS + vLLMfull-control GPU serving; Karpenter for GPU node autoscaling
- Inferentia2 / Trainiumcustom-silicon price-performance play via the Neuron SDK
Google Cloud
- Vertex AI Gemini APIpay-as-you-go managed inference with context caching
- Vertex AI Provisioned Throughputreserved throughput for Gemini at predictable latency
- Model Garden → Vertex endpointsdeploy open models onto managed GPU endpoints, vLLM-based containers
- GKE + vLLMself-managed serving; GKE Inference Gateway adds LLM-aware load balancing
- Cloud Run GPUsserverless L4 GPUs with scale-to-zero for bursty small-model serving
- Cloud TPUalternative accelerator for supported stacks (vLLM TPU, JetStream)
flowchart TD
R["New GenAI workload"] --> Q1{"Frontier-model quality required?"}
Q1 -->|"yes"| M["Managed API on Bedrock or Vertex AI"]
Q1 -->|"open weights pass evals"| Q2{"Steady traffic at high utilization?"}
Q2 -->|"no, spiky or unproven"| M2["Managed endpoints or serverless GPUs, pay as you go"]
Q2 -->|"yes"| Q3{"Team ready to own GPU operations?"}
Q3 -->|"yes"| SH["Self-host vLLM on EKS or GKE"]
Q3 -->|"no"| PT["Provisioned throughput or managed endpoints"]
M --> RT["Re-evaluate quarterly as volume and models shift"]
SH --> RT
9. TTFT, TPOT, throughput — benchmarking without self-deception
Three metrics, three different masters. TTFT (time to first token) is queueing plus prefill — it gates perceived responsiveness in chat and agent loops. TPOT (time per output token, a.k.a. inter-token latency) is the decode rhythm — it gates streaming readability and total completion time. Throughput (aggregate tokens/s per replica) is the cost side. They trade against each other through batch depth: deeper batches raise throughput and TPOT together, and a benchmark that reports one without the others is marketing.
| Metric | Bound by | Improved by | Report as |
|---|---|---|---|
| TTFT | queueing + prefill compute | chunked prefill, prefix caching, more replicas, shorter prompts | p50 / p95 / p99 at a stated request rate |
| TPOT | memory bandwidth per weight pass | quantization, speculative decoding, shallower batches, faster HBM | p50 / p99 per token, streaming |
| Throughput | KV capacity + compute ceiling | continuous batching, paged KV, FP8, longer batch depth | sustained tokens/s at SLO-compliant load (goodput) |
- Fix the workloadSample real prompt/output length distributions — a 6K-in / 300-out RAG trace behaves nothing like ShareGPT chat. Same tokenizer accounting across engines.
- Sweep loadRamp concurrency or request rate stepwise; at each level record TTFT and TPOT percentiles plus aggregate tokens/s.
- Find the kneeIdentify where p95 TTFT or TPOT breaches SLO; the sustained rate just below it is your capacity number.
- Price itCost per million tokens at the knee — not at peak throughput — feeds the build-vs-buy model of Section 7.
- Re-run on changeNew model, quant rung, engine version, or context policy re-runs the sweep; keep results versioned like eval results.
The dishonesty patterns to name in an interview: benchmarking at batch 1 and deploying at batch 64; warm prefix caches flattering TTFT for prompts your real traffic never repeats; comparing engines with different tokenizers so “tokens/s” measures the tokenizer, not the engine; quoting mean latency where tail latency gates the product; and quoting a managed API's demo-time latency without measuring its variance under your regional, peak-hour traffic. A candidate who volunteers these earns instant credibility.
Interview playbook
For any inference-performance question, run BUDGET — size the physics before naming products:
- B — Bytes: weights at the chosen precision, plus KV per token × context × concurrency. Does it fit, and on how many GPUs?
- U — Utilization regime: is the phase compute-bound (prefill, high batch) or bandwidth-bound (decode, low batch)? The regime picks the optimization.
- D — Demand shape: traffic pattern, prompt/output length distribution, prefix reuse, latency SLOs per product surface.
- G — Gains ladder: continuous batching → paged KV/prefix cache → quantization rung → speculative decoding → disaggregation, each gated by evals.
- E — Economics: cost per million tokens at SLO-compliant load; compare managed API vs provisioned vs self-host at the actual utilization.
- T — Test honestly: load-swept percentiles on a realistic trace, re-run on every change.
Common traps
- Citing O(n²) attention as why long context is expensive, instead of KV memory pressure and prefill interference.
- Treating quantization as free compression — no calibration story, no task-slice evals, no rollback rung.
- Recommending speculative decoding for a saturated high-batch cluster where it burns the compute headroom it needs.
- Comparing self-hosting at 100% assumed utilization against a managed API's list price.
- Benchmark numbers without percentiles, load levels, or the prompt-length distribution attached.
- Naming vLLM features without being able to explain what PagedAttention actually fixed (fragmentation and over-reservation of KV memory).
Question bank
These are the questions senior GenAI platform and architect loops actually ask about inference. Practice deriving, not reciting.
Q1Walk me through what happens between an HTTP request hitting your endpoint and the first streamed token.
Strong answer outline
- Gateway: auth, rate limit, route to a replica; request enters the engine scheduler queue.
- Admission: scheduler checks free KV blocks, allocates pages, joins the request into the running batch at the next iteration.
- Prefill: all prompt tokens processed in parallel, KV cache written per layer; this compute plus the queue wait is TTFT.
- First token sampled from the LM head, detokenized, streamed; decode loop begins at one token per pass.
Follow-up probes
- Where can this request be preempted, and what happens to its KV?
- What changes if the prompt shares a 2K system prefix with other traffic?
Pass if queueing, KV allocation, and the prefill/decode split are all present; fail if the answer jumps from “request” to “the model generates.”
Q2Derive the KV-cache memory for a 70B GQA model at 32K context. What does it imply for concurrency?
Strong answer outline
- State the formula: 2 × layers × KV heads × head_dim × dtype bytes per token.
- Compute: 2 × 80 × 8 × 128 × 2 ≈ 320 KB/token → ~10.5 GB per 32K sequence at FP16.
- Subtract weights from HBM (140 GB FP16 on 640 GB node) → bound concurrent sequences; show how FP8 KV or MLA changes the bound.
- Conclude: KV capacity, not compute, usually caps batch size and thus throughput.
Follow-up probes
- Why does GQA divide this by 8 relative to MHA?
- What operational metric tells you KV pressure is biting?
Pass if the arithmetic is done live and tied to concurrency; fail if the formula is recited without consequences.
Q3Why is TTFT compute-bound and TPOT bandwidth-bound, and what follows for optimization?
Strong answer outline
- Prefill: thousands of tokens per weight read → high arithmetic intensity → limited by TFLOPS.
- Decode: one token per full weight read at low batch → ~1 FLOP/byte against a GPU needing hundreds → limited by HBM bandwidth.
- Therefore TTFT improves with compute, chunked prefill, prefix caching, shorter prompts; TPOT improves with fewer bytes — quantization, speculation, faster HBM.
- Batching moves decode toward compute-bound, trading TPOT for throughput.
Follow-up probes
- Estimate the batch-1 decode ceiling for an 8B FP16 model on an H100.
- Why does H200's bandwidth matter more than its FLOPS for serving?
Pass if arithmetic intensity is explained and each optimization maps to a phase; fail on “prefill is parallel, decode is serial” alone.
Q4Compare MQA, GQA, and MLA. Why did the industry converge on GQA, and what does MLA change?
Strong answer outline
- All three shrink KV per token; only cache size and quality differ, not O(n²) prefill.
- MQA: 1 KV head, maximum saving, measurable quality loss on some tasks; GQA: grouped middle ground with near-MHA quality — the pragmatic winner.
- MLA caches a low-rank latent instead of full K/V — order-of-magnitude smaller cache, extra projection compute and kernel complexity.
- Impact channel: cache bytes → concurrency → cost/M tokens.
Follow-up probes
- Can you convert a trained MHA model to GQA after the fact?
- Where does MLA's saving matter most — chat or long-context RAG?
Pass if the answer prices variants in KV bytes and concurrency; fail if it is an acronym tour.
Q5What problem did PagedAttention actually solve, and what became possible because of it?
Strong answer outline
- Before: contiguous KV allocation at max sequence length → 60–80% of KV memory lost to fragmentation and over-reservation.
- PagedAttention: fixed-size KV blocks with per-sequence page tables → waste under ~4%, recovered memory becomes batch depth.
- Enabled: copy-on-write prefix sharing, cheap preemption/swap, and practical continuous batching at scale.
- Distinguish from RadixAttention: automatic prefix-tree reuse across requests.
Follow-up probes
- What is the analogue of a TLB miss here, and does block size matter?
- How does prefix caching change your benchmark design?
Pass if fragmentation and the memory-to-throughput conversion are explicit; fail if PagedAttention is described as “an attention optimization.”
Q6Continuous batching raised our throughput 5× but p99 TTFT got worse. Explain and fix it.
Strong answer outline
- Diagnose interference: long prefills of admitted requests stall the decode loop; deep batches lengthen queue waits at peak.
- Instrument: queue time vs prefill time vs decode jitter; preemption counts; KV pool occupancy.
- Fix ladder: chunked prefill, admission priorities by prompt length or tenant, KV headroom targets, separate long-context pool, or prefill/decode disaggregation.
- Redefine success as goodput at SLO, not peak tokens/s.
Follow-up probes
- Which fix would you try first and how would you verify?
- When is adding replicas the wrong answer?
Pass if the interference mechanism and a measured fix sequence are given; fail on “tune the batch size” alone.
Q7Design the quantization strategy for self-hosting a 70B model. Which rung and how do you validate?
Strong answer outline
- Start from bottleneck and hardware: Hopper + high batch → FP8 W8A8; single-GPU fit or low batch → AWQ/GPTQ 4-bit weight-only; add FP8 KV if concurrency is KV-capped.
- Calibration matters: use domain-representative calibration data, not generic web text.
- Validate on your task harness with slices (math, code, multilingual, long context), plus a latency/throughput sweep — quality and speed together.
- Ship with a rollback rung and monitor drift and complaint rates.
Follow-up probes
- Why can 4-bit weight-only lose to FP8 at batch 64?
- Where does NF4 belong — and why not in serving?
Pass if regime → format mapping and eval gating are both present; fail if a single format is recommended unconditionally.
Q8When does speculative decoding help, when does it hurt, and how does it preserve output quality?
Strong answer outline
- Mechanism: drafter proposes k tokens; target verifies in one parallel pass; rejection sampling keeps exactly the target distribution — quality is provably unchanged.
- Helps when: spare compute (low batch), high acceptance (predictable text, aligned drafter) → 1.5–3× TPOT.
- Hurts when: compute-saturated high batch, or domain-shifted traffic drops acceptance — extra work, worse throughput.
- Ops: monitor live acceptance rate; prefer EAGLE/Medusa-style heads to avoid running a second model.
Follow-up probes
- What acceptance rate breaks even at draft length 4?
- Does speculation change sampling temperature semantics?
Pass if the exactness guarantee and both regimes are stated; fail if it is pitched as a universal speedup.
Q9What is different about serving a mixture-of-experts model like Mixtral?
Strong answer outline
- Memory prices at total parameters (all experts resident); compute prices at active parameters per token.
- Batch effect: at low batch, routed-expert reads save bandwidth; at high batch most experts activate across the batch, eroding the saving.
- Multi-GPU: expert parallelism adds all-to-all communication; hot-expert imbalance creates stragglers.
- Net: MoE trades memory footprint and ops complexity for per-token compute — evaluate against a dense model at equal quality, not equal parameter count.
Follow-up probes
- How would you place experts across an 8-GPU node?
- What breaks if one expert receives 40% of tokens?
Pass if the total-vs-active distinction and the batch-size erosion are both explained; fail on “MoE is cheaper” without conditions.
Q10Product wants 128K context. What actually breaks, and what is your plan?
Strong answer outline
- Cost the KV: linear growth per token (e.g., ~42 GB per 128K sequence for a 70B GQA model at FP16) — concurrency collapses first.
- Prefill: 128K prefill is a multi-second compute event that starves co-located decodes; needs chunked prefill or a separate pool.
- Quality: RoPE scaling (PI, NTK-aware, YaRN) extends positions, but verify with retrieval-in-context and task evals, not the spec sheet.
- Mitigate demand: retrieval, summarization memory, prefix/context caching — cheapest long-context token is the one not sent (Chapter 04).
Follow-up probes
- How does FP8 KV change the capacity math?
- Would you price long-context requests differently?
Pass if memory, interference, and quality verification all appear; fail if the answer is only “use YaRN.”
Q11Choose between vLLM, TGI, TensorRT-LLM, and SGLang for a given workload. What drives the decision?
Strong answer outline
- Establish workload: model churn rate, prefix reuse, structured output needs, peak throughput target, team skill, hardware fleet.
- Map: vLLM as default and fastest model coverage; TensorRT-LLM for maximum tokens/s on a stable NVIDIA-only model list at the price of engine builds; SGLang for prefix-heavy agentic and constrained-decoding traffic; TGI for HF-centric simplicity.
- Note convergence: features cross-pollinate; the choice is reversible behind an OpenAI-compatible API.
- Commit to a normalized benchmark on your trace before deciding.
Follow-up probes
- What breaks when you upgrade the engine version in place?
- How do you compare engines with different tokenizers fairly?
Pass if the decision is workload-conditional with a benchmark gate; fail if it is a leaderboard ranking.
Q12Build a cost-per-million-tokens model for self-hosting versus a managed API. Where do these models usually lie?
Strong answer outline
- Formula: GPU-hour cost ÷ SLO-compliant sustained tokens per hour, at the load knee — not peak benchmark throughput.
- Apply utilization: diurnal traffic at 20–40% average multiplies effective cost 2.5–5×; add redundancy, failover headroom, and engineering time.
- Compare against managed per-token pricing, which embeds the provider's utilization pooling; include provisioned-throughput commitments as the middle option.
- Present break-even volume and the reversal conditions.
Follow-up probes
- How do input vs output token prices change the comparison for RAG traffic?
- What utilization assumption would flip your recommendation?
Pass if utilization is the pivotal variable and numbers are labeled as examples; fail if list prices are compared at assumed 100% usage.
Q13On AWS: Bedrock on-demand vs Provisioned Throughput vs SageMaker endpoints vs EKS with vLLM — give the decision framework.
Strong answer outline
- Bedrock on-demand: frontier and partner models, per-token, zero capacity ops — default for spiky or exploratory traffic; batch mode for offline volume.
- Bedrock Provisioned Throughput: steady high volume on Bedrock models needing predictable latency; commitment risk.
- SageMaker + LMI containers: open or custom weights with managed endpoints, autoscaling, VPC control — the middle altitude.
- EKS + vLLM: maximum control and lowest unit cost at sustained utilization, highest ops burden; justify with volume, residency, or customization needs.
Follow-up probes
- Where does Bedrock Custom Model Import fit between these?
- Map the equivalent ladder on GCP.
Pass if each option gets a workload condition and a named cost/ops trade; fail if services are listed without decision criteria.
Q14How do you benchmark an inference deployment honestly? Name the ways teams fool themselves.
Strong answer outline
- Fix a realistic trace: real prompt/output length distributions, realistic prefix reuse, consistent tokenizer accounting.
- Sweep load stepwise; record p50/p95/p99 TTFT and TPOT plus tokens/s at each level; find the SLO knee.
- Report goodput and cost at the knee; version results and re-run on any model/engine/config change.
- Name traps: batch-1 numbers for a batch-64 deployment, warm prefix caches, mean instead of tails, cross-tokenizer tokens/s, off-peak API latency samples.
Follow-up probes
- How many requests do you need for a stable p99?
- How would you benchmark a managed API you cannot instrument server-side?
Pass if load-swept percentiles on a realistic trace plus at least three traps are given; fail on a single tokens/s figure.
Q15Your p99 TTFT tripled after enabling 64K contexts for one tenant. Debug it live.
Strong answer outline
- Hypothesize the two mechanisms: giant prefills monopolizing iterations (head-of-line for other requests) and KV pressure causing queueing/preemption.
- Check evidence: queue-time vs prefill-time decomposition, KV pool occupancy, preemption counters, correlation with the tenant's request timestamps.
- Mitigate in order: enable/tune chunked prefill, cap admitted prompt length per iteration, tenant-level concurrency limits, dedicated long-context pool or disaggregated prefill.
- Verify with the same load sweep and add a regression alert on p99 TTFT by tenant class.
Follow-up probes
- Why might adding a replica not fix this?
- What would you have load-tested before the rollout?
Pass if both mechanisms are named with the telemetry to distinguish them; fail if the answer is generic “scale up.”
Proof artifact: an honest inference benchmark and cost model
Build a small, reproducible study you can defend line by line: one open model, one engine, a load harness, and a one-page decision memo. This artifact backs answers across Sections 2–9 and pairs with the evaluation harness of Chapter 06.
Steps
- Deploy an 8B-class instruct model with vLLM on a single cloud GPU (an L4 or A10G-class instance keeps example cost low); record exact model, engine version, dtype, and max context.
- Build a load generator that replays a realistic trace: sampled prompt lengths (include a long-prompt slice), realistic output lengths, streaming enabled, fixed random seed.
- Sweep concurrency (e.g., 1, 2, 4, 8, 16, 32, 64) and record p50/p95/p99 TTFT and TPOT plus aggregate tokens/s at each level; chart the saturation curve and mark the SLO knee.
- Compute cost per million output tokens at the knee from the instance's hourly price; then redo the number at 25% assumed utilization and compare with two managed-API list prices as reference points.
- Quantize to AWQ (or FP8 if the GPU supports it), re-run both the load sweep and a ~100-item task eval with slices; record the quality delta next to the speed delta.
Deliberate failures
- Drive KV exhaustion with many long-context sessions; capture preemption/queueing behavior and the client-visible symptom.
- Inject one 30K-token prompt into a busy interactive load; show the TPOT jitter on other streams, then enable chunked prefill and show the repair.
- Repeat a shared system prefix across requests with prefix caching on and off; quantify the TTFT delta and note how it could flatter a dishonest benchmark.
- Compare the quantized model on the aggregate eval and on a math/code slice; show a slice regression that the average hides.
What to present
One saturation chart (TTFT/TPOT percentiles vs load), the KV arithmetic for your model done by hand, a cost table at three utilization assumptions, one failure trace with its fix, and a half-page build-vs-buy recommendation with explicit reversal conditions. State clearly that all figures are from your own small-scale study — that honesty is itself a senior signal.
Chapter review
Inference is a memory system wearing a model's clothes. Prefill spends compute and sets TTFT; decode spends memory bandwidth and sets TPOT; the KV cache converts context length and concurrency into bytes that compete with weights for HBM. Continuous batching and PagedAttention recover wasted capacity, the quantization ladder trades verified quality for bytes, speculation trades spare compute for latency, and every choice terminates in one number — cost per million tokens at SLO-compliant load — which decides build-versus-buy across Bedrock, Vertex AI, and self-hosted GPU fleets.
Glossary
- Prefill
- Parallel processing of all prompt tokens that fills the KV cache; compute-bound; determines TTFT.
- Decode
- One-token-per-pass generation that re-reads weights and KV each step; bandwidth-bound; determines TPOT.
- KV cache
- Per-layer keys/values stored per token: 2 × layers × KV heads × head_dim × dtype bytes per token.
- TTFT / TPOT
- Time to first token; time per output token thereafter. Report as percentiles at stated load.
- GQA / MLA
- Attention variants that shrink KV per token — grouped KV heads, or a cached low-rank latent.
- Continuous batching
- Iteration-level scheduling: sequences join and leave the batch at every decode step.
- PagedAttention
- Block-based KV allocation with page tables; kills fragmentation, enables prefix sharing.
- Chunked prefill
- Splitting long prefills into slices interleaved with decode to bound interference.
- Speculative decoding
- Draft-then-verify generation that preserves the target distribution exactly via rejection sampling.
- RoPE scaling
- Rescaling rotary position frequencies (PI, NTK, YaRN) to extend context beyond training length.
- MoE
- Sparse expert routing: memory priced at total parameters, compute at active parameters.
- Goodput
- Sustained tokens/s delivered while meeting TTFT/TPOT SLOs — the honest capacity number.
Mastery checklist
- I can derive KV bytes per token for a named model and turn it into a concurrency bound.
- I can explain arithmetic intensity and estimate a batch-1 decode ceiling from bandwidth and model size.
- I can describe what PagedAttention fixed, with the before/after memory-waste numbers.
- I can pick a quantization rung from the bottleneck regime and defend the eval gate around it.
- I can state both preconditions for speculative decoding to pay off — and its exactness guarantee.
- I can explain MoE serving economics: total vs active parameters, and the batch-size erosion effect.
- I can choose a serving engine conditionally and design a fair cross-engine benchmark.
- I can build a cost-per-million-tokens model where utilization is the pivotal variable.
- I can name the managed / provisioned / self-hosted ladder on both AWS and GCP with concrete services.
- I can benchmark TTFT, TPOT, and goodput with load-swept percentiles and name five benchmark deceptions.
Primary sources
Links checked . GPU specs, engine features, and cloud pricing move quickly; verify against the current official pages before quoting numbers in an interview.
- Vaswani et al. — Attention Is All You Need
- Su et al. — RoFormer: rotary position embeddings
- Shazeer — multi-query attention
- Ainslie et al. — GQA: grouped-query attention
- DeepSeek-V2 — multi-head latent attention
- Dao et al. — FlashAttention
- Kwon et al. — PagedAttention and vLLM
- Yu et al. — Orca: iteration-level scheduling for transformer serving
- Frantar et al. — GPTQ
- Lin et al. — AWQ: activation-aware weight quantization
- Dettmers et al. — LLM.int8
- Dettmers et al. — QLoRA and NF4
- Leviathan et al. — fast inference via speculative decoding
- Chen et al. — accelerating LLM decoding with speculative sampling
- Peng et al. — YaRN context extension
- Jiang et al. — Mixtral of Experts
- vLLM — official documentation
- TensorRT-LLM — official documentation
- Text Generation Inference — official documentation
- SGLang — official documentation
- Amazon Bedrock — user guide
- Amazon Bedrock — Provisioned Throughput
- Amazon SageMaker AI — documentation, incl. large model inference containers
- Vertex AI — generative AI documentation
- Vertex AI — Provisioned Throughput
- Cloud Run — GPU configuration
- NVIDIA — H100 specifications
CHAPTER 03 · PRIORITY 0
Prompting, Context Engineering & Model Adaptation
26 min read · 12 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- Apply the adaptation ladder—prompting → RAG → PEFT → full fine-tune → distillation—with honest criteria: knowledge vs. behavior, freshness, latency, cost, data volume.
- Design system prompts, few-shot selection, and structured-output contracts as versioned engineering artifacts.
- Explain when chain-of-thought helps, when reasoning models make it redundant, and the token-cost implications.
- Engineer context deliberately: budgets, prefix stability, prompt-cache economics on Anthropic/OpenAI/Bedrock/Vertex, and mitigations for lost-in-the-middle and context rot.
- Reason about LoRA/QLoRA mechanics—rank, alpha, quantized bases—and multi-adapter serving.
- Compare RLHF, DPO, and RLAIF, and defend eval-before/after discipline against catastrophic forgetting.
- Map every rung to AWS (Bedrock customization, SageMaker) and GCP (Vertex AI tuning, context caching), including cost-floor gotchas.
1. The adaptation ladder: decide before you tune
The most common senior-level failure here is reaching for fine-tuning to solve a knowledge problem. Fine-tuning reliably changes behavior—format, tone, task framing, tool-selection habits. It is a poor, unauditable way to inject facts, and it cannot keep them fresh: retrieval delivers knowledge with provenance and an update path measured in minutes, while a tuning run's knowledge is frozen and cannot cite a source. Start every adaptation conversation by classifying the gap.
flowchart TD
A["Capability gap identified"] --> B{"Knowledge gap or behavior gap?"}
B -->|"knowledge that changes"| C["RAG and retrieval grounding - see Chapter 04"]
B -->|"behavior, format, tone, task style"| D{"Does prompting plus few-shot pass the eval bar?"}
D -->|"yes"| E["Ship prompt + context engineering"]
D -->|"no, and 100s to 1000s of labeled examples exist"| F["PEFT such as LoRA or QLoRA"]
F --> G{"Quality still short after PEFT sweep?"}
G -->|"yes, with large data and budget"| H["Full fine-tune of a smaller open model"]
B -->|"cost or latency gap at acceptable quality"| I["Distill to a smaller student model"]
H --> I
Prompting + context
Zero training cost, instant iteration, fully reversible. Ceiling: instruction-following limits, long-instruction token cost, drift across model upgrades.
RAG
Fresh, auditable knowledge with citations and access control; does not fix style, refusals, or format. Depth is Chapter 04—here it is the rung to rule out before tuning.
PEFT (LoRA/QLoRA)
Behavior shaping from hundreds of examples; megabyte adapters; multi-tenant serving. Ceiling: limited deep-domain capacity; base upgrades orphan adapters.
Full fine-tune
Maximum plasticity for deep domain shift on smaller open models with tens of thousands of examples. Costs: GPU budget, forgetting risk, a full model copy per variant.
Distillation
A cost/latency rung, not a quality rung: a teacher labels data, a small student learns the narrow task. Use after quality is proven.
Two criteria interviewers listen for. Data reality: prompting needs zero examples, few-shot 3–10, LoRA is credible from a few hundred vetted pairs, full fine-tuning wants an order of magnitude more. With 40 examples, the answer is prompting plus an eval set, not a training job. Reversal cost: a prompt rollback is a config change; a tuned-model rollback is a deployment with provisioned-capacity implications (Section 8). The ladder is an option-value argument—buy the cheap, reversible option first and let the eval harness (Chapter 06) say when to climb.
2. Prompting as engineering, not incantation
A production system prompt is an API contract: role, capabilities, refusal policy, output format, tool-use policy, and precedence when instructions conflict. Treat it like code—version-controlled, diff-reviewed, regression-tested on every change and model upgrade. The instruction hierarchy matters because user input and retrieved documents are untrusted: system policy outranks user requests, and retrieved content is data, never instructions (injection defense continues in Chapters 05 and 11).
- Structure beats prose — labeled blocks for role, policies, tools, examples, output spec; models follow them more reliably and reviewers can diff them.
- Positive instructions — "respond only with JSON matching the schema" beats piles of "do not" clauses; specify desired behavior plus one fallback (how to abstain).
- Few-shot is bias injection — exemplars anchor format, label distribution, and length; ordering effects are real, so shuffle-test during eval.
- Static vs. dynamic exemplars — per-query nearest-neighbor exemplars lift heterogeneous tasks but destroy cache prefix stability (Section 5); measure whether the lift beats the cache loss.
- Upgrade drift — a prompt is an implicit dependency on a checkpoint; pin versions and gate upgrades on eval regressions, not changelogs.
Chain-of-thought, and when reasoning models retire it
Chain-of-thought prompting (Wei et al., 2022) elicits intermediate reasoning that improves multi-step tasks on models not trained to reason by default. In 2026 the landscape is split: reasoning-first models (OpenAI o-series, Claude extended thinking, Gemini thinking) already run an internal reasoning phase; prepending "think step by step" is at best redundant, at worst interference—and reasoning tokens bill as output even when hidden. The senior position: CoT remains a lever for small or non-reasoning models and for auditable rationales; with reasoning models the lever is the thinking-budget parameter, tuned like any latency/cost knob. Never treat emitted rationales as faithful traces of the computation—they are artifacts, not ground truth.
| Situation | Reasoning approach | Why |
|---|---|---|
| Small/distilled model, multi-step task | Explicit CoT or few-shot rationales | Model won't reason unprompted; rationale tokens buy accuracy. |
| Reasoning model, hard task | Set thinking budget; no CoT boilerplate | Reasoning is trained; budget is the control surface. |
| Reasoning model, trivial task at scale | Minimal budget or non-reasoning tier | Reasoning tokens dominate cost with no gain. |
| Regulated flow needing audit trail | Structured rationale field in the schema | You need a stored artifact, not musing. |
3. Structured output and tool-schema design
Most enterprise LLM calls are machine-to-machine. Three enforcement tiers, in increasing strength: (1) formatting instructions plus a validator and retry-with-repair; (2) provider "JSON mode," guaranteeing syntactic JSON but not your schema; (3) constrained decoding against a declared schema—OpenAI structured outputs (docs), Gemini responseSchema on Vertex, strict tool schemas on Anthropic and Bedrock—where the sampler masks tokens that would violate the grammar.
Constrained decoding removes parse failures, not semantic failures: the model can emit a schema-perfect object with wrong values, and over-tight schemas can degrade content quality by forcing commitments before evidence. Keep the schema as loose as the consumer allows, make uncertainty representable (nullable fields, an explicit abstain reason), and keep application-level validation as the final authority.
Tool schemas are prompts with types
Function-calling quality is dominated by schema design, because the model chooses tools by reading names and descriptions. Rules that hold across Anthropic tool use (docs), Bedrock Converse, and Vertex function calling:
- Few, sharp tools — selection error grows with catalog size and overlap; merge near-duplicates and say when to use each tool and when not to.
- Enums over free strings — every free-text parameter is a hallucination surface.
- Flat over nested — deep optional nesting multiplies invalid-combination states.
- Declare side effects — read-only vs. mutating drives retries and confirmation gates (Chapter 05).
- Schemas cost tokens every call — a 20-tool catalog can be thousands of prompt tokens; exactly the stable prefix caching amortizes.
4. Context engineering: budgets, rot, compaction, memory
A 200K–2M token window is capacity, not a strategy. Context engineering decides what earns a place in the window, in what order, and what happens when the conversation outlives it. Two empirical failure modes drive the discipline. Lost-in-the-middle: models retrieve best from the beginning and end of long contexts, with a U-shaped accuracy curve over position (Liu et al., 2023). Context rot: performance degrades as the window fills—well below the advertised limit—because distractors dilute attention; needle-in-a-haystack benchmarks are saturated and don't predict this, so test with realistic multi-fact tasks at your operating lengths.
The budget above (illustrative, not a standard) encodes two rules: stable content first, so the prefix is byte-identical across calls and cacheable; volatile content last, for cache mechanics and because the end of context is a high-attention position for the current question. Trigger compaction at an explicit threshold—say 60–70% of the window (a heuristic)—because quality decays before capacity runs out and you must reserve output headroom.
flowchart TD
H["Conversation history grows"] --> C{"Context usage above budget threshold?"}
C -->|"no"| ASM["Assemble prompt - stable prefix first, volatile tail last"]
C -->|"yes"| SUM["Compact older turns into a structured summary block"]
SUM --> MEM["Persist durable facts and decisions to external memory"]
MEM --> ASM
ASM --> CALL["Model call"]
CALL --> H
Compaction is lossy summarization under a contract: preserve decisions, constraints, open questions, and identifiers; drop pleasantries and superseded drafts; log what was dropped—silent compaction is a debugging nightmare. Across sessions, memory splits into a scratchpad (working notes), episodic memory (past sessions, retrieved like RAG), and semantic memory (distilled stable facts, editable and inspectable). Memory retrieval inherits every relevance and access-control problem from Chapter 04—cross-tenant memory leakage is a security incident, and stale memory is a quality bug only eval traces catch.
5. Prompt caching: mechanics and cache-hit economics
Prompt caching stores the computed KV state (Chapter 02) of a prompt prefix so requests sharing that exact prefix skip most prefill compute. It is the highest-leverage cost/latency optimization in prompt-heavy systems—agents with big tool catalogs, long system prompts, fixed-document Q&A. Provider mechanics differ, and the differences are interview material.
sequenceDiagram
participant C as "Client"
participant R as "API frontend"
participant K as "Prefix cache"
participant G as "GPU prefill"
C->>R: Request 1 with cache breakpoint after tools
R->>K: Look up prefix hash
K-->>R: Miss
R->>G: Full prefill of entire prompt
G-->>K: Store KV state for prefix
G-->>C: Response billed at write rate for prefix
C->>R: Request 2 with identical prefix plus new question
R->>K: Look up prefix hash
K-->>R: Hit within TTL
R->>G: Prefill only the new suffix tokens
G-->>C: Faster response with discounted prefix tokens
| Platform | Control model | Economics (as of checked date — verify pricing pages) |
|---|---|---|
| Anthropic API | Explicit cache_control breakpoints; ~1024-token minimum prefix; TTL refreshes on hit (docs) | Writes ~1.25× base input (5-min TTL; 1-hour tier costs more); reads ~0.1× base input |
| OpenAI API | Automatic for prompts ≥1024 tokens; prefix stability is your only lever (docs) | No write premium; cached input discounted 50–75% by model |
| Amazon Bedrock | Cache checkpoints on Claude and Nova families; works with Converse (docs) | AWS cites up to ~90% cost and ~85% latency reduction on cached tokens |
| Vertex AI (Gemini) | Implicit caching by default plus explicit CachedContent for fixed corpora (docs) | Cached tokens ~75% discount; explicit caches add per-token-hour storage |
Do the break-even aloud in an interview. With Anthropic-style example rates (1.25× write, 0.1× read), N reuses cost 1.25 + 0.1·(N−1) versus N uncached—caching wins from the second use within the TTL, if the prefix repeats byte-for-byte. That conditional is where systems fail:
- Cache busters — a timestamp, request ID, or user name interpolated early invalidates every prefix; put dynamic content after the last breakpoint.
- Dynamic few-shot — per-query exemplars change the prefix every call; pin per segment or accept the loss knowingly.
- TTL vs. traffic cadence — a 5-minute TTL amortizes in a busy agent session and never hits for hourly visitors; match TTL tier to arrival patterns.
- Hit rate is an SLO — providers return cached-token counts; a deploy that reorders prompt sections can silently zero the hit rate and double the bill. Alert on it.
6. LoRA, QLoRA, and adapter serving
LoRA (Hu et al., 2021) freezes base weights and learns a low-rank update: tuned behavior is W + (α/r)·B·A, with A and B thin matrices of rank r. The bet—empirically sound for behavior shaping—is that task adaptation lives in a low-dimensional subspace. Trainable parameters drop to ~0.1–1%; the artifact is megabytes. Rank r (commonly 8–64) sets capacity; α scales the update, and practitioners tune the α/r ratio rather than either alone. More rank helps until the task's intrinsic dimension is covered, then buys overfitting risk; which modules you adapt (attention projections vs. all linear layers) often moves results more than r.
QLoRA (Dettmers et al., 2023) makes the economics accessible: quantize the frozen base to 4-bit NF4, backpropagate into full-precision adapters, use paged optimizers—the paper fine-tuned a 65B model on a single 48GB GPU. The subtlety to name: you trained against a quantized base, so evaluate in the exact serving configuration you'll deploy; train/serve quantization mismatch is a real regression source (quantized serving is Chapter 02).
| Dimension | Full fine-tune | LoRA | QLoRA |
|---|---|---|---|
| Trainable params | 100% | ~0.1–1% | ~0.1–1% (4-bit frozen base) |
| GPU memory | Highest | Moderate | Lowest — single-GPU for mid-size models |
| Artifact | Full copy per variant | MBs per adapter | MBs per adapter |
| Forgetting risk | Highest | Lower (base frozen) | Lower (base frozen) |
| Serving | Dedicated deployment | Multi-adapter over shared base | Multi-adapter over shared base |
| Best for | Deep domain shift, open models | Behavior/format/tone | Same, tight GPU budgets |
flowchart LR
RQ["Requests tagged with adapter id"] --> SCH["Continuous-batch scheduler"]
SCH --> BASE["Shared frozen base weights on GPU"]
subgraph SG1["Adapter pool in GPU and host memory"]
A1["LoRA - support triage"]
A2["LoRA - SQL generation"]
A3["LoRA - claims summaries"]
end
SCH -->|"attach per request"| A1
SCH -->|"attach per request"| A2
SCH -->|"attach per request"| A3
BASE --> OUT["Batched decode across tenants"]
A1 --> OUT
A2 --> OUT
A3 --> OUT
Multi-LoRA serving usually decides the PEFT-vs-full-FT debate on multi-tenant platforms: systems in the S-LoRA line (Sheng et al., 2023) and engines like vLLM batch requests for different adapters through one shared base, paging adapters between host and GPU memory. Fifty tenant variants become fifty small files on one fleet, not fifty deployments—at the cost of small per-token overhead, cold-swap latency for rare adapters, and a registry mapping tenant → adapter version → base version. That mapping is the sleeper issue: adapters couple to the exact base checkpoint, so a base upgrade is a coordinated retrain-and-reeval event across every adapter.
7. Preference tuning, distillation, and the forgetting problem
SFT teaches what a good answer looks like when a gold answer exists. Preference tuning teaches choosing between plausible answers—helpfulness, tone, safety—where quality is comparative. RLHF as productionized by InstructGPT (Ouyang et al., 2022) trains a reward model on preference pairs, then optimizes the policy with PPO under a KL penalty to the reference. It works and is operationally heavy: a second model to train, RL instability, and reward hacking—the policy exploiting reward-model blind spots such as verbosity and sycophancy.
flowchart LR
S["SFT checkpoint"] --> G["Sample candidate responses"]
G --> L["Human or AI preference labels"]
L --> PP["Preference pairs - chosen vs rejected"]
PP -->|"RLHF path"| RM["Train reward model"]
RM --> PO["PPO updates with KL penalty to reference"]
PP -->|"DPO path"| DL["Direct preference loss - no reward model, no RL loop"]
PO --> AM["Aligned model"]
DL --> AM
AM --> EV["Win-rate evals plus general-capability regression suite"]
DPO (Rafailov et al., 2023) collapses the pipeline: a closed-form objective optimizes directly on preference pairs—no reward model, no RL loop, far simpler to reproduce. The trade: no reusable reward model for filtering or online scoring, and quality bounded by the static pair set rather than on-policy sampling. RLAIF—AI feedback replacing human labels, as in Constitutional AI (Bai et al., 2022)—scales label volume while inheriting labeler bias. Interview depth is which knob solves which problem: SFT for task format, preference methods for comparative qualities, DPO as the pragmatic first choice because its failures are visible in data rather than RL dynamics.
Distillation: buying back cost and latency
Distillation transfers a narrow capability from a large teacher to a small student—training on teacher outputs and rationales (Hsieh et al., 2023), or on logits where weights allow. It sits last on the ladder because it presumes you know what good looks like: the teacher sets the ceiling, and the student inherits teacher errors invisibly unless evals cover the tails. Managed offerings (Bedrock Model Distillation; teacher generation plus Vertex tuning) compress the workflow, but the eval obligation stays yours—and distilling proprietary outputs into a competitor model is a licensing question.
Catastrophic forgetting and eval discipline
Every gradient that makes the model better at your task makes it different at everything else. Forgetting shows up as degraded instruction following, lost multilingual ability, or safety drift after a narrow tune—invisible if you only measure the target task. The non-negotiable discipline:
- Baseline firstFreeze a task eval and a general regression suite (instruction following, safety refusals, benchmark slice); score the base model.
- Train with mitigationsPrefer PEFT; mix general instruction data into the tuning set; early-stop on validation.
- Score both suites afterAccept only if task lift clears the bar AND regressions stay within budget; record both in the model card.
- Canary in servingSmall traffic slice with task metrics and safety monitors (Chapter 11) before ramp.
Synthetic data feeds every rung and needs its own gates: generate with a strong teacher from seeds and real failures; filter with programmatic verifiers then a calibrated LLM judge (Chapter 06); deduplicate; decontaminate against eval sets. A small vetted-human set beats a large unfiltered synthetic dump; the ratio you can defend with an ablation is the senior answer.
8. Cloud mapping: tuning and caching on AWS and GCP
Platform and forward-deployed interviews expect the ladder landed on real services—and the cost-model fine print that changes the recommendation. Full stack tours are Chapters 08 and 09; this is the adaptation-specific mapping.
AWS
- Bedrock custom modelsmanaged fine-tuning and continued pre-training
- Bedrock Model Distillationmanaged teacher→student distillation
- Bedrock Custom Model Importserve weights tuned elsewhere
- SageMaker training jobs / HyperPodfull-control LoRA/QLoRA/full FT on open weights
- Bedrock prompt cachingcache checkpoints on Claude and Nova families
Google Cloud
- Vertex AI supervised tuningmanaged LoRA-based tuning for Gemini
- Vertex AI custom trainingGPU/TPU jobs for open-weight PEFT and full FT
- Model Gardenopen-weight models with tuning and deployment recipes
- Vertex context cachingimplicit caching plus explicit CachedContent API
- Gemini responseSchemastructured-output enforcement at the API level
The portable logic: managed tuning when you want provider models and minimal MLOps; roll-your-own on SageMaker/Vertex custom training for open weights, custom losses (DPO is often DIY territory), multi-LoRA economics, or portability. Either way, eval-before/after is yours—no managed service owns your regression suite.
Interview playbook
For any "should we fine-tune?" question, walk the LADDER:
- L — Locate the gap: knowledge vs. behavior vs. cost/latency; tuning is for behavior, retrieval for knowledge.
- A — Attempt the cheapest rung: prompt + context engineering against a frozen eval with a stated pass bar.
- D — Data audit: vetted example count, labeling ownership, licensing, synthetic-augmentation defensibility.
- D — Decide with exit criteria: "LoRA if the baseline stalls below X with ≥500 examples; full FT only for deep domain shift on open weights."
- E — Evaluate before and after: task eval plus general regression to catch forgetting; canary rollout.
- R — Runtime and cost model: adapter vs. full-copy serving, provisioned floors, cache hit-rate impact, base-upgrade coupling.
Senior signals
- Quantified caching arguments: write premium vs. read discount, TTL vs. cadence, hit rate as a monitored SLO.
- Prompts and schemas as versioned, regression-tested artifacts pinned to model versions.
- Reasoning models move CoT from prompt text to a thinking-budget parameter—with the output-cost implication.
- Naming adapter/base version coupling and provisioned-capacity floors before recommending tuning.
Common traps
- Fine-tuning to "teach the model our docs"—a knowledge/freshness problem retrieval solves with provenance.
- Claiming constrained decoding guarantees correct output—it guarantees shape, not semantics.
- A timestamp at the top of the system prompt, and a mysteriously doubled cache bill.
- Reporting only target-task metrics after tuning, with no regression suite.
- "DPO is better than RLHF" as an absolute rather than a complexity/control trade-off.
Question bank
Questions actually asked for prompting, context engineering, and model adaptation at senior level.
Q1A product team wants to fine-tune a model on the company wiki so it "knows our products." Respond.
Strong answer outline
- Classify the gap: knowledge that changes—tuning bakes stale facts with no citations or access control.
- Propose RAG for knowledge with provenance and freshness; reserve tuning for behavior gaps.
- Offer the ladder with exit criteria and the eval set that would justify climbing.
Follow-up probes
- When would tuning on the wiki ever be right?
- How would you prove RAG is sufficient?
Pass if knowledge-vs-behavior is the spine and an eval bar is named; fail on a generic "RAG is cheaper."
Q2Design the prompt-caching strategy for an agent with a 20-tool catalog and a long system prompt.
Strong answer outline
- Order by stability: system prompt → tool schemas → static exemplars → memory → volatile history/query; breakpoints after stable blocks.
- Eliminate cache busters (timestamps, request IDs) from the prefix; pin exemplars or justify the trade.
- Match TTL to session cadence; monitor cached-token counts as an SLO; estimate savings with write/read arithmetic.
Follow-up probes
- What changes on OpenAI's automatic caching vs. Anthropic's explicit breakpoints?
- A deploy halves your hit rate—first thing you inspect?
Pass if ordering, busters, TTL, and monitoring all appear; fail if the answer is "turn on caching."
Q3Do the math: when does prompt caching pay for itself?
Strong answer outline
- Example rates: write ~1.25× base input, read ~0.1× (Anthropic-style; OpenAI has no write premium).
- N uses cost 1.25 + 0.1(N−1) cached vs. N uncached → any real reuse within TTL wins.
- The dominant variable is hit rate—prefix churn and TTL misses destroy the economics; caching cuts TTFT, not decode.
Follow-up probes
- How does a per-token-hour storage fee (Vertex explicit caches) change the model?
- What hit rate would you demand before relying on caching in capacity planning?
Pass if arithmetic is performed and hit rate identified as dominant; fail on vague "caching saves money."
Q4What are lost-in-the-middle and context rot, and how do you design around them?
Strong answer outline
- Lost-in-the-middle: U-shaped accuracy over position (Liu et al.); context rot: quality decays as the window fills, below the advertised limit.
- Mitigations: retrieve/select rather than stuff, critical evidence and the question near the end, compact at a utilization threshold.
- Test with realistic multi-fact tasks at operating lengths—needle-in-a-haystack is saturated.
Follow-up probes
- Does a bigger window remove the need for RAG?
- How would you detect rot in production traces?
Pass if both phenomena are distinguished with positional and budgetary mitigations; fail on "use a bigger model."
Q5Explain LoRA and QLoRA: what do rank and alpha control, and what risk does QLoRA add?
Strong answer outline
- Freeze W, learn ΔW = (α/r)·B·A; ~0.1–1% trainable params; MB artifacts. r sets capacity (8–64 typical); tune the α/r ratio; target-module choice often matters more than r.
- QLoRA: 4-bit NF4 frozen base, full-precision adapters, paged optimizers—single-GPU fine-tuning for mid-size models.
- Risk: train/serve quantization mismatch—evaluate in the exact deployment configuration.
Follow-up probes
- Why might doubling r not improve results?
- You'll serve the base in 8-bit—does the QLoRA adapter transfer cleanly?
Pass if low-rank intuition, serving consequences, and the mismatch risk all appear; fail if it's only "efficient fine-tuning."
Q6Design serving for 50 tenant-specific model variants.
Strong answer outline
- Multi-LoRA over one shared frozen base (S-LoRA/vLLM pattern): 50 adapters as small files, one fleet, batching across adapters.
- Operational needs: adapter registry (tenant→adapter→base version), cold-swap latency budget, per-tenant eval gates.
- Name the coupling: a base upgrade forces coordinated retrain/reeval of all adapters; compare cost with 50 deployments.
Follow-up probes
- What breaks if two tenants need different base models?
- How do you canary one tenant's new adapter?
Pass if shared-base economics and version coupling both appear; fail if the answer is 50 endpoints.
Q7RLHF vs. DPO: differences, and which would you run first?
Strong answer outline
- RLHF: reward model + PPO with KL leash—powerful, reusable reward model, operationally heavy, reward-hacking risk.
- DPO: closed-form loss on pairs—no RM, no RL loop, reproducible; bounded by static pair quality, no reusable scorer.
- Default DPO first for applied teams; mention RLAIF for scaling labels with inherited bias.
Follow-up probes
- What is reward hacking and how do you detect it?
- When is the reusable reward model worth RLHF's complexity?
Pass if the trade-off is complexity/control, not "DPO is newer"; fail on absolutes.
Q8How do you detect and mitigate catastrophic forgetting after a fine-tune?
Strong answer outline
- Detection: frozen general regression suite (instruction following, safety refusals, benchmark slice) scored before and after, next to the task eval.
- Mitigation: prefer PEFT, mix general instruction data, early-stop on validation.
- Governance: accept within an agreed regression budget; canary with safety monitors.
Follow-up probes
- Task accuracy up 6 points, refusal rate down 8—ship it?
- Why does LoRA reduce but not eliminate forgetting?
Pass if the before/after dual-suite discipline is explicit; fail if forgetting is defined but not operationalized.
Q9Compare JSON mode, constrained decoding, and tool-calling for structured extraction.
Strong answer outline
- JSON mode: syntactic JSON only. Constrained decoding (structured outputs, responseSchema): grammar-level schema enforcement. Tool-calling: schema enforcement plus multi-action framing.
- Constrained decoding removes parse failures, not semantic errors; over-tight schemas can hurt content quality.
- Application validator as final authority, one repair retry, dead-letter path.
Follow-up probes
- How do you represent "not found" without inviting hallucinated values?
- What breaks with an unbounded free-string field?
Pass if shape-vs-semantics is crisp; fail if constrained decoding is called a correctness guarantee.
Q10Is chain-of-thought prompting still relevant with reasoning models?
Strong answer outline
- Yes for small/non-reasoning models and audit-trail rationales; largely no as boilerplate for reasoning models, where the control is the thinking budget.
- Reasoning tokens bill as output—budget them per task tier like a latency/cost knob.
- Emitted rationales aren't faithful traces; use as artifacts, verify with evals.
Follow-up probes
- How would you set thinking budgets across a mixed-difficulty workload?
- When would you route to a non-reasoning tier entirely?
Pass if the answer is model-conditional with a cost dimension; fail on a blanket yes or no.
Q11Compare the tuning paths and cost gotchas on Bedrock vs. Vertex AI.
Strong answer outline
- AWS: Bedrock customization (fine-tune/distill) with the provisioned-throughput serving floor historically attached to custom models—do utilization math; SageMaker for open-weight control.
- GCP: Vertex supervised tuning is adapter-based; tuned Gemini serves at base per-token pricing per current docs—different break-even.
- Both: eval-before/after is customer-owned; verify current docs since serving options evolve.
Follow-up probes
- At 100K requests/day vs. 1K/day, does the recommendation change?
- When does Custom Model Import beat tuning inside Bedrock?
Pass if the provisioned-floor vs. per-token difference is named; fail on a feature-list comparison.
Q12Design a synthetic data pipeline for fine-tuning, with quality controls.
Strong answer outline
- Seed from real failures and vetted human examples; generate variations with a strong teacher under coverage targets.
- Filter: programmatic verifiers first (schema, execution), calibrated LLM judge second; deduplicate; decontaminate against eval sets.
- Validate with an ablation (synthetic vs. mixed vs. human-only); check teacher-output licensing.
Follow-up probes
- How do you prevent the student learning the teacher's systematic errors?
- What human-to-synthetic ratio would you defend, and how?
Pass if filtering, dedup, and decontamination all appear; fail if volume is the only lever.
Proof artifact: an adaptation-ladder bake-off
Build a public, reproducible comparison of three rungs on one task—structured extraction from messy support tickets (public or synthetic data). All results are portfolio measurements, not claims about Purnendu's production history.
Steps
- Freeze the task: extraction JSON schema, 300–500 labeled examples split train/validation/eval, plus a general regression suite (instruction following + safety refusals).
- Variant A — engineered prompt: versioned system prompt, static few-shot, constrained decoding, validator with one repair retry; iterate to plateau, logging every version's score.
- Variant B — dynamic few-shot: nearest-neighbor exemplars; measure the accuracy delta AND the cache hit-rate/cost delta vs. A.
- Variant C — QLoRA tune of an 8B-class open model on the train split via a SageMaker training job or Vertex custom job; serve on a vLLM container with LoRA support.
- Benchmark identically: per-field accuracy, schema-violation rate, p50/p95 latency, cost per 1,000 requests including cache effects, regression deltas for C.
- Write a one-page memo recommending a rung at 1K, 50K, and 1M requests/day with break-even arithmetic shown.
Metrics
Per-field precision/recall, exact match, schema-violation and repair-retry rates, cached vs. uncached token counts and TTFT, cost per 1,000 requests per variant, tuning-cost amortization curve, before/after regression scores for the tuned model.
Deliberate failures
- Insert a timestamp at the top of the system prompt; show the cache hit-rate collapse and cost delta.
- Over-tighten the schema and measure content-quality degradation vs. the loose schema.
- Overtrain the QLoRA run (no early stopping) and show the regression suite catching instruction-following decay while task accuracy still looks fine.
- Evaluate the adapter against a differently quantized serving base and document the silent quality drop.
What to present
The ladder diagram with measured numbers per rung, the cost-vs-volume break-even chart, the cache-buster incident graph, and the forgetting demonstration. The arc—"prompting won at low volume, tuning won at 1M/day, here is the crossover"—is exactly the judgment senior interviews probe.
Chapter review
Adaptation is an economics-and-evidence problem. Classify the gap, buy the cheapest reversible option first, and climb only when a frozen eval says the rung below failed. Engineer prompts and schemas as versioned artifacts; order context for stability and cache economics; treat LoRA as the default tuning tool; and never accept a tuned model without before/after scores on both the task and a regression suite.
Glossary
- Adaptation ladder
- Escalation—prompting, RAG, PEFT, full fine-tune, distillation—governed by gap type, data volume, and reversal cost.
- Context rot
- Task-quality degradation as the window fills, well below the advertised token limit.
- Lost in the middle
- U-shaped retrieval accuracy over position in long contexts; beginnings and ends are privileged.
- Prompt caching
- Reuse of a prefix's computed KV state across requests, discounting cached tokens and cutting TTFT.
- LoRA
- Frozen base weights plus a trained low-rank update (α/r)·B·A; megabyte-scale artifacts.
- QLoRA
- LoRA over a 4-bit-quantized frozen base with paged optimizers; single-GPU fine-tuning of mid-size models.
- Multi-LoRA serving
- Batching requests for many adapters through one shared base model on the same GPUs.
- DPO
- Direct preference optimization: closed-form loss on preference pairs, no reward model or RL loop.
- Catastrophic forgetting
- Degradation of general capabilities from narrow fine-tuning; detected only by regression suites.
- Constrained decoding
- Sampling restricted by a schema/grammar—shape guarantee, not semantic correctness.
Mastery checklist
- I can classify a requirement as knowledge, behavior, or cost/latency and pick the rung with exit criteria.
- I can design system prompts and tool schemas as versioned, regression-tested artifacts.
- I can say when CoT helps, when a thinking budget replaces it, and the billing implication.
- I can order a context window for cache stability and positional attention, with a compaction threshold.
- I can do prompt-cache break-even arithmetic and name the top three cache busters.
- I can explain rank, alpha, and target modules in LoRA and the QLoRA train/serve quantization risk.
- I can design multi-LoRA serving for many tenants and name the base-version coupling.
- I can compare RLHF, DPO, and RLAIF as complexity/control trade-offs.
- I can run eval-before/after and defend a regression budget.
- I can map every rung to Bedrock/SageMaker and Vertex AI, including provisioned-throughput vs. per-token serving.
Primary sources
Links checked . Pricing and model support change frequently; verify current pricing pages before quoting numbers.
- Anthropic — prompt caching: breakpoints, TTLs, and pricing multipliers
- Anthropic — tool use and schema design
- OpenAI — automatic prompt caching
- OpenAI — structured outputs and strict schemas
- Amazon Bedrock — prompt caching
- Amazon Bedrock — model customization
- Amazon Bedrock — model distillation
- Amazon SageMaker — training jobs
- Vertex AI — context caching overview
- Vertex AI — Gemini model tuning overview
- Wei et al. — Chain-of-Thought Prompting Elicits Reasoning in LLMs
- Liu et al. — Lost in the Middle: How Language Models Use Long Contexts
- Hu et al. — LoRA: Low-Rank Adaptation of Large Language Models
- Dettmers et al. — QLoRA: Efficient Finetuning of Quantized LLMs
- Sheng et al. — S-LoRA: Serving Thousands of Concurrent LoRA Adapters
- Ouyang et al. — Training language models to follow instructions (InstructGPT/RLHF)
- Rafailov et al. — Direct Preference Optimization
- Bai et al. — Constitutional AI: Harmlessness from AI Feedback
- Hsieh et al. — Distilling Step-by-Step
CHAPTER 04 · PRIORITY 0
Retrieval, Vector Search & Production RAG
36 min read · 16 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- Decompose a RAG request into candidate generation, ranking, context construction, generation, and verification—and locate failures at the right stage.
- Explain dense, sparse, and hybrid retrieval—including RRF fusion and reranking—using concrete failure cases and held-out judgments.
- Reason about vector geometry, HNSW, filtered ANN, and quantization as measured trade-offs, not defaults.
- Apply 2026-era upgrades with judgment: contextual retrieval, late chunking, GraphRAG, multimodal retrieval, and agentic retrieval loops.
- Design freshness pipelines, semantic answer caches, and a cost model that identifies the dominant spend in a RAG system.
- Map a retrieval architecture onto managed AWS and GCP services and defend an honest build-vs-managed decision.
- Operate or migrate a production search service with incremental indexing, tenancy, backups, canary rollout, and monitoring.
1. Start with the retrieval contract
A production RAG system is not “an LLM plus a vector database.” It is an evidence-selection system followed by a constrained answer generator. Define the retrieval contract before choosing an embedding model: given a query, authorization context, freshness boundary, and latency budget, return a ranked set of evidence units with stable identifiers and provenance. Generation may then answer only from those units—or abstain.
flowchart LR
subgraph SG1["Ingest path"]
S1["Source of truth"] --> P1["Parse"]
P1 --> C1["Segment + enrich"]
C1 --> E1["Embed + index"]
E1 --> V1["Versioned index"]
end
subgraph SG2["Serve path"]
Q1["Query"] --> A1["Authorize"]
A1 --> R1["Retrieve candidates"]
R1 --> F1["Fuse + rerank"]
F1 --> K1["Pack context"]
K1 --> G1["Generate + cite"]
end
V1 --> R1
G1 --> OBS["Judgments + traces"]
R1 --> OBS
OBS -->|"failure analysis"| C1
This decomposition creates useful fault boundaries. If the correct passage never enters the candidate set, prompt tuning cannot recover it. If a relevant passage is retrieved but buried, inspect fusion or reranking. If strong evidence is packed but the answer contradicts it, inspect generation and grounding controls (Chapter 6 covers the evaluation machinery). If the answer is faithful but stale, inspect indexing freshness. If a cross-tenant passage appears, treat it as a security incident, not a relevance defect.
Define the evidence unit
A “document” is rarely the right ranking unit. A policy page may contain a definition, exceptions, a table, and an effective date. Store chunks with document ID, section path, source URI, version, access-control attributes, offsets, content hash, and timestamps. Stable IDs enable deduplication, citation repair, incremental updates, and evaluation across re-indexes. Keep the primary content store authoritative; a search index is usually a derived projection.
2. Design chunks, queries, and context together
Chunk size is not a universal token constant. It mediates two competing risks: small chunks lose the conditions that make a statement true; large chunks dilute the matching signal and consume the generation budget. Start from semantic boundaries—headings, paragraphs, table rows, code symbols, ticket threads—then measure. Store parent relationships so retrieval can find a small unit and context expansion can include the surrounding section.
Fixed windows
Simple and fast; useful as a baseline. They can split tables, procedures, or definitions from exceptions. Overlap reduces boundary loss but increases duplicates and index cost.
Structure-aware
Preserves sections, lists, code symbols, or table units. Parsing is harder, and malformed documents need fallbacks, but citations and context coherence improve.
Parent-child
Index compact child units and expand to a parent after ranking. It separates match granularity from reading granularity, at the cost of another packing decision.
Multi-vector
Represent one object with body, title, summary, image, or late-interaction vectors. Recall may improve while storage, query fan-out, and evaluation complexity rise.
Query transformation is a hypothesis
Rewriting can normalize spelling, resolve a conversational reference, expand an acronym, or convert a question into search-oriented language; decomposition can retrieve evidence for separate subquestions. Each transformation can also erase a product code, invent intent, or leak unauthorized conversation context. Preserve the original query and protected literals (IDs, dates, jurisdictions, negations), log and version transformations, cap fan-out, and evaluate transformed and untransformed variants on the same judgments. A rewrite that produces fluent language while dropping “EU” or “2026” is a regression, not an improvement.
Pack context as a budgeted ranking problem
Do not concatenate the first k chunks blindly. Deduplicate near-identical passages; group adjacent units; favor coverage of distinct subquestions; reserve tokens for instructions and response; and include provenance outside the quoted text. A simple packing heuristic can maximize reranker score plus subtopic coverage minus redundancy and token cost. Evaluate whether evidence survives packing, not merely whether retrieval found it—packing is also the dominant lever on per-answer generation cost (Section 7).
3. Contextual retrieval and late chunking
Classic chunking has a structural defect: a chunk is embedded in isolation, so pronouns, abbreviations, and section-scoped conditions lose their referents. “The termination clause above does not apply to contractors” embeds poorly when “above” is in a different chunk. Two 2024–2025 techniques attack this at ingest time and are now standard interview material.
Contextual retrieval, described by Anthropic, uses an LLM at ingestion to generate a short document-situating preamble for each chunk (“This clause is from the 2026 EU contractor policy, section 4, on notice periods…”), prepends it before embedding, and does the same for the BM25 index. Anthropic reports roughly a 49% reduction in top-20 retrieval failure rate for contextual embeddings plus contextual BM25, and about 67% when a reranker is added—vendor-reported numbers, but directionally consistent with public replications. The cost is one LLM call per chunk at ingest; prompt caching (Chapter 2) makes this cheap because the full document is the cached prefix and only the chunk varies.
Late chunking, introduced by Jina AI, inverts the order of operations: run the whole document through a long-context embedding model once, then mean-pool token embeddings per chunk boundary afterward. Each chunk vector is conditioned on the surrounding text without any extra LLM calls—but it requires an embedding model with a long context window and does not improve the lexical index.
flowchart TD
D["Full document"] --> N1["Naive: split, then embed each chunk alone"]
D --> C2["Contextual: prepend LLM-generated document context, then embed"]
D --> L3["Late chunking: embed whole document, then pool token vectors per chunk"]
N1 --> X1["Chunk vector loses global references"]
C2 --> X2["Chunk carries entity, section, and date context"]
L3 --> X3["Chunk conditioned on surrounding text at no LLM cost"]
| Technique | Ingest cost | Helps lexical index? | When it pays |
|---|---|---|---|
| Naive chunk embedding | Embedding only | n/a | Self-contained units: FAQs, tickets, short articles. |
| Contextual retrieval | One cached LLM call per chunk | Yes (contextual BM25) | Reference-heavy corpora: policies, contracts, codebases, long reports. |
| Late chunking | Long-context embedding pass | No | Long documents where re-embedding budget is tight and a long-context embedder is available. |
4. Dense, sparse, hybrid, and reranked retrieval
Sparse lexical retrieval rewards token overlap and is especially strong for identifiers, names, rare terms, and exact phrases. BM25-style scoring balances term frequency, document frequency, and length normalization. Dense retrieval maps queries and passages into a vector space, recovering conceptual similarity and paraphrase. Neither dominates across all query types. Hybrid retrieval builds independent candidate lists and combines them.
| Query | Likely strength | Characteristic failure |
|---|---|---|
ERR_AUTH_0417 | Sparse/exact | Dense representation smooths away a rare identifier. |
| “Why does login work locally but fail behind the proxy?” | Dense | Lexical search misses passages framed as forwarded-header configuration. |
| “EU leave carryover 2026” | Hybrid + filters | Dense misses year/entity; sparse misses paraphrased policy language. |
| Broad comparison with many constraints | Hybrid + reranker | Cheap retrievers cannot jointly reason over all constraints. |
Fuse ranks before comparing incompatible scores
Dense cosine scores and sparse scores do not share a calibrated scale. A raw weighted sum can be dominated by whichever retriever emits larger values. Reciprocal rank fusion (RRF) uses positions instead: for each document, add 1 / (c + rank) across result lists, where c dampens top-rank differences. Qdrant’s Query API supports hybrid and multi-stage retrieval with prefetches and RRF/DBSF fusion; Elasticsearch likewise documents RRF as a hybrid-search fusion option. Tune weighted fusion on held-out judgments, not intuition.
# Illustrative, one-based ranks and c = 60
dense = {"policy-A": 1, "faq-B": 2, "policy-C": 3}
sparse = {"policy-C": 1, "policy-A": 2, "memo-D": 3}
def rrf_score(doc_id):
ranks = [ranking[doc_id] for ranking in (dense, sparse)
if doc_id in ranking]
return sum(1 / (60 + rank) for rank in ranks)
# policy-A and policy-C gain support from both lists.
RRF is robust but discards score magnitude. Score-distribution normalization or learned fusion can exploit more information, but it needs validation and drift monitoring. Always retrieve deeper than the final k; fusion and reranking cannot select a document absent from every candidate list.
Rerank selectively
A cross-encoder or LLM reranker jointly inspects query and candidate and can resolve nuanced constraints. It adds cost and tail latency, and a reranker cannot repair missing candidates. Cache only when the query, corpus/index version, authorization scope, and reranker version make reuse safe. Test a cheap deterministic reranker or metadata boost as a baseline before adding another model call.
5. Vector geometry and approximate nearest neighbors
The similarity function must match model training and stored-vector treatment. Cosine compares direction; dot product combines direction and magnitude; Euclidean distance measures geometric separation. For unit-normalized vectors, ranking by cosine and dot product is equivalent, and squared Euclidean distance is monotonically related. Do not normalize reflexively if magnitude carries trained meaning. Record model, dimensions, preprocessing, normalization, and distance metric as one versioned contract.
Exact search is the quality oracle
An exact scan computes distances against all eligible vectors and provides the reference neighbor set for measuring approximate recall. It is often practical for small or tightly filtered subsets. Approximate nearest-neighbor (ANN) indexes trade perfect recall for lower latency and resource use. Keep an exact path in the benchmark environment; without it, you cannot tell whether missed results come from the embedding or the index.
flowchart TD
subgraph SG3["Layer 2 - sparsest"]
EPT["Entry point"] --> H1["Greedy hop toward query"]
end
subgraph SG4["Layer 1 - denser"]
H2["Descend and refine locally"]
end
subgraph SG5["Layer 0 - full graph"]
H3["Explore ef_search candidates"] --> RES["Top-k approximate neighbors"]
end
H1 --> H2
H2 --> H3
Hierarchical Navigable Small World (Malkov & Yashunin) constructs a multi-layer proximity graph. m controls graph connectivity and therefore memory/build/search behavior; ef_construct expands the build-time candidate pool; ef_search (often surfaced as hnsw_ef or ef) expands the query-time search. Higher values commonly improve recall but cost build time, memory, or latency. Benchmark the actual filtered workload rather than repeating defaults. Google’s ScaNN takes a different route—partitioning plus anisotropic quantization—and underlies Vertex AI Vector Search and AlloyDB’s ANN index (Section 10); the tuning story is different, the recall-versus-latency discipline identical.
Filtering changes the graph problem
A post-filter may leave too few candidates; a strict filter can also make graph traversal ineffective. Build indexes for frequent metadata filters and test selectivity slices. Qdrant documents a filterable HNSW approach and recommends creating payload indexes before ingestion so filter-aware graph edges can be built. In pgvector, approximate-index filtering is applied after the scan; its README describes iterative scans as a way to search farther when filtering removes candidates. These implementation differences belong in a database decision and benchmark.
Quantization is an end-to-end trade
Quantization compresses vector representations to reduce memory and often accelerate distance work, at the cost of approximation error. Preserve original vectors when a two-stage search can rescore a larger compressed candidate set. Qdrant documents scalar, product, and binary approaches and explicitly frames the choice as accuracy, storage, and speed. Measure relevance, ANN recall against exact search, p50/p95/p99 latency, build time, and resident memory—not only compression ratio. The arithmetic is a first-class cost lever: see Section 7.
6. Beyond flat retrieval: GraphRAG, multimodal, and agentic loops
Flat top-k retrieval assumes the answer lives in a handful of independently rankable passages. Three query families break that assumption, and by 2026 interviewers expect you to know which upgrade fixes which family—and what each one costs.
GraphRAG: when relationships are the evidence
Microsoft’s GraphRAG uses an LLM at ingest to extract entities and relations into a knowledge graph, clusters it into communities, and pre-summarizes each community. “Local” queries traverse from an entity through its neighborhood; “global” queries (“what are the recurring risk themes across these filings?”) map over community summaries—questions flat retrieval simply cannot answer because no single passage contains the answer. The price is steep: LLM extraction over the whole corpus at ingest, graph maintenance on every update, and a much harder evaluation problem. The honest default remains flat hybrid retrieval; reach for graphs when queries are genuinely multi-hop or aggregative and the corpus is entity-dense (compliance, biomedical, org knowledge, incident histories).
flowchart LR
DOCS["Corpus"] --> EX["LLM entity + relation extraction"]
EX --> KG["Knowledge graph"]
KG --> CM["Community detection + summaries"]
QL["Local query about one entity"] --> KG
QG["Global query about themes"] --> CM
KG --> AN1["Neighborhood evidence"]
CM --> AN2["Corpus-level synthesis"]
Multimodal RAG: tables, figures, and ColPali
Enterprise answers hide in tables, charts, and scanned diagrams that text parsers mangle. Two viable strategies: (a) parse-and-describe—extract tables as structured units, generate text summaries of figures, embed both alongside the source crop; (b) skip parsing entirely with vision retrievers like ColPali, which embeds page screenshots as grids of patch vectors via a vision-language model and scores queries with late interaction (MaxSim over multivectors). ColPali-style retrieval is remarkably robust on visually rich PDFs and eliminates the parser as a failure mode—but multivector storage is an order of magnitude larger per page, and your engine must support multivector comparators natively (Qdrant does; see Section 9). Retrieval returns page images, so the generator must be a vision-capable model, which raises per-answer token cost.
Agentic retrieval: iterate only when single-shot fails
Single-shot retrieval fails on ambiguous, multi-hop, or under-specified questions where first-pass recall is inherently low. Agentic (iterative) retrieval—in the spirit of Self-RAG and FLARE—lets the model assess evidence, reformulate, pivot to another index or tool, and stop when coverage is sufficient. It typically multiplies latency and token cost by 2–5× and compounds error if the assessment step is weak, so bound it: a fixed step budget, deterministic stop conditions, and abstention on budget exhaustion. Chapter 5 covers the surrounding agent machinery; here the retrieval-side rule is that every iteration must run through the same authorization and evaluation contract as the first.
flowchart LR
U["User question"] --> PL["Plan retrieval step"]
PL --> RT["Retrieve via contract"]
RT --> AS["Assess evidence coverage"]
AS -->|"sufficient"| ANS["Answer with citations"]
AS -->|"gap found"| RF["Reformulate or pivot source"]
RF -->|"budget left"| PL
RF -->|"budget exhausted"| AB["Abstain or partial answer"]
Flat hybrid RAG
Default. Cheapest, fastest, easiest to evaluate. Choose unless a measured query slice proves it insufficient.
GraphRAG
Multi-hop entity questions and corpus-level synthesis. Pay LLM ingest cost and graph maintenance; evaluate local and global modes separately.
Multimodal / ColPali
Visually rich documents where parsers lose the evidence. Pay multivector storage and vision-model generation cost.
Agentic loops
Ambiguous or compositional queries with low first-pass recall. Pay 2–5× latency/cost; require budgets and abstention.
7. Freshness, semantic caching, and RAG cost engineering
Three production concerns dominate senior RAG interviews in 2026 and rarely appear in tutorials: keeping the index true to a moving corpus, not paying for the same answer twice, and knowing where the money actually goes.
Freshness is an SLO, not a batch job
Treat “time from source change to retrievable” as a first-class SLO. The reference pattern: change data capture on the source of truth → queue → re-parse and re-chunk only changed documents (content hashes decide) → upsert by deterministic ID → tombstone deletes → verify visibility. Deletions are the part teams forget: a revoked document that still answers queries is a compliance incident. Track freshness lag as a monitored metric with alerting, and record the embedding model version alongside content version so a model migration and a content update cannot be confused.
- Freshness lag — measure per-document time from source commit to index visibility; alert on the p95, not the mean.
- Idempotent upserts — deterministic point IDs plus content hashes make reprocessing safe and reconciliation possible.
- Deletes and ACL changes — propagate with higher priority than inserts; stale permissions are security bugs.
- Re-embed selectively — hash-gated incremental embedding avoids full-corpus re-embeds on every pipeline run.
Semantic caching of answers
Support and internal-helpdesk traffic is heavily repetitive; a semantic cache embeds incoming queries and serves a stored answer when a previous query is similar enough. Done naively it is a correctness and security hazard: paraphrases with different constraints (“2025” vs “2026”, negations, tenants) collide, and corpus updates silently invalidate cached answers. The guards are the design: scope cache keys by tenant, corpus version, and policy version; require a high similarity threshold plus a cheap lexical constraint check on protected literals; TTL tied to the freshness SLO; invalidate entries whose cited sources changed; and cache only high-confidence answers. Hit rates of 20–40% on repetitive support workloads are a realistic example planning number—measure your own. (This is distinct from provider-side prompt/prefix caching, covered in Chapter 2, and from reranker result caching in Section 4.)
flowchart LR
Q["Incoming query"] --> EMB["Embed query"]
EMB --> LK["Similarity lookup in answer cache"]
LK -->|"hit above threshold"| GD["Guards: tenant, corpus version, literals, TTL"]
GD -->|"pass"| CA["Serve cached answer"]
GD -->|"fail"| FULL["Full retrieve + generate"]
LK -->|"miss"| FULL
FULL --> WR["Write answer + query vector back"]
WR --> LK
Where the money goes
Run the arithmetic before optimizing. Example only: 10M chunks at 1024 dimensions in float32 is 10M × 1024 × 4 B ≈ 41 GB of raw vectors—int8 scalar quantization cuts it to ~10 GB, binary to ~1.3 GB with rescoring, which decides whether the index fits RAM or needs disk-backed storage. Embedding 10M chunks of ~300 tokens at an example $0.02/M tokens is roughly $60 one-time—ingest embedding is rarely the problem. Generation dominates steady-state: a 6k-token packed prompt at an example $3/M input is about $0.018 per answer, hundreds of times the marginal vector-query cost on a warm node. So the levers, in order of typical impact: cache answers, pack fewer/better tokens (reranking pays for itself here), route easy queries to cheaper models (Chapter 2), then quantize and tier storage.
8. Benchmark relevance and performance together
A golden set contains representative queries plus graded or binary relevance judgments over evidence units. Sample navigational, exact-identifier, semantic, multi-hop, filtered, multilingual, fresh-content, long-document, and “no answer” cases. Split tuning from final validation so fusion weights and chunk sizes are not optimized on the score you report. Version queries, judgments, corpus snapshot, parsing, embedding, index configuration, and code.
| Metric | Question answered | Blind spot |
|---|---|---|
| Precision@k | What fraction of the top k is relevant? | Does not reward finding all relevant material. |
| Recall@k | What fraction of known relevant items appears by k? | Requires reasonably complete judgments. |
| MRR | How early is the first relevant result? | Ignores additional relevant items. |
| nDCG@k | Are highly relevant items ranked early, using graded judgments? | Depends on judgment quality and cutoff. |
| Evidence coverage | Are all answer-required facts present after packing? | Needs task-specific annotation. |
| Abstention precision/recall | Does the system refuse when evidence is absent? | Thresholds depend on failure cost. |
Report macro averages and slices. A 2-point nDCG gain that hides a severe regression on one tenant, language, or exact-ID query is not a safe improvement. Inspect per-query deltas and categorize failures: parse loss, chunk boundary, stale index, ACL/filter error, candidate miss, fusion error, reranker error, packing loss, or generator misuse. New variants from this chapter—contextual retrieval, GraphRAG, agentic loops, the semantic cache—enter the harness as configurations, never as unconditioned defaults.
Measure under concurrency
Benchmark isolated stage latency and end-to-end latency. Warm and cold behavior differ. Include index build and freshness lag, throughput, error rate, CPU, memory, I/O, and cost per successful answer. Run controlled sweeps—candidate depth, ef, quantization, reranker depth—with fixed corpus and query set. Then load-test the best few configurations because tail latency can change under resource contention.
9. Qdrant as a production case study
Qdrant’s core data model is a collection of points, where a point has an ID, one or more vectors, and optional JSON payload. Named vectors allow different representations on the same point, and multivector fields with a MaxSim comparator support ColPali-style late-interaction retrieval natively—one reason it pairs well with the multimodal patterns in Section 6. Distance and dimensions are configured per vector. Payload fields support filtering; index frequent, security-relevant fields such as tenant or visibility before ingestion. The collection documentation notes that point and indexed-vector counters can be approximate during optimization, so do not use them as an exact ingestion ledger.
Collection and tenancy choices
A collection per tenant gives strong operational separation but can create excessive collection/index overhead and complicate fleet-wide updates. A shared collection with tenant payload and a mandatory filter is efficient for many smaller tenants but makes authorization enforcement and noisy-neighbor testing critical. Dedicated collections may be appropriate for very large tenants, incompatible schemas, independent scaling, or embedding migrations. Put authorization-derived filters in trusted server code, not model output or user-provided query text.
Storage, index, and availability
Choose in-memory versus on-disk vectors and HNSW, quantization, and rescoring from the measured working set—the Section 7 memory arithmetic decides which regime you are in. Qdrant’s optimization guide documents configurations for low memory, high speed, and high precision; these are starting scenarios, not automatic recommendations. Sharding increases capacity and parallelism; replication improves availability and can increase read capacity, but multiplies storage and write work. The distributed-deployment documentation notes that self-hosted shard balancing is operational work and recommends a load balancer so replicas and coordinators are not stranded behind one entry node. On Kubernetes, treat it as a stateful system: persistent volumes, anti-affinity or topology spread, resource limits, disruption budgets, snapshot/restore drills, and an upgrade/rollback runbook—stateful recovery and shard placement, not a green Pod, define readiness.
Writes, backups, and migrations
Use deterministic point IDs and content hashes so reprocessing is idempotent. Write a source version into payload. Reconcile source-of-truth counts and hashes, not approximate index counters. Exercise snapshot creation and restore; a backup that has never restored is only an assumption. For an embedding change, build a new named vector or collection, backfill from the authoritative content, dual-read a shadow sample, compare quality and latency, switch an alias or routing layer, and retain rollback until freshness and parity checks pass. Qdrant provides official guides for snapshots and zero-downtime embedding-model migration; validate version-specific mechanics before execution.
10. Managed retrieval on AWS and GCP — and when to build instead
Both clouds now sell the entire Figure 1 pipeline as a service. A senior candidate is expected to know what each managed layer actually does, where its control surface ends, and how to argue build-vs-managed without ideology. Chapters 8 and 9 cover the full stacks; here is the retrieval slice.
AWS
- Bedrock Knowledge Basesmanaged ingest → chunk → embed → retrieve, with RetrieveAndGenerate and citation support
- OpenSearch Serverlessvector engine for hybrid lexical + ANN at scale
- Amazon Kendraconnector-rich enterprise search with ACL-aware result trimming; usable as a Bedrock retriever
- Aurora / RDS + pgvectorSQL-consolidated vectors; a supported Knowledge Bases backend
- S3 Vectorslow-cost object-storage vector tier for cold or massive corpora (verify current GA status/regions)
Google Cloud
- Vertex AI Searchend-to-end managed search and grounding with connectors and ACLs
- Vertex AI Vector SearchScaNN-based ANN service (formerly Matching Engine)
- Vertex AI RAG Enginemanaged corpus, chunking, and retrieval orchestration for Gemini grounding
- AlloyDB AIPostgreSQL-compatible with pgvector plus a ScaNN index option
- BigQuery vector searchvector similarity inside the warehouse for analytical joins
Build vs managed, honestly
Fully managed (Bedrock KB, Vertex AI Search)
Wins when the corpus is standard formats, connectors and ACLs matter more than ranking control, the team is small, and time-to-value is the constraint. Accept opaque ranking and coarse chunking control; keep your own golden-set evaluation anyway.
Managed store, custom orchestration
The common senior middle path: OpenSearch/Vector Search/pgvector as the engine, your own parsing, chunking, hybrid fusion, reranking, and packing. Most of the quality levers with far less undifferentiated ops.
Self-hosted engine (Qdrant)
Wins on multivector/late-interaction features, filtered-HNSW control, cost at high sustained scale, and portability across clouds. You inherit sharding, backups, upgrades, and capacity planning (Section 9).
Database consolidation (pgvector / AlloyDB)
Wins when vectors join transactional data, scale is moderate, and one fewer system beats peak ANN performance. Revisit when the working set outgrows the instance.
The deciding questions are always the same: Is retrieval quality your product differentiator or a commodity? Can the managed chunking/ranking be overridden where your judgments show it failing? What does exit cost look like—can you re-derive the index from the authoritative store you kept? Answer those with benchmark evidence on your corpus, and the choice usually makes itself.
11. Select, migrate, and debug the whole system
Choose a search engine by workload, not category labels. A dedicated vector system is attractive for vector-native filtering, multivector search, and independent scaling. Elasticsearch/OpenSearch can consolidate mature lexical search, aggregations, and hybrid retrieval. PostgreSQL with pgvector can minimize operational surface when transactional metadata and scale fit one system. Apache Solr still anchors many Lucene-based estates; a Solr migration must translate analyzers, synonyms, boosts, faceting, and operational SLAs—relevance behavior and recovery procedures, not just stored documents. Include team expertise, recovery, tenancy, write patterns, compliance, cost, and migration reversibility in every comparison.
| Pressure | First evidence to inspect | Common wrong fix |
|---|---|---|
| Relevant document absent | Parser output, chunk IDs, source/index version, exact retrieval | Increase prompt length |
| Exact codes fail | Sparse analyzer/tokenization and hybrid candidate list | Swap dense model only |
| Filtered query returns few hits | Filter selectivity, payload index, ANN candidate depth, exact filtered result | Raise final k blindly |
| p99 spikes during ingest | CPU/I/O saturation, optimizer/index activity, segment state | Add model retries |
| Citations resolve incorrectly | Stable IDs, offsets, version mapping, context-packer transforms | Ask the generator to invent better citations |
| Fresh document not found | Pipeline checkpoint, queue lag, upsert acknowledgement, index visibility | Tune HNSW |
| Stale answer served fast | Semantic-cache scope keys, TTL, invalidation on source change | Disable caching everywhere |
A safe migration sequence
- FreezeVersion the retrieval contract, judgments, and a replayable traffic sample.
- BackfillLoad the target from the authoritative source with deterministic IDs; reconcile counts and hashes.
- ShadowCompare result overlap, relevance, filters, latency, and errors without affecting users.
- Dual-writeOr capture a change log; monitor freshness divergence between old and new.
- CanaryRoute by tenant/query class with automatic rollback gates on relevance and p95 deltas.
- Cut overVerify restore/DR and dashboards; retire the old index only after the rollback window.
Search quality monitoring in production needs proxies plus sampled judgments. Track empty/low-confidence results, reformulations, citation clicks, answer abstentions, retrieval overlap by version, semantic-cache hit/invalidation rates, freshness lag, and user feedback. Do not equate click-through with relevance: position, presentation, and user urgency confound it. Convert investigated failures into the offline dataset.
Interview playbook
For a retrieval design question, use RANKED:
- R — Requirements: users, corpus, relevance definition, freshness SLO, ACLs, scale, latency, and cost per answer.
- A — Authoritative data: source, parsing, stable IDs, versions, lineage, and deletion behavior.
- N — Nomination: lexical/dense candidate generators, contextual enrichment, filters, depths, and exact baseline.
- K — Keep order: fusion, reranking, deduplication, parent expansion, and packing.
- E — Evaluate and expose: judgments, slices, metrics, traces, load tests, and error taxonomy.
- D — Deploy safely: idempotent writes, sharding/replication, caching guards, backups, canary, rollback, and SLOs.
Then earn senior credit by escalating deliberately: name the flat-RAG baseline first, and justify each upgrade—contextual retrieval, GraphRAG, multimodal, agentic loops, semantic caching—by the query slice it fixes and the cost it adds. On cloud questions, show you know what Bedrock Knowledge Bases or Vertex AI Search actually manage, and where their control surface ends.
Common traps
- Calling cosine similarity “accuracy,” or conflating ANN recall with relevance recall.
- Tuning on a few memorable queries and reporting the same queries as validation.
- Combining raw dense and sparse scores without calibration, instead of rank fusion.
- Applying tenant filters after retrieval, allowing unauthorized candidates into context or traces.
- Assuming a larger chunk, larger k, or larger context window monotonically improves answers.
- Proposing GraphRAG or agentic loops before showing that flat hybrid retrieval fails a measured slice.
- Adding a semantic cache with no tenant/version scoping—an availability feature that becomes a correctness or security bug.
- Choosing managed vs self-hosted by ideology rather than control-surface, cost, and exit analysis.
- Describing a migration as “reindex and switch” with no change capture, parity check, canary, or rollback.
Question bank
These prompts test retrieval reasoning, not memorized product vocabulary.
Q1When will BM25 outperform dense retrieval?
Strong answer outline
- Describe lexical strength on rare identifiers, names, exact phrases, and domain tokens.
- Contrast dense paraphrase recovery and embedding domain mismatch.
- Propose query slices and a hybrid baseline instead of declaring a universal winner.
Follow-up probes
- How do analyzers affect product codes?
- How would you detect query-class drift?
Pass if examples, failure modes, and measurement are present; fail if the answer is “keywords versus semantics” only.
Q2Why use reciprocal rank fusion instead of adding dense and sparse scores?
Strong answer outline
- Explain incompatible, query-varying score scales.
- Show that RRF combines rank support without score calibration.
- Name limitations: it discards magnitude and still needs candidate-depth and weight tuning on held-out judgments.
Follow-up probes
- When would normalized score fusion be preferable?
- What does the RRF constant change?
Pass if the scale problem and validation plan are clear; fail if RRF is described as guaranteed superior.
Q3How would you choose a chunking strategy for policy PDFs?
Strong answer outline
- Preserve headings, clauses, tables, effective dates, and exception relationships.
- Index focused units with parent links and stable offsets.
- Compare fixed-window baseline and structure-aware variants on evidence coverage, citations, latency, and index size.
Follow-up probes
- What happens with a malformed PDF?
- How do you handle repeated headers?
Pass if parsing failures and a benchmark are included; fail if a universal token size is asserted.
Q4What problem do contextual retrieval and late chunking solve, and when is each worth it?
Strong answer outline
- Name the defect: chunks embedded in isolation lose referents, entities, and section-scoped conditions.
- Contrast mechanisms: contextual retrieval prepends an LLM-generated situating preamble (helps dense and BM25; one cached LLM call per chunk); late chunking pools token embeddings from a long-context pass (no LLM cost; dense only).
- Commit to measuring on your own judgments—gains are corpus-dependent and re-index cost rises.
Follow-up probes
- How does prompt caching change contextual retrieval economics?
- Which corpora would show near-zero gain?
Pass if both mechanisms and their cost asymmetry are explained; fail if vendor-reported percentages are recited as universal truths.
Q5Explain HNSW tuning without relying on defaults.
Strong answer outline
- Describe layered graph traversal and the roles of connectivity, build exploration, and search exploration.
- Keep exact results as the ANN oracle.
- Sweep parameters against recall, latency, memory, build time, filters, and concurrency.
Follow-up probes
- Why might higher
efnot repair relevance? - What changes under strict filters?
Pass if index recall is separated from application relevance; fail if “higher equals better” is the whole answer.
Q6Why can metadata filtering reduce vector-search recall?
Strong answer outline
- Contrast pre-, in-, and post-filter behavior and candidate depletion.
- Discuss selectivity, payload indexes/filter-aware traversal, and exact filtered baselines.
- Test by filter slice and raise search effort only with measured bounds.
Follow-up probes
- How would tenant filtering differ from a preference filter?
- When is exact search cheaper?
Pass if security filters are mandatory and implementation-specific behavior is acknowledged; fail if filters are treated as a UI detail.
Q7Design a golden set for enterprise search and defend your metric choices.
Strong answer outline
- Sample real intent and important query classes, including no-answer and ACL cases; define the relevance unit and a graded rubric.
- Map metrics to behavior: MRR for navigational, recall/nDCG for multi-evidence, evidence coverage after packing, abstention quality.
- Version corpus/judgments, separate tuning from validation, and report slices alongside macro averages.
Follow-up probes
- How do you handle incomplete relevance judgments?
- How do production failures enter the set?
Pass if provenance, slices, and leakage prevention are concrete and metric choice follows user behavior; fail if the dataset is just generated questions plus recited formulas.
Q8The correct passage is retrieved but the answer is wrong. What next?
Strong answer outline
- Verify it survived deduplication and packing with sufficient surrounding conditions.
- Inspect answer trace, instruction hierarchy, citation mapping, and conflicting evidence.
- Run a controlled answer test with fixed context before changing retrieval.
Follow-up probes
- How do you test faithfulness?
- When should the system abstain?
Pass if component isolation precedes tuning; fail if the embedding model is changed immediately.
Q9When does GraphRAG beat flat retrieval, and what does it cost?
Strong answer outline
- Identify the failing query families: multi-hop entity questions and corpus-level synthesis where no single passage holds the answer.
- Describe the pipeline—LLM entity/relation extraction, community detection, pre-summarization—and the local vs global query modes.
- Weigh costs: LLM ingest over the whole corpus, graph maintenance on updates, harder evaluation; keep flat hybrid as default.
Follow-up probes
- How do incremental document updates propagate into the graph?
- How would you evaluate a global-synthesis answer?
Pass if the answer names the query slice that justifies the graph and its maintenance cost; fail if GraphRAG is pitched as a general upgrade.
Q10How would you build RAG over scanned, table-heavy PDFs?
Strong answer outline
- Contrast parse-and-describe (structured table units plus figure summaries with source crops) against vision retrieval (ColPali-style page-screenshot multivectors with late interaction).
- Name the costs: parser fragility on one side; multivector storage blow-up, engine support, and vision-model generation cost on the other.
- Propose a benchmark on judged visual queries before committing, and a hybrid where parsed text handles clean documents.
Follow-up probes
- What does late interaction (MaxSim) buy over a single page vector?
- How do citations work when evidence is an image region?
Pass if both strategies and their storage/generation cost asymmetry are explicit; fail if “use a multimodal model” is the whole answer.
Q11When should retrieval be agentic/iterative rather than single-shot?
Strong answer outline
- Identify low first-pass-recall slices: ambiguous, compositional, or multi-source questions.
- Design the loop with bounded steps, evidence-coverage assessment, deterministic stop conditions, and abstention on budget exhaustion.
- Quantify the 2–5× latency/cost multiplier and require every iteration to pass the same authorization contract.
Follow-up probes
- How do you keep the assessment step from compounding errors?
- What telemetry proves the loop earns its cost?
Pass if iteration is justified by a measured slice with explicit budgets; fail if agentic retrieval is presented as strictly better.
Q12Is semantic caching of answers safe? Design the guards.
Strong answer outline
- Name the hazards: paraphrase collisions on differing constraints, staleness after corpus updates, cross-tenant leakage.
- Design guards: scope keys by tenant/corpus-version/policy-version, high similarity threshold plus lexical checks on protected literals, TTL tied to freshness SLO, invalidation when cited sources change.
- Measure hit rate, false-hit rate on a judged paraphrase set, and cost saved per answer.
Follow-up probes
- How does this differ from provider prompt caching?
- What is your invalidation path when one document is revoked?
Pass if correctness and security guards precede the cost win; fail if similarity threshold is the only control mentioned.
Q13How would you migrate to a new embedding model without downtime?
Strong answer outline
- Version representation contracts and build a new vector/collection from authoritative data.
- Capture changes, shadow read, compare relevance/latency, and canary.
- Switch routing with rollback, verify freshness, then retire after a defined window.
Follow-up probes
- How do dimensions and distance change the plan?
- What if rankings improve but citations break?
Pass if parity, incremental writes, and rollback are explicit; fail if only the backfill is described.
Q14Bedrock Knowledge Bases / Vertex AI Search versus building your own pipeline—how do you decide?
Strong answer outline
- Frame the axis: connectors, ACLs, and time-to-value versus control over chunking, fusion, reranking, and debuggability.
- State what each manages and where the control surface ends—opaque ranking, coarse chunking options, limited retrieval introspection.
- Decide on evidence: run the same golden set through the managed path and a custom path; include exit cost via the authoritative content store.
Follow-up probes
- Which failure classes can you not debug in the managed path?
- What would trigger migrating off the managed service?
Pass if the answer is conditional, benchmark-driven, and names concrete control-surface limits; fail if it is vendor cheerleading or reflexive build-it-yourself.
Q15How should multi-tenant vector data be modeled?
Strong answer outline
- Compare shared collection with mandatory tenant payload against dedicated collections and hybrid tiers.
- Keep authorization outside model control and index security filters.
- Test noisy neighbors, backup/restore, deletion, migration, and cross-tenant adversarial cases.
Follow-up probes
- What if one tenant is 1,000 times larger?
- How do you prove isolation?
Pass if security and operational cardinality drive the choice; fail if tenant ID is merely added to metadata.
Q16Your RAG bill doubled. Walk through the cost model and your first three levers.
Strong answer outline
- Decompose cost per answer: generation tokens (usually dominant), reranker calls, vector query compute/memory, embedding amortization, ingest LLM enrichment.
- Instrument before acting: packed tokens per answer, cache hit rate, reranker depth, index residency, query mix drift.
- Apply levers in impact order: guarded answer caching, tighter packing/reranking, model routing, then quantization and storage tiering.
Follow-up probes
- When does quantization move the needle and when is it noise?
- How do you keep cost cuts from silently regressing quality?
Pass if the answer starts with measured decomposition and ties each lever to a quality gate; fail if it jumps straight to a cheaper model.
Proof artifact: a reproducible hybrid-search benchmark
Build a small, public-data search service that makes retrieval quality, filtered performance, and operational behavior inspectable. All results are portfolio measurements, not claims about Purnendu’s production experience.
Steps
- Select a legally usable corpus with meaningful structure. Freeze a corpus manifest containing source URI, content hash, version, and parser result.
- Create 150–300 queries with binary or graded evidence judgments. Include exact IDs, paraphrases, strict metadata filters, multiple required passages, stale versions, and no-answer cases. Split tuning and validation.
- Implement five fixed variants: sparse baseline; dense exact/ANN; hybrid RRF; hybrid plus reranking; hybrid plus contextual retrieval. Keep parsing and corpus constant.
- Sweep chunking, candidate depth, HNSW search effort, filter selectivity, and optional quantization. Record every configuration.
- Add a guarded semantic answer cache and replay a repetitive traffic sample; report hit rate, false-hit rate on judged paraphrases, and cost saved per answer.
- Run single-request and concurrent load profiles on declared hardware. Capture stage spans, p50/p95/p99, throughput, errors, CPU, memory, index size, and freshness lag.
- Run the same golden set through one managed path (Bedrock Knowledge Bases or Vertex AI Search) and write a decision memo comparing it and Qdrant against an adjacent option such as pgvector, including a reversal condition.
Metrics
Report precision@5, recall@20, MRR, nDCG@10, evidence coverage, no-answer behavior, ANN recall against exact search, filtered-query slices, latency percentiles, index/build time, memory, cache hit/false-hit rates, and cost per successful answer. Show per-query deltas, confidence intervals or paired resampling when practical, and the error taxonomy—not only averages.
Deliberate failures
- Remove the payload index for a frequent strict filter and observe quality/latency under load.
- Lower ANN search effort until exact-neighbor recall visibly fails, then distinguish ANN loss from embedding relevance.
- Corrupt a parser boundary so an exception is separated from a policy statement; confirm the golden set catches it.
- Pause incremental indexing while source versions advance; ensure freshness monitoring and a user-visible policy respond.
- Update a document cited by cached answers without invalidating the semantic cache; show the stale-answer detection catching it.
- Make the reranker unavailable; verify timeout, fallback ranking, trace status, and bounded latency.
What to present
Present one pipeline diagram, the dataset card, a Pareto chart of nDCG versus p95 latency, the cost-per-answer decomposition, two failure traces, and the one-page engine decision. Demonstrate one query where sparse wins, one where dense wins, one where contextual retrieval rescues an isolated chunk, one where a filter breaks naïve ANN, and one where the system correctly abstains.
Chapter review
Production retrieval is a chain of contracts. Preserve authoritative content and provenance, generate complementary candidates, fuse and rerank deliberately, pack evidence within a budget, and measure component as well as end-to-end behavior. Escalate beyond flat retrieval—contextual enrichment, graphs, vision retrievers, agentic loops, semantic caches—only when a measured query slice justifies the added cost, and choose between managed platforms and self-hosted engines on control-surface, cost, and exit evidence rather than ideology.
Glossary
- ANN recall
- Fraction of exact nearest neighbors recovered by an approximate index at a cutoff; distinct from judged relevance recall.
- BM25
- Lexical ranking family using term frequency, inverse document frequency, and document-length normalization.
- Contextual retrieval
- Prepending an LLM-generated document-situating preamble to each chunk before embedding and lexical indexing.
- Late chunking
- Embedding a whole document with a long-context model, then pooling token vectors per chunk boundary afterward.
- GraphRAG
- Retrieval over an LLM-extracted knowledge graph with community summaries, enabling local entity and global synthesis queries.
- Late interaction
- Scoring queries against multivector representations (for example MaxSim over ColPali patch vectors) instead of one pooled vector.
- HNSW
- Hierarchical proximity-graph ANN structure with build, memory, latency, and recall trade-offs.
- RRF
- Reciprocal rank fusion, which combines result lists using document positions rather than raw score scales.
- Reranker
- Later-stage scorer that evaluates a query and candidate more jointly than a first-stage retriever.
- Semantic cache
- Answer cache keyed by query-embedding similarity, guarded by tenant, corpus version, literals, and TTL.
- Freshness lag
- Time from a source-of-truth change to that change being retrievable; an SLO, not an accident.
- Evidence coverage
- Whether packed context contains all facts required to answer a task, not merely one relevant chunk.
Mastery checklist
- I can isolate ingestion, candidate, rank, pack, and generation failures.
- I can give a query where sparse wins and one where dense wins.
- I can calculate a simple RRF result and explain when it is insufficient.
- I can explain contextual retrieval and late chunking, including their cost asymmetry and when each pays.
- I can separate embedding relevance, ANN recall, and end-to-end answer quality.
- I can explain how HNSW parameters and strict filters change the workload.
- I can name the query slices that justify GraphRAG, ColPali-style retrieval, and agentic loops—and their costs.
- I can design a guarded semantic cache and a freshness pipeline with deletion handling.
- I can decompose RAG cost per answer and order the levers by impact.
- I can map a retrieval design onto Bedrock/OpenSearch/Kendra and Vertex AI Search/Vector Search/AlloyDB and defend build-vs-managed.
- I can outline a shadowed, canaried, reversible search migration.
Primary sources
Links checked . Vendor behavior is version-sensitive; verify the documentation for the deployed release.
- Qdrant — collections, points, vectors, distance, and configuration
- Qdrant — hybrid and multi-stage queries, RRF, and fusion
- Qdrant — payload indexes, filterable HNSW, and index parameters
- Qdrant — scalar, product, and binary quantization trade-offs
- Qdrant — distributed deployment, shard movement, and load balancing
- Qdrant — zero-downtime embedding-model migration
- Anthropic — Introducing Contextual Retrieval
- Günther et al. — Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
- Edge et al. — From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Faysse et al. — ColPali: Efficient Document Retrieval with Vision Language Models
- Asai et al. — Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- Malkov and Yashunin — original HNSW paper
- Guo et al. — ScaNN: Accelerating Large-Scale Inference with Anisotropic Vector Quantization
- pgvector — official project documentation for exact, HNSW, IVFFlat, and filtered search
- Elastic — official hybrid-search overview
- AWS — Amazon Bedrock Knowledge Bases user guide
- AWS — OpenSearch Serverless vector search collections
- AWS — Amazon Kendra enterprise search
- Google Cloud — Vertex AI Search
- Google Cloud — Vertex AI Vector Search overview
- Google Cloud — Vertex AI RAG Engine overview
- Google Cloud — AlloyDB AI
CHAPTER 05 · PRIORITY 0
Agentic Systems & LLM Application Engineering
34 min read · 16 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- Decide when an agent is warranted and when deterministic code, retrieval, or a form is safer.
- Choose among router, planner–executor, supervisor/worker, pipeline, debate, and swarm topologies — and argue when a single agent with good tools beats all of them.
- Design narrow tool contracts with schema validation, authorization, idempotency, and interpretable failures, and explain what MCP and A2A standardize versus what stays your responsibility.
- Make multi-step executions resumable with checkpointing (LangGraph) or deterministic replay (Temporal), and reason about idempotent side effects under replay.
- Separate working context, workflow state, and long-term memory; design episodic/semantic memory with compaction, provenance, and deletion.
- Design human-in-the-loop approval bound to exact expiring actions, plus sandboxing for code-executing and computer-use agents.
- Compare AWS Bedrock Agents/AgentCore and GCP Vertex AI Agent Builder/Agent Engine/ADK against a self-built LangGraph stack, with concrete trade-offs.
1. Earn the right to be agentic
An agent is a system in which a model chooses at least part of the action sequence at runtime. That flexibility is useful when the task is open-ended, the correct path depends on observations, and tool selection cannot be exhaustively encoded. It also expands the state space: more trajectories, model calls, permissions, latency, cost, and failure combinations. "Agentic" is therefore a design choice, not a maturity level — a point Anthropic's Building effective agents makes explicitly: use the simplest composable pattern that solves the task.
Deterministic function
Use when inputs and rules are known: calculations, schema transformations, authorization, validation, and irreversible side effects. It is cheap, testable, and explainable.
Fixed workflow
Use when steps are known but some steps need model judgment: classify → retrieve → draft → validate. Explicit control flow makes recovery and evaluation tractable.
Bounded agent
Use when the next information-gathering action depends on prior results. Constrain tools, steps, budget, scopes, and terminal outcomes.
Human decision
Use when policy, accountability, or irreversible impact requires judgment that the system is not authorized to make.
Apply the uncertainty–consequence test. Agent value rises with path uncertainty: research, diagnosis, codebase exploration, or heterogeneous support requests. Required control rises with consequence: moving money, deleting data, changing production, contacting a customer, or disclosing sensitive information. High uncertainty plus high consequence calls for a bounded agent that prepares evidence and a plan, then an explicit approval before action.
Define success and terminal outcomes
"Helpful response" is not an operational contract. Define allowed terminal states such as completed, needs_user_input, awaiting_approval, blocked_by_policy, budget_exhausted, and dependency_failed. For each, specify what the user sees and whether resumption is possible. A loop without a terminal-state model is a reliability bug waiting for traffic.
2. Choose an orchestration pattern from the control problem
Framework names matter less than who decides the next step and where state lives. Model the workflow as states, events, guarded transitions, side effects, and terminal states. Then choose a framework — or plain code — that expresses this model clearly.
| Pattern | Use when | Primary risk | Control |
|---|---|---|---|
| Router | One request maps to one specialist path | Misrouting or category drift | Confidence threshold, fallback, labeled confusion matrix |
| State machine | Allowed transitions and recovery must be explicit | State explosion | Small typed state, invariants, terminal states |
| Planner–executor | Task path depends on intermediate evidence | Stale or impossible plans | Plan validation, step cap, replan trigger |
| Supervisor/worker | Several specialist capabilities must be coordinated | Extra calls and opaque delegation | Narrow roles, shared outcome schema, central budget |
| Pipeline | Stages have distinct contracts and can be validated between steps | Error propagation without repair | Inter-stage schemas, gate checks, bounded repair loops |
| Parallel fan-out / swarm | Independent evidence can be gathered concurrently | Duplicate work and merge conflict | Branch budget, dedupe, deterministic aggregation |
| Multi-agent debate | Distinct perspectives have measurable value | Expensive agreement theater | Independent evidence, calibrated judge, stop rule |
Multi-agent topologies — and when not to use them
The supervisor/worker topology is the workhorse: a coordinator decomposes the task, delegates to workers with narrow tool grants, and merges typed results under a central budget. It pays off when subtasks genuinely parallelize (breadth-first research, multi-source evidence gathering) or when permission separation matters — a read-only research worker cannot mutate anything even if compromised by injected content. The cost is real: multi-agent systems multiply token spend because each worker re-establishes context, and coordination failures (duplicate work, contradictory partial results, lost context at handoff) become your dominant defect class.
flowchart TD
U["User request"] --> S["Supervisor: decompose, delegate, merge"]
S -->|"subtask + budget slice"| W1["Research worker (read-only tools)"]
S -->|"subtask + budget slice"| W2["Data worker (scoped SQL)"]
S -->|"subtask + budget slice"| W3["Writer worker (no tools)"]
W1 -->|"typed result"| S
W2 -->|"typed result"| S
W3 -->|"typed result"| S
S --> V["Deterministic validation + merge"]
V --> R["Final answer with source attribution"]
Debate topologies — agents critiquing each other before a judge decides — show measurable factuality gains in research settings (Du et al., 2023), but only when critics have independent evidence or genuinely different capabilities. Swarms of homogeneous agents sharing a scratchpad are the least controllable topology: emergent coordination is emergent failure. Default heuristics: a single agent with well-designed tools beats a multi-agent system for most sequential tasks; add agents only for parallelism, specialization with different tool grants, or context isolation (keeping a 200k-token research dump out of the main thread). If the subtasks never run concurrently and share all permissions, you likely want functions, not agents.
Worked example: bounded refund investigation
Consider a support workflow that may inspect an order and draft a refund, but cannot issue it without policy checks and approval above a threshold. A useful graph:
flowchart TD
A["Intake"] --> B["Validate input"]
B -->|"missing input"| N["Terminal: needs_user_input"]
B --> C["Lookup order (read tool)"]
C -->|"not found"| N
C --> D["Model proposes refund + rationale"]
D --> P["Deterministic policy check"]
P -->|"deny"| F["Terminal: blocked_by_policy"]
P -->|"low impact"| E["Execute with idempotency key"]
P -->|"high impact"| H["Human approval"]
H -->|"approved"| E
H -->|"rejected"| F
E --> T["Terminal: completed"]
ALLOWED = {
"validated": {"looked_up", "needs_user_input"},
"looked_up": {"proposed", "not_found"},
"proposed": {"policy_denied", "awaiting_approval", "approved"},
"awaiting_approval": {"approved", "rejected"},
"approved": {"completed", "dependency_failed"},
}
def transition(state, next_status):
if next_status not in ALLOWED.get(state["status"], set()):
raise ValueError("invalid workflow transition")
return {**state, "status": next_status}
The model may propose an action and rationale, but deterministic code verifies policy and authorization. The approval stores the exact proposed action, resource ID, amount, policy version, and expiry. Execution uses an idempotency key. The system never interprets approval as permission for a later, altered action.
3. Treat tool calls as untrusted requests
Function calling is a protocol round trip: the application describes tools; the model emits a structured request; application code validates and authorizes it; the application executes the operation; and a structured result returns to the model. OpenAI, Anthropic, and Google all document this separation — the model proposes arguments while client-side code performs the function. Never let the model's selection bypass normal service controls. Providers differ in schema subsets, tool-call message formats, and stop reasons; the durable design is a narrow internal tool interface plus versioned provider adapters, tested against malformed arguments, unknown tools, duplicate calls, timeouts, and partial streaming failures.
A strong tool contract
- Narrow intent —
get_order_statusis safer and easier to select thanrun_api_request. - Constrained schema — enums, bounded lengths/ranges, required fields; reject unknown fields where supported.
- Trusted identity — derive tenant, actor, and scopes from authenticated context; never accept them as model-controlled arguments.
- Clear effects — distinguish read, reversible write, irreversible write, and external communication.
- Idempotency — mutation tools accept or derive a key tied to workflow and semantic operation.
- Typed result — separate
ok, retryable error, permanent error, policy denial, and not-found; keep user-safe and operator detail distinct. - Limits — server-enforced timeout, pagination, result-size cap, rate/quota, and redaction.
{
"name": "prepare_refund",
"description": "Create a reviewable refund proposal; does not issue funds.",
"input_schema": {
"type": "object",
"properties": {
"order_id": {"type": "string", "pattern": "^ord_[A-Za-z0-9]+$"},
"reason": {"type": "string", "maxLength": 500},
"amount_minor": {"type": "integer", "minimum": 1}
},
"required": ["order_id", "reason", "amount_minor"],
"additionalProperties": false
}
}
Structured-output support can guarantee or improve schema conformance depending on provider and mode, but schema validity is not semantic validity. A perfectly formed order ID can belong to another tenant; a valid amount can exceed the refundable balance. Revalidate business invariants at execution time. Keep tool descriptions accurate and treat third-party descriptions as untrusted metadata.
Return errors the orchestrator can act on
A timeout, invalid input, expired credential, policy denial, and missing record need different control flow. Avoid returning a prose blob that the model must reinterpret. Use stable error codes, retryability, safe user message, and correlation ID. Do not expose tokens, stack traces, raw database errors, or private tool results to the model unnecessarily.
4. MCP and A2A: standard wiring, not delegated security
Before the Model Context Protocol, every agent framework re-implemented tool integration against every backend — an M×N adapter problem. MCP collapses it to M+N: hosts (the agent application) run one MCP client per connection to an MCP server, speaking JSON-RPC over stdio locally or streamable HTTP remotely. A server exposes three primitives with different control ownership, and the client/server negotiate capabilities at initialization. The practical consequence for platform teams: tool catalogs become independently deployable, versioned, governable artifacts — an internal API can be wrapped once as an MCP server and consumed by every agent runtime in the company, regardless of framework or model vendor.
| Primitive | Controlled by | What it is | Design duty |
|---|---|---|---|
| Tools | Model-invoked | Executable actions with JSON schemas | Validation, authorization, idempotency, audit — same as any tool |
| Resources | Application-controlled | Readable context (files, records, docs) addressed by URI | Tenancy filtering, redaction, freshness |
| Prompts | User-invoked | Parameterized prompt templates the client surfaces | Versioning, injection review |
flowchart LR
subgraph HA["Host application (agent runtime)"]
M["Model loop"]
C1["MCP client A"]
C2["MCP client B"]
end
M --> C1
M --> C2
C1 -->|"JSON-RPC over stdio"| S1["MCP server: internal order API"]
C2 -->|"JSON-RPC over streamable HTTP"| S2["MCP server: SaaS connector"]
S1 -->|"tools, resources, prompts"| C1
S2 -->|"tools, resources, prompts"| C2
S1 --> D1["Order service (server-held credentials)"]
S2 --> D2["Third-party API (scoped OAuth)"]
MCP is interoperability, not an authorization shortcut. The specification treats tool execution and arbitrary data access as high-risk and emphasizes consent and control; its authorization spec requires resource-bound OAuth tokens for HTTP transports and forbids token passthrough. An MCP server still needs authentication, per-tool scopes, tenant isolation, input validation, output redaction, rate limits, audit logs, and downstream credentials distinct from inbound credentials. New attack surface comes with the standard: malicious or mutated tool descriptions (a server can change its advertised tools after approval), confused-deputy patterns where a broadly-scoped server acts for a narrowly-authorized user, and supply-chain risk in community servers. Pin server versions, review descriptions like code, and put an MCP gateway with policy enforcement between agents and third-party servers.
A2A: agent-to-agent across trust boundaries
The Agent2Agent protocol (initiated by Google, now under the Linux Foundation) standardizes the layer above MCP: opaque agents delegating to each other across team or vendor boundaries. An agent publishes an Agent Card — capability metadata for discovery — and peers exchange long-running tasks with lifecycle states, messages, and artifacts, with streaming and push-notification support for work that takes hours. The mental model that lands in interviews: MCP connects an agent to its tools; A2A connects an agent to other agents it does not trust with its internal state, memory, or credentials. Reach for A2A when delegation crosses an organizational boundary where sharing tool credentials is impossible; inside one team and process, in-process multi-agent orchestration is simpler, cheaper, and easier to trace.
5. Design for interruption and replay
Multi-step systems fail between steps. A process can crash after an external write but before recording success; a user can approve hours later; a provider can time out after completing a request. Durable execution requires explicit persisted state and replay-safe side effects — not merely "retry three times." The industry has converged on two architectures for this.
Checkpointing versus deterministic replay
Checkpointing (LangGraph's model): the orchestrator persists a snapshot of graph state at every superstep to a checkpointer backed by Postgres or similar. LangGraph's persistence docs describe thread-scoped checkpoints, pending writes, fault recovery, and time-travel debugging; its interrupt mechanism pauses a thread for human input and resumes from saved state — with the explicit caveat that side effects before an interrupt must be idempotent because the node re-runs on resume. Deterministic replay (Temporal's model, shared by Azure Durable Functions): workflow code must be deterministic; every side effect (model call, tool call, timer) runs in an activity whose result is recorded in an event history. After a crash, the engine re-executes the workflow function and feeds back recorded results, reconstructing exact state without re-running effects. Retries, timeouts, heartbeats, and multi-day waits are engine primitives.
| Dimension | LangGraph checkpointing | Temporal replay |
|---|---|---|
| Unit of persistence | State snapshot per superstep | Append-only event history |
| Nondeterminism | Tolerated in nodes; you own idempotency | Forbidden in workflow code; isolated in activities |
| Built for | LLM-native graphs, streaming, human interrupts | Mission-critical, months-long business workflows |
| Ops burden | You run the checkpoint store and workers | Cluster or Temporal Cloud; heavier but battle-tested |
| Sweet spot | Agent loops with minutes-to-days human gates | Agent steps embedded in transactional business processes |
sequenceDiagram
participant O as "Orchestrator"
participant CP as "Checkpoint store"
participant T as "Refund tool"
O->>CP: persist state before execute step
O->>T: execute refund with idempotency key K1
T-->>O: committed
Note over O: crash before success is recorded
O->>CP: reload thread on restart
CP-->>O: last checkpoint shows execute pending
O->>T: replay call with same key K1
T-->>O: already applied, same result returned
O->>CP: record success and advance
Retry ownership
Place retries at one layer whenever possible. SDK, HTTP client, orchestrator, queue, and tool service each retrying can multiply calls. Retry only transient failures, with exponential backoff, jitter, a deadline, and a total attempt budget. Respect provider retry hints. Do not retry policy denials or deterministic validation failures. If a request has ambiguous outcome, reconcile using an operation key or read-before-retry.
| Failure | Response | Why |
|---|---|---|
| Rate limit with retry hint | Bounded delayed retry or alternate capacity | Likely transient; avoid synchronized retry storm. |
| Malformed model arguments | Return validation detail; one bounded repair attempt | Repeated sampling can loop without new information. |
| Policy denial | Terminal denial or human policy path | Technical retries must not override governance. |
| Mutation timed out | Query by idempotency key before retry | The remote side may have committed. |
| Dependency outage | Checkpoint, degrade or pause, show recoverable status | Preserves user trust and avoids runaway cost. |
Termination is a product feature
Enforce maximum wall time, model calls, tool calls, repeated identical calls, replans, tokens, and monetary budget. Detect no-progress cycles using normalized action/result fingerprints. On exhaustion, preserve a partial result and missing requirements where safe. A hard "something went wrong" after ten hidden retries wastes both evidence and trust.
6. Separate context, state, and memory
Conversation history is not a database, and a vector store is not automatically memory. Use three distinct concepts:
- Working context: bounded messages, evidence, tool results, and instructions needed for the current model call.
- Workflow state: authoritative typed fields required to resume and enforce transitions.
- Long-term memory: intentionally retained facts or summaries available across sessions, with provenance, consent, correction, and deletion.
Episodic and semantic memory, and the compaction pipeline
Long-term memory splits along the same lines as human memory research. Episodic memory records what happened: specific interactions, decisions, and outcomes, time-stamped and immutable ("user rejected the summary format on 2026-07-12"). Semantic memory stores distilled facts and preferences ("user prefers bullet summaries; account tier is enterprise"), each carrying provenance back to its episodic sources and a confidence level. A background consolidation job — run at session end or on a schedule, not inline on the hot path — extracts candidate facts from episodes, deduplicates against existing memory, resolves conflicts by recency and evidence, and expires stale entries. This is the same design implemented by managed offerings: Bedrock AgentCore Memory's short-term/long-term strategies and Vertex AI Agent Engine's Memory Bank both separate raw session events from extracted durable facts.
Keep raw authoritative values in workflow state; format them into prompts at call time. Summaries are lossy and should carry source references and versions. Treat retrieved memory as untrusted context: it can be stale, incorrectly attributed, or deliberately poisoned by earlier injected content. Enforce tenant/user boundaries before retrieval, never after. Do not silently promote a model inference into a user fact, and do not store secrets or sensitive tool results merely because they might help a later response.
Manage the context window as a budget
Reserve space for system/tool schemas, current request, evidence, and output. Drop irrelevant history; compact older turns into summaries that explicitly preserve unresolved commitments, pending constraints, and open questions — the three things naive summarization loses first. Retrieve only task-relevant memory and cap tool outputs. Compaction failures are silent, so evaluate long-running conversations and resume cases specifically. Cache stable prefixes only when provider semantics, privacy, and version invalidation are understood (Chapter 2 covers the serving mechanics; Chapter 3 covers context engineering in depth).
7. Put authority outside the model
Prompt injection is a control-flow attack: untrusted content attempts to redefine instructions or induce tool use. Label and delimit external content, but do not depend on prompting alone. Authorization must be enforced by trusted code at the tool boundary. Give each workflow an explicit capability set and derive scopes from the authenticated actor, tenant, environment, and approved purpose.
| Action class | Default control | Example |
|---|---|---|
| Read, low sensitivity | Scoped authorization + logging | Read public product documentation |
| Read, sensitive | Least privilege + purpose/tenant check + redaction | Retrieve a customer record |
| Reversible write | Preview + idempotency + bounded auto-execution policy | Create a draft ticket |
| Irreversible/high impact | Exact human approval + separation + audit | Issue funds or delete production data |
| External communication | Recipient/content preview + approval or explicit policy | Send email to a customer |
Human-in-the-loop approval as a state machine
Request approval for the exact action, not a vague plan. Show target, effect, sensitive fields, cost/amount, environment, and why the action is requested. Bind the approval to a content hash of the arguments, a policy version, the approving actor, and an expiry; invalidate it whenever material arguments change. Separate proposer from executor where risk warrants. Design the reviewer experience deliberately: batch low-risk approvals to avoid alert fatigue, surface diffs rather than raw payloads, and make rejection a first-class state that carries feedback back into the workflow rather than a dead end.
stateDiagram-v2
[*] --> Proposed
Proposed --> AutoExecute: within auto policy
Proposed --> AwaitingApproval: high impact action
AwaitingApproval --> Approved: reviewer approves exact args
AwaitingApproval --> Rejected: reviewer rejects with reason
AwaitingApproval --> Expired: approval TTL elapsed
Approved --> Invalidated: arguments changed
Invalidated --> Proposed
Approved --> Executed: idempotent execution
AutoExecute --> Executed
Executed --> [*]
Rejected --> [*]
Expired --> [*]
Sandboxing code-executing agents
Any agent that runs generated code, shells, or browsers needs an execution boundary stronger than a prompt. The standard stack, from inside out:
- Isolation — run generated code in a microVM or gVisor-class sandbox per session, never in the orchestrator process; managed runtimes (AgentCore Runtime, Agent Engine code execution) provide session-isolated sandboxes for exactly this reason.
- Egress control — default-deny outbound network; allowlist specific domains. Data exfiltration via an innocent-looking HTTP call is the primary injection payoff.
- No ambient credentials — the sandbox holds no long-lived secrets; tools broker scoped, short-lived tokens server-side.
- Resource caps — CPU, memory, disk, wall time, and process count limits; a runaway loop is a denial-of-wallet.
- Audit — record commands, file mutations, and network attempts; sample into security review.
8. Computer-use and browser agents: the integration of last resort
Computer-use agents perceive a screen (screenshots, sometimes accessibility trees or DOM) and act through clicks, keystrokes, and scrolls; browser agents are the web-scoped variant, often driving Playwright against the DOM instead of pixels. They matter because the long tail of enterprise work lives in UIs without APIs — legacy ERP screens, partner portals, internal admin consoles. They are also the least reliable agent class in production: every step is a vision-model round trip (seconds of latency, real token cost), and errors compound across the 20–50 steps a nontrivial task takes. On OS-level benchmarks like OSWorld and realistic web benchmarks like WebArena, even frontier agents remain far below human success rates — improving fast, but not a reliability profile you build unattended irreversible actions on.
flowchart LR
G["Goal + constraints"] --> M["Model plans next UI action"]
M --> V["Action validator (domain allowlist, action policy)"]
V -->|"allowed"| B["Sandboxed browser or VM"]
V -->|"blocked"| HIL["Escalate to human"]
B -->|"sensitive step (login, payment)"| HIL
B --> SS["Screenshot + DOM observation"]
SS --> M
HIL -->|"human completes or approves"| B
Production guardrails follow directly from the threat model. The rendered page is untrusted input, so a webpage can inject instructions into the agent — treat every observation as adversarial. Run the browser in a disposable, sandboxed profile with a domain allowlist and default-deny egress. Never give the model raw credentials: inject secrets at a trusted proxy or have a human complete login steps, so screenshots and traces never contain passwords. Gate payments, sends, deletes, and permission changes on human confirmation. Record the full action/screenshot trace for audit and replay. And check the decision order: if an API or MCP server exists for the target system, use it — UI automation costs roughly an order of magnitude more per task in latency and tokens and breaks on every front-end redesign. Anthropic's tool-use documentation ships computer use with equivalent cautions.
9. Managed agent platforms: AWS and GCP versus building it yourself
Everything in sections 4–8 — durable sessions, tool gateways, memory infrastructure, identity propagation, sandboxes, observability — is undifferentiated heavy lifting that both clouds now sell. On AWS, Bedrock Agents is the opinionated managed orchestrator (instructions, action groups from OpenAPI/Lambda, knowledge bases, return-of-control for client-side execution), while Bedrock AgentCore is the framework-agnostic platform layer: Runtime (serverless sessions with microVM isolation, long-running executions), Gateway (turns existing APIs and Lambda functions into MCP tools), Memory (short/long-term with extraction strategies), Identity (OAuth token vault so agents act on behalf of users), plus managed Code Interpreter and Browser tools — running any framework and any model. On GCP, Vertex AI Agent Builder is the umbrella: the open-source Agent Development Kit (ADK) for code-first multi-agent development with built-in evaluation, and Agent Engine as the managed runtime with sessions, Memory Bank, sandboxed code execution, and native A2A support — deployable from ADK, LangGraph, or LangChain.
AWS
- Bedrock Agentsmanaged orchestrator: action groups, KBs, return-of-control
- AgentCore Runtimeserverless sessions, microVM isolation, long executions
- AgentCore Gatewayexisting APIs/Lambda exposed as MCP tools
- AgentCore Memoryshort-term events + long-term extraction strategies
- AgentCore IdentityOAuth token vault, delegated user authority
- Step Functionsdeterministic workflow backbone around agent steps
Google Cloud
- Vertex AI Agent Builderumbrella console + governance for agents
- Agent Development Kitopen-source code-first framework, multi-agent, evals
- Vertex AI Agent Enginemanaged runtime: sessions, Memory Bank, sandboxes
- A2A supportagent-to-agent interop, Agent Cards
- Workflows / Cloud Rundeterministic backbone or self-hosted runtime
Fully managed orchestrator
Bedrock Agents or Agent Builder console agents. Fastest to demo; least control over the loop, prompt assembly, and failure semantics. Fits standard tool-plus-RAG assistants owned by small teams.
Your framework on managed runtime
LangGraph or ADK deployed to AgentCore Runtime / Agent Engine. You own graph logic and contracts; the cloud owns session isolation, scaling, identity, memory. The current default for serious teams.
Fully self-built
LangGraph plus your own Postgres checkpointer, sandboxes, and gateway on ECS/Cloud Run. Maximum control and portability; you staff the undifferentiated infrastructure and its security reviews.
The trade-off conversation interviewers want: managed platforms buy you session isolation, identity, and memory infrastructure you would otherwise build badly under deadline, at the price of a thicker lock-in surface (memory schemas, identity flows, and gateway configs are harder to port than model APIs) and less visibility when the loop misbehaves. Self-built LangGraph maximizes control and portability but makes you the security and reliability owner for sandboxes, token handling, and checkpoint storage. The middle path — your graph, their runtime — is winning because it splits the lock-in: your orchestration logic stays portable code while the cloud absorbs the parts auditors ask about.
10. Operate the agent as a distributed system
One user request may span router, model, retriever, several tools, approval wait, and final synthesis. Give it a trace ID and create spans for each meaningful step. Capture workflow/step name, model and prompt version, tool name, attempt, status, latency, token usage, cache outcome, budget remaining, and safe error category. Do not record raw prompts or tool payloads by default; OpenTelemetry's GenAI semantic conventions warn that message content can contain sensitive information. Chapter 11 covers the full LLMOps stack; here, own the agent-specific signals.
Observe outcomes and trajectories
End-to-end task success alone hides inefficient or unsafe routes. Measure correct tool selection, argument validity, authorization denials, unnecessary calls, repeated calls, plan changes, step success, approval rate/time, recovery success, and terminal-state distribution. Pair quality with latency and cost. Sample failed, expensive, long, denied, and novel traces into evaluation datasets (Chapter 6 covers trajectory evaluation methodology).
Use hierarchical budgets
Start with an end-to-end deadline and cost ceiling. Allocate child timeouts per dependency with room for response construction. Enforce model-call, tool-call, parallel-branch, output-token, and retry budgets. Cancel abandoned work when the client disconnects or the outcome becomes terminal. Streaming improves perceived latency but complicates error semantics: distinguish provisional progress from committed result, and never stream a claim of success before a side effect is confirmed.
Degrade by preserving the user's goal
- If the planner fails, fall back to a fixed supported workflow or ask a targeted question.
- If a nonessential tool fails, return a partial result with the missing source named.
- If the primary model is unavailable, use a validated fallback only for routes it passed; fallback must preserve the contract — schema, safety policy, tool permissions.
- If approval infrastructure is unavailable, pause; do not silently auto-approve or discard the proposal.
- If a mutating tool has ambiguous outcome, reconcile before telling the user to retry.
Interview playbook
Use AGENTS to structure a design answer:
- A — Aim and authority: user outcome, risk, actor, tenant, and actions the system may never take.
- G — Graph and state: deterministic baseline, model decisions, topology choice, transitions, invariants, terminal states.
- E — Execution contracts: narrow tools (in-process or MCP), schemas, validation, idempotency, typed results, deadlines.
- N — Non-happy paths: retries, ambiguous writes, dependency failure, no progress, rejection, crash-and-resume semantics.
- T — Telemetry and tests: traces, versions, trajectory/outcome evals, injection cases, release gates.
- S — Spend and safe rollout: token/tool/time budgets, shadow mode, approvals, canary, fallback; managed versus self-built runtime.
Common traps
- Calling a prompt chain an "agent" without identifying any runtime decision.
- Letting model-generated tenant IDs, URLs, SQL, or scopes reach a tool unvalidated.
- Assuming structured output means business-correct or authorized output.
- Retrying a timed-out mutation without idempotency or reconciliation.
- Using conversation history as authoritative workflow state.
- Proposing a multi-agent swarm where one agent with better tools and context isolation would do.
- Presenting MCP or A2A as a security mechanism rather than an interoperability layer.
- Describing human approval without binding it to exact, expiring action arguments.
- Recommending a computer-use agent for a system that has an API.
- Tracing sensitive content by default or omitting model/prompt/tool versions.
Question bank
Answer with control boundaries, measurable failure behavior, and the smallest justified architecture.
Q1When should a workflow become an agent?
Strong answer outline
- Identify runtime path uncertainty that fixed rules cannot economically cover.
- Compare value against added trajectory, safety, cost, and evaluation complexity.
- Keep deterministic invariants and propose a bounded agent with terminal states and an exit condition.
Follow-up probes
- What is the fixed-workflow baseline?
- What evidence would remove the agent?
Pass if agenticity is an earned trade-off; fail if natural-language input alone is treated as justification.
Q2Compare router, planner–executor, and supervisor patterns.
Strong answer outline
- Define who selects one route, creates a changing plan, or delegates among specialists.
- Map each to misrouting, stale plans, or opaque/expensive delegation.
- Give a concrete workload and a simpler baseline for each.
Follow-up probes
- When may branches run in parallel?
- How do you evaluate the supervisor itself?
Pass if control flow and failure modes differ clearly; fail if patterns are only framework class names.
Q3How do you prevent an agent from looping forever?
Strong answer outline
- Define terminal states and progress invariants.
- Cap wall time, calls, tokens, replans, repeated action/result fingerprints, and cost.
- Checkpoint and return a safe partial outcome or targeted question on exhaustion.
Follow-up probes
- What counts as progress?
- Can the user resume after budget exhaustion?
Pass if both static budgets and dynamic no-progress detection appear; fail if only "max iterations" is named.
Q4What makes a good tool schema?
Strong answer outline
- Narrow intent, discriminating description, constrained fields, and no model-supplied identity.
- Separate proposal/read/mutation tools and declare effects.
- Server-side business validation, authorization, limits, idempotency, and typed errors.
Follow-up probes
- How do you evolve a schema without breaking traces?
- What if arguments validate but are semantically wrong?
Pass if syntax, semantics, and authority are distinct; fail if JSON Schema is treated as the whole boundary.
Q5A payment tool times out. Should the agent retry?
Strong answer outline
- Classify the outcome as ambiguous, not failed.
- Query or reconcile by stable idempotency/operation key.
- Retry only if the remote contract is idempotent; otherwise pause/escalate and never claim completion.
Follow-up probes
- Where is the operation key generated and stored?
- What if status lookup is also unavailable?
Pass if duplicate side effects are explicitly prevented; fail if backoff alone is proposed.
Q6Compare LangGraph checkpointing with Temporal-style durable execution for agents.
Strong answer outline
- Explain checkpointing: state snapshots per superstep, thread-scoped resume, interrupts — with you owning node idempotency.
- Explain deterministic replay: event history, side effects isolated in activities, engine-owned retries/timers.
- Pick by workload: LLM-native loops with human gates versus agent steps inside long transactional business processes; note managed runtimes internalize session persistence.
Follow-up probes
- Why must side effects before a LangGraph interrupt be idempotent?
- Why can't a model call live in Temporal workflow code?
Pass if replay semantics and nondeterminism constraints are concrete; fail if "both persist state" is the depth.
Q7How should human approval work for a high-impact tool?
Strong answer outline
- Present exact target, arguments, effect, rationale, and policy evidence.
- Bind decision to actor, argument hash, policy version, expiry, and one operation; invalidate on change.
- Design the reviewer workflow: batching, diffs, rejection with feedback, audit, and safe resume.
Follow-up probes
- How do you prevent approval fatigue?
- What happens during an approval-service outage?
Pass if approval cannot be reused for an altered action; fail if "human in the loop" is a generic UI step.
Q8How do you defend against prompt injection in retrieved or browsed content?
Strong answer outline
- Treat content as data; preserve instruction hierarchy but assume the model will sometimes comply with injections.
- Contain via capability allowlists, server-side identity/scope, egress control, and approval on consequence.
- Add injection test cases to CI, trace denials safely, and review memory writes for poisoning.
Follow-up probes
- Can a classifier solve injection?
- How does injected content exfiltrate data through tool output or URLs?
Pass if containment survives a model mistake; fail if prompt wording is the only defense.
Q9Distinguish working context, workflow state, and long-term memory. How would you build agent memory?
Strong answer outline
- State is authoritative typed data for transitions/resumption; context is per-call and bounded; memory is intentionally retained across sessions.
- Split memory into episodic events and semantic facts with provenance and confidence; consolidate in the background, not on the hot path.
- Cover tenant-scoped retrieval, conflict resolution, staleness, deletion propagation, and poisoning defenses.
Follow-up probes
- Can a summary ever be authoritative?
- When does compaction lose a pending commitment, and how do you test for it?
Pass if the three stores have different contracts and memory has a write policy; fail if every prior message is called memory.
Q10When is a multi-agent system justified, and which topology would you pick?
Strong answer outline
- Require parallelism, specialization with different tool grants, permission separation, or context isolation — otherwise one agent with tools.
- Map topologies: supervisor/worker for decompose-and-merge, pipeline for staged contracts, debate only with independent evidence, swarm rarely.
- Budget inter-agent calls; evaluate contribution by ablation; name coordination failures (duplication, handoff loss) as the new defect class.
Follow-up probes
- How does token cost scale with worker count?
- How do you stop agreement theater in debate?
Pass if agents add measurable value beyond personas and a "when not" is stated; fail if complexity is the objective.
Q11What does MCP standardize, and what does it deliberately not solve?
Strong answer outline
- Describe host/client/server roles, JSON-RPC transports, and the tools/resources/prompts primitives with capability negotiation.
- Explain the M×N to M+N integration collapse and governable, framework-independent tool catalogs.
- State what remains yours: authorization, tenancy, validation, redaction, audit; cite resource-bound tokens and the token-passthrough prohibition; name tool-description mutation and confused-deputy risks.
Follow-up probes
- When is a direct in-process function simpler than an MCP server?
- How do you govern third-party MCP servers?
Pass if interoperability is separated from security policy; fail if MCP is called a secure tool bus by default.
Q12How does A2A differ from MCP, and when do you actually need it?
Strong answer outline
- MCP is agent-to-tool; A2A is agent-to-agent across trust boundaries with opaque internals.
- Describe Agent Cards for discovery and task lifecycle with streaming/push for long-running work.
- Justify A2A only for cross-org/cross-vendor delegation where credential sharing is impossible; prefer in-process orchestration within one team.
Follow-up probes
- How do you authenticate and rate-limit a peer agent?
- What do you log when the remote agent is a black box?
Pass if the trust-boundary framing is explicit; fail if A2A is proposed for two agents in the same process.
Q13Design a browser/computer-use agent for a legacy portal without an API.
Strong answer outline
- Confirm no API/MCP path exists; frame UI automation as the integration of last resort with compounding per-step error and cost.
- Architecture: sandboxed browser profile, domain allowlist, default-deny egress, credential injection at a proxy, action validator, human gates on login/payment/destructive steps.
- Operations: full action/screenshot audit trail, replayable traces, benchmark-informed reliability expectations, fallback to human completion.
Follow-up probes
- How does a malicious page attack the agent?
- What breaks when the portal redesigns its front end?
Pass if the page is treated as adversarial input and secrets never reach the model; fail if reliability is assumed.
Q14Bedrock Agents/AgentCore versus Vertex AI Agent Engine/ADK versus self-built LangGraph — how do you choose?
Strong answer outline
- Decompose the platform problem: runtime isolation, tool gateway, memory, identity, observability.
- Map offerings: Bedrock Agents / console agents as managed orchestrators; AgentCore and Agent Engine as framework-agnostic runtimes; self-built as maximum control with owned security burden.
- Recommend by team maturity, compliance, and portability: often your graph on their runtime, with lock-in analyzed at the memory/identity layer, not the model layer.
Follow-up probes
- Where exactly is the lock-in in each option?
- How would you migrate memory between platforms?
Pass if trade-offs are layer-by-layer with a contextual recommendation; fail if it is vendor cheerleading or reflexive build-it-yourself.
Q15How would you sandbox an agent that executes generated code?
Strong answer outline
- Isolate per session in a microVM/gVisor-class sandbox, never in the orchestrator process.
- Default-deny egress with domain allowlist; no ambient credentials — broker short-lived scoped tokens server-side.
- Cap CPU/memory/time/processes, audit commands and network attempts, and destroy the sandbox after the session.
Follow-up probes
- Why is egress control the highest-value guardrail?
- What changes when the sandbox needs package installation?
Pass if exfiltration and credential theft are the named threats; fail if a Docker container with open network is called a sandbox.
Q16Give your view on the limits of autonomous agents in production today.
Strong answer outline
- Acknowledge value in uncertain, reversible information work with human gates on consequence.
- Name brittleness: compounding step errors, injection exposure, opaque trajectories, cost variance, and accountability gaps; cite computer-use benchmark gaps as evidence.
- Advocate bounded autonomy: deterministic invariants, approval by consequence class, trajectory evals, durable execution, and gradual rollout.
Follow-up probes
- Which measurable capability change would expand your autonomy budget?
- Where would you deploy full autonomy today?
Pass if the position is nuanced and operational; fail if it is categorical hype or dismissal.
Proof artifact: a resumable, approval-bound agent
Build a support workflow against a fake order service. It may read an order, retrieve a policy, propose a refund, and — only after deterministic policy checks — execute a low-value simulated refund or request human approval. Expose the read tools through a small MCP server to demonstrate protocol fluency. Use synthetic data and fake funds. The artifact demonstrates controls, not production outcomes.
Steps
- Authority firstWrite the authority matrix and terminal states; mark transitions as deterministic, model-selected, or human-controlled.
- ContractsImplement typed state and narrow read/proposal/execution tools; serve reads via MCP; inject actor and tenant server-side; add idempotency keys to execution.
- DurabilityPersist LangGraph checkpoints, version prompt/model/tool schemas, and support resume after process kill — including an approval interrupt held overnight.
- TelemetryAdd end-to-end traces with safe metadata: per-step timing/tokens, attempts, budget, approval, and terminal reason.
- EvaluationBuild a scenario set for route choice, argument accuracy, policy result, trajectory length, injection resistance, and recovery.
- Rollout drillRun shadow mode over synthetic scenarios, then enable only the reversible fake action behind a feature flag.
Metrics
- Task completion and correct terminal-state rate by scenario.
- Tool selection precision/recall, argument validity, and business-invariant pass rate.
- Unauthorized action attempts and cross-tenant disclosures — both must be zero in the test suite.
- Median/p95 model calls, tool calls, tokens, wall time, and example cost per terminal state.
- Duplicate side effects after crash-and-replay — must be zero with idempotent fake execution.
- Checkpoint recovery and user-visible safe-degradation success rate.
Deliberate failure injection
Crash after the fake refund service commits but before the workflow records success; resume and prove no duplicate. Time out policy retrieval; exhaust the model-call budget; change a proposal after approval and verify approval invalidation; return malicious instructions inside a policy document served over MCP; submit a model-generated tenant ID; and make the approval service unavailable. Capture trace and terminal behavior for each.
What to present
Present the state diagram, authority matrix, one tool schema, the MCP server manifest, a crash-and-resume trace, an injection-denial trace, evaluation results by slice, and a short argument for which steps were deliberately kept non-agentic — plus what you would delegate to AgentCore or Agent Engine in a production version. Label every metric as a synthetic artifact measurement.
Chapter review
Reliable agent engineering is control engineering around probabilistic decisions. Keep authority in trusted code, express state and termination explicitly, make tools narrow and replay-safe, standardize integration with MCP without outsourcing security to it, persist checkpoints or event histories for resumability, bound every resource, sandbox anything that executes, and evaluate trajectory as well as outcome. Add autonomy — and additional agents — only where runtime uncertainty creates measured value.
Glossary
- Agent
- A system in which a model chooses part of the action sequence at runtime.
- Capability
- An explicitly granted operation available to a workflow, distinct from what a model requests.
- Checkpoint
- A persisted workflow boundary from which execution can be inspected or safely resumed.
- Idempotency key
- A stable operation identifier that lets a service return the same result without repeating the effect.
- Interrupt
- A deliberate pause that persists state and awaits external input such as approval.
- MCP
- Model Context Protocol: standard host/client/server wiring exposing tools, resources, and prompts.
- A2A
- Agent2Agent protocol for delegating tasks between opaque agents across trust boundaries.
- Episodic memory
- Time-stamped records of what happened in past sessions; the raw input to consolidation.
- Semantic memory
- Distilled durable facts with provenance and confidence, extracted from episodes.
- Durable execution
- Running workflows so crashes resume from persisted state or replayed event history, not from scratch.
- Trajectory
- The ordered sequence of model decisions, tool calls, observations, and transitions.
- Sandbox
- An isolated execution environment with egress control, no ambient credentials, and resource caps.
Mastery checklist
- I can justify every model-selected step against a deterministic baseline.
- I can draw states, guarded transitions, terminal outcomes, and recovery paths.
- I can pick a multi-agent topology — or reject multi-agent — from parallelism, permissions, and context isolation.
- I can design a tool whose schema, authority, effect, and errors are explicit, in-process or over MCP.
- I can explain MCP primitives and A2A's trust-boundary role without outsourcing security to either.
- I can contrast checkpointing and deterministic replay, and state their idempotency obligations.
- I can design episodic/semantic memory with consolidation, provenance, and deletion.
- I can bind human approval to one exact, expiring action and design the reviewer workflow.
- I can sandbox code-executing and browser agents against exfiltration and credential theft.
- I can compare AgentCore, Agent Engine/ADK, and self-built LangGraph layer by layer.
Primary sources
Links checked . Provider APIs and protocol revisions change; pin versions and re-check deployed semantics.
- Model Context Protocol — specification and safety principles
- Model Context Protocol — authorization and resource-bound tokens
- A2A — Agent2Agent protocol specification and Agent Cards
- LangGraph — checkpoints, threads, pending writes, and recovery
- LangGraph — interrupts and human-in-the-loop resumption
- Temporal — durable execution, workflows, and activities
- AWS — Amazon Bedrock Agents user guide
- AWS — Amazon Bedrock AgentCore (Runtime, Gateway, Memory, Identity)
- Google Cloud — Vertex AI Agent Builder
- Google Cloud — Vertex AI Agent Engine overview
- Google — Agent Development Kit documentation
- Anthropic — Building effective agents
- Anthropic — tool-use execution boundary and computer use
- OpenAI — function calling lifecycle and tool definitions
- OpenTelemetry — semantic conventions, including GenAI instrumentation
- Du et al. — Improving Factuality and Reasoning through Multiagent Debate (arXiv)
- OSWorld — benchmarking computer-use agents in real environments (arXiv)
- WebArena — a realistic web environment for autonomous agents (arXiv)
CHAPTER 06 · PRIORITY 0
Evaluation & AI Quality Engineering
33 min read · 16 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- Translate product value and failure consequences into component, end-to-end, safety, and operational quality criteria — and into an explicit release decision.
- Build a versioned, representative golden set with provenance, grouped splits, and contamination controls, and explain why public benchmark scores overstate capability on your task.
- Apply the RAG triad and a layered evaluation pyramid: deterministic checks first, calibrated LLM judges for semantics, humans for ground truth and high-risk ambiguity.
- Evaluate agents on terminal state, step correctness, tool-call accuracy, trajectory efficiency, and
pass^kreliability — not just final prose. - Name and mitigate LLM-judge failure modes: position bias, verbosity bias, self-preference, and judge-targeted prompt injection.
- Design CI release gates, a red-team program, and online experiments (shadow, interleaving, canary with guardrail metrics) that fail closed.
- Compare open-source and managed evaluation tooling on AWS and GCP, and run an eval-driven development loop your team actually follows.
1. Quality is a decision system, not a score
An evaluation system exists to make decisions: continue iterating, merge a change, canary it, expand rollout, roll back, investigate a slice, or escalate a risk. Start by writing the decision and the consequence of a false pass or false fail. Only then choose metrics and thresholds. A single “quality score” cannot represent correctness, evidence use, safety, latency, cost, and user value without hiding important trade-offs.
flowchart TD
P["Product outcome"] --> E["End-to-end task success"]
P --> C["Component quality (retrieve / generate / tools)"]
P --> I["Hard invariants (auth / schema / policy / citations)"]
P --> G["Operations (latency / cost / errors / recovery)"]
E --> D["Release decision: ship / canary / hold / rollback"]
C --> D
I -->|"any failure blocks"| D
G -->|"guardrails"| D
Build a quality tree
For an enterprise policy assistant, the top outcome might be “authorized users resolve policy questions accurately and quickly.” Decompose it into: relevant current evidence retrieved; answer claims supported; citations resolvable; uncertainty handled; unauthorized information never disclosed; correct escalation on ambiguous or high-risk questions; and acceptable latency/cost. Attach at least one measure and one failure example to each leaf.
| Criterion | Measure | Decision role |
|---|---|---|
| Required evidence is present | Recall@k / evidence coverage | Diagnose retriever and packer |
| Claims follow evidence | Human or calibrated claim-level faithfulness | Generation quality gate |
| Answer resolves task | Task-specific rubric / execution success | End-to-end comparison |
| Tenant boundary holds | Deterministic adversarial test | Non-negotiable release blocker |
| User experience is timely | p50/p95/p99 end-to-end and stage latency | Guardrail / capacity decision |
| Economics are viable | Cost and tokens per successful task | Route/rollout decision |
2. Engineer the dataset before the evaluator
A golden set is a versioned collection of inputs, expected properties, metadata, and judgments. “Golden” means reviewed, traceable, and stable enough for comparison — not perfect or frozen. Every serious platform (LangSmith, Langfuse, Vertex AI, Bedrock) is organized around this object; if your dataset is an untracked spreadsheet, no tool downstream can save you.
Four complementary sources
- Curated core: expert-written canonical and boundary cases.
- Production traces: privacy-reviewed samples of common, failed, costly, uncertain, and novel interactions.
- Adversarial cases: injection, leakage, malformed input, unavailable tools, contradictory sources, and no-answer examples.
- Synthetic expansion: reviewed, provenance-labeled paraphrases or rare combinations; a supplement, never proof of representativeness.
Each record should include a stable ID, input, reference evidence or expected behavior, rubric, risk, source/provenance, created/reviewed dates, language, tenant/data class, intent, difficulty, and applicable evaluators. Preserve an immutable raw trace reference separately when allowed. Redact or synthesize sensitive values before placing cases in developer-visible stores. Split by how data can leak: group by source document, user/thread, template, or time so near-duplicates never cross development, validation, and sequestered test splits. Repeated optimization makes validation data de facto training data; preserve a final untouched set. Version everything that changes meaning — dataset, judgments, rubrics, corpus/parser/retrieval, prompt/model/tools, evaluators, sampling, environment — with hashes and a changelog. A score without dataset and evaluator versions is not reproducible evidence.
Benchmark contamination: why public leaderboards mislead
Public benchmarks (MMLU, HumanEval, GSM8K and successors) circulate on the open web, which means they leak into pretraining corpora. The GSM1k study (arXiv:2405.00332) rebuilt grade-school math problems of matched difficulty from scratch and found some model families dropped by double-digit accuracy points versus their GSM8K scores — evidence of memorization, not reasoning. For an interview, the takeaway is a posture: treat vendor benchmark claims as marketing until reproduced on your task distribution, and treat any public test set as presumptively contaminated.
- Private, post-cutoff data — author fresh cases from your own domain; prefer material created after the model's training cutoff.
- Canary strings — embed unique GUIDs (the BIG-bench convention) in eval files so future contamination is detectable.
- Overlap checks — run n-gram and embedding-similarity dedup between eval items and any corpus you fine-tune on.
- Rotation — refresh sequestered sets on a schedule; retire items once they have influenced many decisions.
3. The evaluation pyramid and the RAG triad
Component scores localize defects; end-to-end scores reveal interactions. Retrieval recall can rise while excess context harms answers; correct tool selection can still carry invalid arguments. Layer evaluators by cost and coverage: cheap deterministic checks run on everything, calibrated model judges on samples, expert humans on a small stratified slice that continuously re-anchors the judges.
flowchart TD
A["Every output: deterministic checks (schema, citations, execution, policy)"] --> B["Every experiment: component metrics (recall@k, nDCG, tool accuracy)"]
B --> C["Sampled: calibrated LLM judges (faithfulness, relevance, rubric scores)"]
C --> D["Small stratified sample: expert human review + adjudication"]
D -->|"labels recalibrate judges"| C
D -->|"new cases join the golden set"| A
The RAG triad, by name
The industry-standard decomposition (popularized by TruLens and mirrored in Ragas, Vertex, and Bedrock metric catalogs) scores three edges of the query–context–answer triangle. Use the names — interviewers listen for them — but always state the rubric behind each, because tools define them differently.
| Triad edge | Question it answers | Typical failure it isolates |
|---|---|---|
| Context relevance (query ↔ context) | Is the retrieved evidence actually about the question? | Retriever/ranker pulls plausible but off-topic chunks |
| Faithfulness / groundedness (context ↔ answer) | Is every material claim supported by the supplied evidence? | Generation hallucinates beyond or against the context |
| Answer relevance (query ↔ answer) | Does the response address the user's actual request? | Faithful summary of evidence that dodges the question |
Faithfulness and correctness differ: an answer can repeat stale evidence faithfully, or be correct yet unsupported. Evaluate both, plus completeness and citation correctness separately — a faithful answer can still omit the decisive policy exception. Complement the triad with retrieval-stage metrics (precision@k, recall@k, MRR, nDCG, evidence coverage after packing) and answer-stage checks (instruction adherence, appropriate abstention). Chapter 4 covers the retrieval mechanics; here your job is choosing which edge a metric belongs to so a regression routes to the right owner.
Deterministic output checks first
Use exact match for known classifications, JSON Schema for structure, parsers/compilers for code or queries, executable tests for calculations, database comparisons for extraction, and policy engines for allowed actions. A deterministic checker is cheaper, faster, repeatable, and easier to debug than a model judge whenever the property is mechanically decidable. Reaching for an LLM judge to validate JSON is a junior tell.
4. Agent evaluation: trajectories, tools, and pass^k
Agents (chapter 5) break response-only evaluation because quality lives in a trajectory of decisions with real side effects. Grade three distinct things and keep them separate: did the world end in the correct state (task completion), were the individual decisions right (step correctness), and was the path economical (trajectory efficiency). An agent can reach the right terminal state through a wasteful or policy-violating path, and it can execute every step plausibly while never finishing the job.
flowchart LR
T["Task instance"] --> R["Agent run: plan, tool calls, observations"]
R --> F["Final-state check: is the environment state correct"]
R --> S["Step grading: tool choice, arguments, ordering, policy"]
R --> Y["Trajectory metrics: calls, retries, tokens, wall time"]
F --> V["Per-run verdict"]
S --> V
Y --> V
V --> K["pass^k across k repeated i.i.d. runs"]
What to measure at each level
- Task completion: correct terminal state verified against the environment (database row, ticket status, calendar entry) — not the agent's claim that it finished.
- Tool-call accuracy: tool-selection precision/recall, argument schema and semantic validity, authorization result, unnecessary-call rate.
- Trajectory match: against a reference trajectory — exact match, in-order match (allows extra steps), any-order match, and precision/recall over reference steps. Vertex AI's evaluation service ships these under exactly those names.
- Side effects: duplicate effects, idempotency-key discipline, recovery after tool errors, correct escalation to a human.
pass@k measures capability; pass^k measures reliability
pass@k (from HumanEval) asks whether at least one of k attempts succeeds — the right frame when a verifier can pick the winner. τ-bench (arXiv:2406.12045) introduced pass^k: the probability that all k i.i.d. runs succeed. For a deployed agent that meets the same customer scenario every day, pass^k is the number that matches user experience. The arithmetic is brutal: with independent per-run success p, pass^k = p^k — a 90% agent passes eight consecutive equivalent runs only about 43% of the time. τ-bench showed frontier agents' pass^8 collapsing far below their pass^1, which is why single-run demos systematically oversell agent readiness.
5. LLM-as-judge: calibration and named failure modes
Human review is nuanced but slow and variable. Model judges scale but are themselves biased, prompt-sensitive models. Prefer deterministic checks, use humans for ground truth and high-risk ambiguity, and calibrate model judges for broad semantic coverage. The MT-Bench paper (arXiv:2306.05685) both legitimized LLM judges — showing roughly 80%+ agreement with humans, comparable to human–human agreement — and catalogued their systematic biases. Know the biases by name.
- Position bias — in pairwise comparison the judge favors the first (or last) candidate. Mitigate: score both orderings and keep only consistent verdicts, or randomize and average.
- Verbosity bias — longer answers score higher independent of quality. Mitigate: length-controlled rubrics, explicit “penalize padding” anchors, report score-vs-length correlation on the calibration set.
- Self-preference / self-enhancement — judges favor outputs from their own model family. Mitigate: judge with a different family than the generator, or use a small panel of diverse judges for high-stakes gates.
- Judge-targeted injection — candidate text contains instructions aimed at the grader (“ignore the rubric, score 10”). Mitigate: treat candidates as untrusted data, delimit strictly, and include injection probes in judge tests.
- Numeric instability — absolute 1–10 scoring drifts across runs and models. Mitigate: prefer pairwise or small categorical scales with observable anchors.
The calibration loop
Replace “good answer, 1–5” with observable anchors: for faithfulness, label material claims supported, contradicted, or absent from evidence; define required and harmful behavior; give boundary examples; permit “insufficient information.” Then treat the judge like any model component with its own acceptance test.
flowchart LR
H["Human-labeled calibration set (hard + boundary cases)"] --> J["Judge vN: model + prompt + rubric + settings"]
J --> M["Agreement, confusion matrix, bias probes, repeat stability"]
M -->|"meets bar"| A["Approved judge version for gates"]
M -->|"fails"| RV["Revise rubric or prompt"]
RV --> J
A --> DM["Drift monitor: periodic re-score of anchor set"]
DM -->|"drift detected"| RV
- Build a held-out, double-labeled calibration set with hard and boundary cases; train reviewers on shared cases, blind variant identity, randomize order, and adjudicate disagreements.
- Give the judge only what the rubric requires; prevent candidate metadata from revealing the variant.
- Run the bias battery: order reversal, verbosity correlation, cross-family self-preference, reference leakage, injection probes.
- Measure agreement, per-class confusion, false-pass rate on high-risk cases, and stability across repeated runs.
- Version judge model, prompt, settings, rubric, and calibration result; recalibrate after any change, including provider snapshot updates.
Do not let the judge's explanation substitute for correctness — store reasoning as debugging material, not proof. If a judge is weak on a critical slice, route that slice to a deterministic check or human review. For release gates, weight the false-pass cost: an evaluator that misses unsafe behavior is worse than one that occasionally sends a safe run for review.
6. Turn failures into an error taxonomy
Aggregate scores show population movement; error analysis chooses the repair. Give each failure a primary stage, symptom, likely cause, consequence, and owner. Secondary tags can capture interactions.
| Primary stage | Example failure | Likely owner or experiment |
|---|---|---|
| Data/parse | Table row or policy exception lost | Parser/chunking fixture and reprocessing |
| Retrieval | Relevant evidence absent from candidates | Embedding, sparse route, filters, ANN depth |
| Ranking/packing | Evidence found then dropped or truncated | Fusion, reranker, dedupe, token allocation |
| Generation | Unsupported claim despite sufficient evidence | Prompt/model/grounding control |
| Tool/control | Wrong tool, invalid argument, repeated effect | Schema, policy, state machine, idempotency |
| Safety/privacy | Injection obeyed or cross-tenant disclosure | Authorization boundary and incident response |
| Operations | Timeout, cost cap, stale version, failed fallback | Budgets, capacity, recovery, routing |
| Evaluation | Label/rubric/judge is wrong | Adjudication and evaluator recalibration |
Slice before celebrating, and inspect paired deltas
Choose slices from risk and plausible causes: intent, language, policy regime, tenant/data class, document type, exact identifier, multi-hop, no-answer, tool, route, cohort, and age. Report counts and uncertainty; define critical slices before the experiment. Then list paired improvements and regressions per case: equal means can hide replacing harmless style errors with one security failure. Inspect the largest negative deltas and every invariant violation; track severity and consequence.
7. Release gates that tolerate variability, not regressions
CI compares immutable baseline and candidate configurations on a pinned dataset/evaluator suite. Record environment and inspectable responses when policy permits. Invalidate caches for every changed prompt, model, retrieval, or tool dimension.
flowchart LR
CH["Change: prompt, model, retrieval, or tool"] --> SM["Smoke suite: fast, includes all high-risk cases"]
SM --> INV["Hard invariants: 100% required"]
INV -->|"any failure"| BL["Block + per-case diff report"]
INV --> RG["Paired regression vs pinned baseline + critical slices"]
RG -->|"delta beyond budget"| BL
RG --> GD["Guardrails: p95 latency, cost per success"]
GD -->|"breach"| BL
GD --> CN["Shadow, then canary"]
CN -->|"gates hold 48h"| RP["Ramp"]
def release_decision(base, candidate):
hard_fail = any(candidate[name] != 1.0 for name in (
"tenant_isolation", "schema_valid", "no_duplicate_effect"
))
quality_drop = candidate["task_success"] < base["task_success"] - 0.02
slow = candidate["p95_ms"] > 1.10 * base["p95_ms"]
costly = candidate["cost_per_success"] > 1.15 * base["cost_per_success"]
return "block" if hard_fail or quality_drop or slow or costly else "canary"
# Thresholds above are illustrative; derive real gates from product risk.
Hard invariants require every applicable case to pass. Comparative metrics need both absolute floors and allowable deltas. Critical slices need their own gates. Cost and latency are guardrails. Treat missing evaluator output as a failure or explicit “inconclusive,” never as a pass. Keep a small smoke suite on every change and a larger suite on scheduled runs or release candidates, while ensuring high-risk cases remain in the fast gate.
Account for stochastic and sampling uncertainty: use paired comparisons and report intervals via paired bootstrap; inspect discordant binary outcomes; repeat a stratified subset to estimate run-to-run variance. Never average away safety failures, and weigh practical — not only statistical — significance.
8. Online experimentation and the offline–online loop
Offline datasets provide controlled repeatability; production provides distribution reality. Instrument traces with application, prompt, model, retrieval, tool, evaluator, and release versions plus safe outcome metadata. Sample common traffic randomly for prevalence, and oversample rare/high-risk signals for discovery — but keep the weighting explicit: a risk-enriched review queue cannot estimate population quality without correcting its sampling design.
flowchart LR
DS["Versioned golden set"] --> OF["Offline experiment"]
OF --> CI["CI release gate"]
CI --> SH["Shadow / interleave / canary"]
SH --> PR["Production traffic"]
PR --> TR["Traces + online scores + user signals"]
TR --> SP["Random + risk-weighted sampling"]
SP --> AJ["Human adjudication"]
AJ --> DS
Experiment designs that fit GenAI
Classic A/B testing works but is sample-hungry, and GenAI quality deltas are often small relative to outcome noise. Two adaptations matter. First, interleaving: at the retrieval/ranking layer, blend results from two rankers in the same session (team-draft interleaving) and score which side earns the click or citation — within-session comparison removes between-user variance and reaches significance with a fraction of the traffic. For full generations, the analogue is paired preference: run both variants on the same prompt (one served, one shadowed) and collect judge or human preferences on the pairs. Second, guardrail metrics as stop rules: predefine p95 latency, cost per session, refusal rate, safety-flag rate, thumbs-down rate, and escalation-to-human rate with automatic stop thresholds, monitored sequentially — you are not waiting for the primary metric to go wrong before pulling an unsafe variant. Randomize by user or session, never by request, to avoid within-user contamination and inconsistent experiences.
- ShadowRun the candidate on mirrored traffic with no user exposure; diff outputs, latency, and cost offline.
- Interleave / paired preferenceWithin-session comparison at the ranking layer, or judged preference on shadowed generation pairs.
- Canary 5%Real exposure gated on guardrail metrics with automatic stop rules and a rollback owner.
- RampPromote when primary and guardrail gates hold for a predefined window; keep the holdback for measurement.
Production signals are evidence, not ground truth. Feedback, completion, abandonment, reformulation, escalation, citation clicks, and corrections are each confounded: clicks reflect position; silence may mean abandonment or satisfaction. Calibrate proxies against reviewed traces before trusting them in a decision, and never expose an unsafe variant merely to gain statistical power.
9. Red teaming and adversarial evaluation
Safety testing is a program, not a checklist pass. Build adversarial cases from actual input surfaces — user text, retrieval, tools, files, connectors, memory, tenant data — using the OWASP GenAI LLM Top 10 as the threat catalog and the NIST Generative AI Profile to structure governance, then translate both into system-specific executable tests.
Test the full threat path
- Prompt injection: direct user instructions and indirect instructions embedded in retrieved pages, documents, tool output, or memory.
- Sensitive information: secrets, PII, hidden prompts, credentials, and private records requested directly or inferred through side channels.
- Tenant isolation: IDs from another tenant, mixed-index candidates, cached responses, trace views, and shared memory.
- Unsafe tool use: unauthorized tool, excessive scope, altered action after approval, malicious URL/arguments, duplicate effect, ambiguous timeout.
- Policy behavior: refusal consistency, over-refusal on benign requests, safe alternatives, correct human escalation.
- Robustness: malformed encoding, extreme length, empty/contradictory evidence, unavailable dependencies, partial streaming.
Measure attack success rate (ASR) per surface, sensitive-data disclosure, unauthorized-action rate, false refusal, escalation precision/recall, time to detection, and recovery behavior. Automated red teaming — attacker LLMs mutating seed attacks, tools like promptfoo's red-team mode or Bedrock Guardrails test suites — expands coverage cheaply, but manually validate what the generator missed, and keep a sequestered attack pool so defenses are not tuned to the public probes. Every successful attack becomes a permanent regression case: the red team feeds the golden set. Hard security boundaries must be enforced deterministically outside the model and pass every applicable test — chapter 11 covers the runtime enforcement side.
Protect the evaluation system itself. Datasets and traces contain your most revealing failures: apply minimization, access control, encryption, retention/deletion, tenant partitioning, and audit. Treat candidate text as untrusted judge input; sandbox code evaluators with minimal permissions; send production content to external eval services only under an approved data contract.
10. Tooling landscape and eval-driven development
Choose tools by data model and exit path, not dashboards: can it represent your datasets, ground truth, evaluator provenance, versions, and CI integration — and can you export everything if you leave? Keep manifests and deterministic evaluators portable regardless of platform.
| Open-source tool | Center of gravity | Questions before adopting |
|---|---|---|
| Ragas | RAG triad + agent metric library | Do its judge prompts and definitions correlate with your domain humans? |
| promptfoo | Config-driven CI evals + automated red teaming | Does declarative YAML cover your trace-level assertions? |
| DeepEval | pytest-style unit tests for LLM outputs | Who calibrates the built-in judges against your labels? |
| Arize Phoenix | OTel-native tracing + eval on traces, self-hostable | Does your OpenTelemetry convention match its semantics? |
| Langfuse | Traces, datasets, scores, annotation queues; self-host option | Deployment operations, retention, export, feature parity across versions? |
| LangSmith | Datasets, experiments, human/code/model/pairwise evaluators | Framework coupling, hosting/data policy, cost at trace volume? |
AWS
- Bedrock Evaluationsmodel + RAG evaluation jobs; LLM-as-judge and human workflows
- SageMaker Clarify / fmevalfoundation-model evaluation, bias and toxicity checks
- Bedrock Guardrailsonline policy enforcement + grounding checks the evals must mirror
- CloudWatchguardrail metrics, canary alarms, rollback triggers
Google Cloud
- Vertex AI Gen AI evaluation servicepointwise/pairwise judges, RAG and agent trajectory metrics
- Vertex AI Experimentsrun/version comparison and lineage
- Model Armorinjection screening the adversarial suite should exercise
- BigQuery + Cloud Monitoringtrace analytics, slice dashboards, stop-rule alerts
Custom Python + pytest
Maximum transparency and exact product contracts; you build dataset UI, annotation, and trace joins yourself. Right for hard invariants and small teams with strong opinions.
OSS library + observability platform
Ragas/DeepEval metrics over Phoenix or Langfuse traces; self-hostable for data-residency constraints. Right when you need trace-level evals and control the stack.
Managed cloud service
Bedrock Evaluations or Vertex eval service; lowest setup cost, native IAM and data governance, judge models on tap. Right when the workload already lives on that cloud and export paths are verified.
Eval-driven development as a culture
The highest-leverage practice is procedural, not technical: write the eval before the fix. A reported failure becomes a reproducing case (plus a neighborhood of variants) before anyone touches the prompt; the change merges only when the new cases pass and the gate holds. Complement that with: no prompt/model change without an experiment link in the PR; domain experts — not only engineers — own rubrics; a weekly error-analysis review that walks new taxonomy entries; and an explicit eval compute budget (teams commonly spend a meaningful fraction of inference spend on evaluation — treat it as an engineering line item, not overhead; exact ratios are product-specific). One current-events note: as checked on 2026-08-04, OpenAI's docs schedule the legacy Evals platform for read-only status in late 2026 in favor of Datasets — a reminder not to design a durable eval program around any surface with a published retirement date.
Interview playbook
Use QUALITY to answer an evaluation-system design prompt:
- Q — Question and consequence: Which release/product decision, and what does a false pass cost?
- U — User distribution: traffic, important slices, risks, and no-answer/edge behavior.
- A — Artifacts and annotations: dataset sources, provenance, rubric, splits, versions, contamination controls, privacy.
- L — Layers of measures: deterministic, component (RAG triad, tool accuracy), end-to-end, human, calibrated judge, safety, operations.
- I — Inspect errors: paired deltas, taxonomy, severity, slices, uncertainty, and evaluator failures.
- T — Threshold and trial: hard invariants, regression gates, shadow/interleave/canary with guardrail stop rules, rollback.
- Y — Yield feedback: production sampling, adjudication, new cases, red-team regressions, ownership, change cadence.
Common traps
- Choosing metrics before defining the product decision and failure cost.
- Using one aggregate judge score with no rubric, calibration, bias battery, or slice analysis.
- Quoting public benchmark scores as evidence of task fitness — contamination makes them upper bounds at best.
- Grading agents on final prose while ignoring terminal state, side effects, and
pass^kreliability. - Calling synthetic questions representative without validation against production.
- Tuning prompt, threshold, and evaluator on the same test set and reporting it as generalization.
- Allowing quality gains to compensate mathematically for security or privacy violations.
- Treating user feedback, clicks, or judge explanations as uncontested ground truth.
- Building dashboards without an owner, alert/action threshold, or rollback path.
Question bank
These questions test whether evaluation evidence can support a production release decision.
Q1How do you define “good” for a RAG assistant?
Strong answer outline
- Start from user task and failure consequences; write the release decision first.
- Decompose into the RAG triad (context relevance, faithfulness, answer relevance) plus retrieval metrics, citations, abstention, safety, latency, cost.
- Mark objectives, guardrails, and hard invariants separately.
Follow-up probes
- Can a faithful answer be wrong?
- Which single metric blocks release?
Pass if quality is a decision-linked hierarchy with named triad edges; fail if “accuracy and helpfulness” are the only criteria.
Q2How would you build a representative golden set?
Strong answer outline
- Combine curated core, privacy-reviewed production traces, adversarial cases, and reviewed synthetic expansion.
- Annotate provenance, risk, slices, evidence, rubric, and expected behavior.
- Group splits to prevent near-duplicate/source leakage; version the manifest with hashes.
Follow-up probes
- How do you find rare failures?
- When is a case removed versus quarantined?
Pass if distribution, leakage, provenance, and maintenance are explicit; fail if size is the main quality claim.
Q3Explain the RAG triad and how it localizes failures.
Strong answer outline
- Context relevance scores query↔context; faithfulness scores context↔answer; answer relevance scores query↔answer.
- Each edge isolates a different owner: retriever/ranker, generation grounding, or instruction following.
- Add completeness and citation correctness separately; use fixed-context experiments to isolate generation from retrieval.
Follow-up probes
- Can all three edges score high while the answer is still wrong?
- Where does staleness show up in the triad?
Pass if each edge routes to a distinct repair; fail if one blended “RAG score” is used.
Q4When should you use an LLM as a judge, and when not?
Strong answer outline
- Use for semantic/subjective criteria not cheaply decidable in code, at a scale humans cannot cover.
- Never for mechanically decidable properties (schema, citations resolving, policy) — deterministic checks are cheaper and exact.
- Always with an anchored rubric, calibration against human labels, and a bias battery.
Follow-up probes
- Why is pairwise often more stable than absolute scoring?
- How do you detect judge drift after a provider snapshot change?
Pass if judge error is measured and governed; fail if a strong model is assumed objective.
Q5Name the known LLM-judge biases and your mitigations.
Strong answer outline
- Position bias: swap candidate order, keep only consistent verdicts.
- Verbosity bias: length-controlled rubrics; monitor score-length correlation on the calibration set.
- Self-preference: judge from a different model family, or a diverse judge panel for high-stakes gates.
- Judge-targeted injection: treat candidates as untrusted data and include injection probes in judge tests.
Follow-up probes
- Which bias did MT-Bench document, and how large was human–judge agreement?
- What is your acceptance bar for approving a judge version?
Pass if biases are named with concrete mitigations and a calibration loop; fail if “we use GPT-x as judge” ends the answer.
Q6Why can an aggregate improvement be unsafe to ship?
Strong answer outline
- Averages weight severity and slices poorly and can hide invariant violations.
- Inspect paired regressions, critical slice gates, and the error taxonomy.
- Give a concrete case such as cross-tenant leakage or strict-filter recall loss behind a rising mean.
Follow-up probes
- How do you choose critical slices in advance?
- What if a critical slice has only 15 cases?
Pass if counts, severity, and uncertainty constrain the decision; fail if slicing is retrospective cherry-picking.
Q7Design a CI gate for a prompt or model change.
Strong answer outline
- Pin baseline/candidate, dataset, corpus, evaluators, versions, and environment.
- Run hard invariants at 100%, overall and critical-slice floors, paired deltas, and latency/cost guardrails.
- Fail closed on missing critical results; publish per-case diffs; canary only after passing.
Follow-up probes
- How do you keep CI affordable at hundreds of cases per run?
- How do you handle flaky judge verdicts?
Pass if reproducibility, diagnostics, and rollout follow the score; fail if one average threshold merges all risks.
Q8How do you account for non-determinism, and what is pass^k?
Strong answer outline
- Paired cases, pinned versions/settings; estimate repeat variance on a stratified subset; report intervals, not points.
- pass@k measures capability (any of k succeeds); pass^k measures reliability (all k succeed) — for agents facing the same scenario repeatedly, pass^k matches user experience.
- With per-run success p, pass^k = p^k: a 90% agent passes 8 consecutive runs ~43% of the time — quantify before shipping.
Follow-up probes
- Should you retry a failed eval case?
- What breaks the i.i.d. assumption behind p^k?
Pass if uncertainty and reliability change the decision procedure; fail if rerunning until green is acceptable.
Q9How do evaluation-set leakage and benchmark contamination differ, and how do you defend against each?
Strong answer outline
- Internal leakage: near-duplicates crossing splits or repeated tuning against validation — defend with grouped splits, sequestered sets, and access control.
- Benchmark contamination: public test data in pretraining corpora inflates scores (GSM1k showed double-digit drops on fresh equivalents for some models).
- Defend with private post-cutoff data, canary strings, overlap checks, and rotation; treat vendor benchmark claims as upper bounds.
Follow-up probes
- Can synthetic paraphrases cross splits?
- How would you test whether a model has memorized your eval set?
Pass if both mechanisms have distinct, concrete defenses; fail if random row splitting is assumed sufficient.
Q10How would you evaluate a tool-using agent?
Strong answer outline
- Grade three layers separately: terminal state verified against the environment, step correctness (tool selection precision/recall, argument validity, authorization), and trajectory efficiency (calls, retries, tokens, wall time).
- Use trajectory-match metrics (exact, in-order, any-order, precision/recall) against reference trajectories where they exist.
- Inject tool errors, ambiguous writes, approval changes, and malicious tool results; check side effects, idempotency, recovery, escalation; report pass^k.
Follow-up probes
- Can two different trajectories both pass?
- How do you grade a partial success with a harmful side effect?
Pass if trace and system effects matter beyond final prose; fail if answer relevance is the primary agent metric.
Q11What is a useful error taxonomy for RAG?
Strong answer outline
- Separate parse/data, retrieval, rank/pack, generation, citation, tool/control, safety, operations, and evaluator errors.
- Assign primary cause, severity, slice, and owner; secondary tags for interactions.
- Review taxonomy coverage and merge/split labels only when actionability improves.
Follow-up probes
- What if multiple stages contribute to one failure?
- How does taxonomy change prioritization?
Pass if labels route to experiments or owners; fail if categories are just “hallucination” and “bad retrieval.”
Q12How do offline and online evaluation work together, and where does interleaving fit?
Strong answer outline
- Offline provides controlled regression comparison; production provides distribution reality and rare failures.
- Rollout ladder: shadow, interleave or paired preference, canary with guardrail stop rules, ramp; randomize by user/session.
- Interleaving gives within-session ranker comparison at a fraction of A/B traffic; adjudicated production traces become new versioned offline cases.
Follow-up probes
- Which online signals are confounded and how do you calibrate them?
- Why randomize by user rather than request?
Pass if there is a closed, privacy-reviewed loop with explicit stop rules; fail if monitoring is called evaluation without labels or action.
Q13Design a red-team program for an enterprise LLM application.
Strong answer outline
- Threat-model per input surface (user, retrieval, tools, memory, connectors) using OWASP LLM Top 10; define ASR and disclosure metrics per surface.
- Combine manual expert attacks with automated attacker-LLM generation; keep a sequestered attack pool.
- Measure effects (tool actions, disclosures), not just refusal text; every successful attack becomes a permanent regression case; enforce hard boundaries deterministically outside the model.
Follow-up probes
- Can an LLM judge grade injection outcomes safely?
- How do you test tool-result (indirect) injection?
Pass if a model mistake is contained by system controls and attacks feed the eval suite; fail if one jailbreak list is the defense.
Q14How do you choose among Ragas, promptfoo, DeepEval, Phoenix, Langfuse, and the managed cloud eval services?
Strong answer outline
- Define needs first: datasets, annotation, traces, online sampling, CI, hosting/data policy, agent metrics, collaboration.
- Map to families: metric libraries (Ragas/DeepEval), CI harness + red team (promptfoo), trace-native platforms (Phoenix/Langfuse/LangSmith), managed (Bedrock Evaluations, Vertex eval service) for native IAM and lowest setup.
- Prototype one workflow; verify judge transparency, versioning, export, and exit path; keep manifests and deterministic evaluators portable.
Follow-up probes
- When is self-hosting worth the operational cost?
- What would make you distrust a platform's built-in faithfulness metric?
Pass if comparison follows architecture and governance; fail if feature count or popularity decides.
Q15A candidate improves quality but increases cost and latency. How do you decide?
Strong answer outline
- Quantify practical quality gain and affected high-value slices with intervals.
- Compare p95/p99 and cost per successful task against product budgets and value per success.
- Consider selective routing, reranking depth, caching, or canary; state reversal/stop conditions in advance.
Follow-up probes
- What if users prefer it online despite the latency hit?
- How do you value fewer severe failures against a slower median?
Pass if the decision uses user value, risk, and Pareto trade-offs; fail if quality always wins or cheapest always wins.
Q16Your team ships prompt changes on vibes. How do you install eval-driven development?
Strong answer outline
- Start from the last incident: turn it into a reproducing case plus a neighborhood of variants, and a small smoke suite in CI within a week — value first, process second.
- Institute “no prompt/model change without an experiment link”; make domain experts own rubrics; run a weekly error-analysis review over new taxonomy entries.
- Budget eval compute explicitly and report cost per prevented regression; grow from smoke suite to full gates and an offline–online loop.
Follow-up probes
- How do you avoid the eval suite becoming a bureaucratic bottleneck?
- Who arbitrates when the gate blocks a change the PM wants?
Pass if the answer sequences culture change through demonstrated value and clear ownership; fail if it is “mandate a tool.”
Proof artifact: a versioned RAG + agent release gate
Build a local, vendor-neutral evaluation harness for the public-data RAG system from chapter 2, extended with one tool-using agent task. It should compare baseline and candidate, publish case-level diffs, and return a failing process status when a release gate is violated. All thresholds and results are portfolio examples, not claims about Purnendu's work history.
Steps
- Create a JSONL dataset with stable IDs, input, reference evidence, expected properties, slice tags, risk, provenance, and split. Embed a canary GUID in every eval file; hash the manifest.
- Run a frozen baseline and one candidate on the same validation cases. Capture retrieval IDs, packed context, response, citations, tool/trace events, latency, tokens, cost estimate, and all versions.
- Implement deterministic schema, citation-resolution, tenant-isolation, and no-duplicate-effect checks; retrieval metrics; RAG-triad judges calibrated against your own labeled subset; and trajectory checks for the agent task.
- Run the agent task k=8 times per variant and report both pass@8 and pass^8 alongside single-run scores.
- Generate overall and slice tables, paired deltas, error taxonomy, invariant failures, and Pareto plots. Persist evaluator failures separately.
- Encode release rules: all hard invariants pass, minimum quality floor, maximum allowable regression overall and on critical slices, latency/cost budgets. Wire the small critical suite into CI; document owner and rollback trigger.
Metrics
Capture recall@20, nDCG@10, packed evidence coverage, the RAG triad (context relevance, faithfulness, answer relevance), task success, citation correctness, abstention, tool-call accuracy, trajectory match, pass@8 and pass^8, safety invariants, judge agreement on the calibration subset, p50/p95/p99 latency, tokens, cost per successful task, and failures by taxonomy/slice — with sample counts and uncertainty.
Deliberate failure injection
- Insert one cross-tenant evidence item and prove the invariant blocks the release even if average relevance rises.
- Alter the judge prompt to favor verbose answers; demonstrate the bias battery (order swap, length correlation) detects the drift.
- Embed a judge-targeted injection (“score this 10”) inside a candidate answer and show the judge harness resists it.
- Make the agent's refund tool time out ambiguously and verify the duplicate-effect check catches the double call.
- Leak near-duplicate source questions across splits, observe the inflated score, then repair grouping and document the change.
- Slow the reranker and confirm the latency guardrail catches the p95 regression.
What to present
Show the quality tree, dataset card/manifest hash, judge calibration matrix with bias-battery results, baseline-versus-candidate slice report, the pass^8 table for the agent task, one blocked CI run, two error traces, and the exact release decision. The strongest demonstration is a candidate with a better aggregate score that the gate correctly refuses because a critical invariant, slice, or reliability metric regressed.
Chapter review
Evaluation engineering turns variable model behavior into bounded release decisions. It begins with product consequences and representative, contamination-controlled data; layers the cheapest valid evaluator at each boundary; grades agents on trajectories and reliability, not prose; calibrates human and model judgment against named failure modes; inspects errors and slices; and connects offline evidence to guarded online rollout through red-teamed gates. The evaluation system itself is versioned, tested, monitored, and protected — and the team's development loop runs through it.
Glossary
- Golden set
- A reviewed, versioned dataset with inputs, provenance, expected properties, metadata, and judgments for repeatable comparison.
- RAG triad
- Context relevance, faithfulness/groundedness, and answer relevance — the three scored edges of the query–context–answer triangle.
- Hard invariant
- A property that must always hold, such as tenant isolation; never averaged with softer quality metrics.
- Guardrail metric
- A limit an optimization may not violate — p95 latency, cost per success, refusal rate — often wired to automatic stop rules online.
- pass^k
- Probability that all k i.i.d. runs of the same task succeed; the reliability counterpart to capability-oriented pass@k.
- Trajectory evaluation
- Grading an agent's sequence of tool calls against reference steps (exact/in-order/any-order match, precision/recall) plus terminal state.
- Position bias
- An LLM judge's preference for a candidate based on presentation order; mitigated by order swapping and consistency filtering.
- Verbosity bias
- An LLM judge's tendency to score longer answers higher independent of quality.
- Self-preference
- A judge favoring outputs from its own model family; mitigated by cross-family judges or panels.
- Benchmark contamination
- Public test data leaking into training corpora, inflating benchmark scores relative to true task capability.
- Interleaving
- Within-session comparison of two rankers by blending their results; far more sample-efficient than between-user A/B for retrieval changes.
- Attack success rate
- Fraction of adversarial attempts per surface that achieve their objective — the core red-team regression metric.
- Paired evaluation
- Comparison of variants on the same cases, enabling per-case deltas and efficient uncertainty analysis.
- Evaluation leakage
- Test information influencing development or related examples crossing splits, inflating apparent generalization.
Mastery checklist
- I can draw a product-specific quality tree and label objectives, guardrails, and invariants.
- I can design dataset sources, provenance, grouped splits, slices, contamination controls, and a version manifest.
- I can name the RAG triad edges and route a regression on each to the right owner.
- I can evaluate an agent on terminal state, step correctness, trajectory efficiency, and pass^k — and explain why pass^k collapses.
- I can name position, verbosity, and self-preference bias with mitigations, and run a judge calibration loop with an acceptance bar.
- I can turn case failures into an actionable taxonomy and paired slice report.
- I can specify CI gates, uncertainty handling, shadow/interleave/canary with guardrail stop rules, and rollback.
- I can design a red-team program whose successful attacks become permanent regression cases.
- I can compare open-source and managed evaluation tooling on AWS and GCP while preserving portability and governance.
- I can install eval-driven development in a team, starting from one incident and one smoke suite.
Primary sources
Links checked . Evaluation tools and hosted surfaces evolve rapidly; verify versions, deprecations, data handling, and metric definitions before adoption.
- Zheng et al. — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (judge biases and agreement)
- Yao et al. — τ-bench: tool-agent-user benchmark introducing pass^k reliability
- Zhang et al. — GSM1k: careful examination of LLM grade-school math performance (contamination evidence)
- Amazon Bedrock — model and RAG evaluation jobs
- Amazon SageMaker Clarify — foundation model evaluation, bias and explainability
- Vertex AI — Gen AI evaluation service overview (pointwise, pairwise, trajectory metrics)
- Ragas — RAG and agent metric catalog
- promptfoo — config-driven evals and automated red teaming
- DeepEval — pytest-oriented component and end-to-end evaluation
- Arize Phoenix — OpenTelemetry-native tracing and evaluation
- Langfuse — traces, datasets, scores, and online evaluation
- LangSmith — datasets, experiments, and evaluator types
- TruLens — the RAG triad framing
- OpenTelemetry — semantic conventions and stability model
- NIST AI 600-1 — Generative AI Profile
- OWASP GenAI Security Project — LLM application risks
CHAPTER 07 · PRIORITY 1
Enterprise Integrations & Backend Engineering
22 min read · 12 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- turn a third-party API into an explicit, versioned contract, not a collection of happy-path calls;
- defend OAuth, service-account, and webhook admission controls for a multi-tenant connector;
- design end-to-end LLM streaming — SSE versus WebSockets, backpressure, mid-stream errors, resumable streams;
- architect an LLM gateway: routing, cross-provider fallback, token metering, budgets, safe caching;
- design token-denominated rate limits, quotas, and per-tenant cost attribution;
- run long-running agent work as durable async jobs with idempotent side effects; and
- present an HRIS connector and a gateway design as senior interview answers with failure evidence.
1. Treat every integration as a changing contract
An enterprise connector is a small distributed system at an organizational boundary: the remote team controls schema, quotas, release cadence, and incident response; your team owns the consequences. Write the contract you need first; isolate the vendor adapter behind it.
REST gives resources, caching, and operational visibility; GraphQL reduces over-fetching but still needs explicit query-cost, pagination, partial-error, and field-authorization handling. Neither removes the need for a canonical internal model: convert remote objects into stable internal types at the edge so provider renames do not spread.
| Contract concern | Decision to make | Failure if omitted |
|---|---|---|
| Identity | Which remote identifier is immutable? Is email only an attribute? | Renames create duplicates or overwrite the wrong record. |
| Pagination | Cursor, offset, or time window; stable ordering; page-size cap | Concurrent changes produce gaps or repeated pages. |
| Null versus absent | Does absence mean “unchanged,” “unknown,” or “clear the value”? | Partial updates erase valid data. |
| Versioning | URL/header version, compatibility window, schema capability | A vendor rollout breaks all tenants at once. |
| Error model | Machine-readable code, retryability, request ID, field errors | Workers retry permanent failures or drop transient ones. |
Design the public API around jobs
A sync that may take minutes should not hold an HTTP connection open. Validate authorization and idempotency, persist a job, enqueue it, return 202 Accepted with a status URL exposing queued → running → succeeded | partially_succeeded | failed | cancelled plus counts, timestamps, and a sanitized error summary. Cancellation is observed at safe checkpoints, not instant. Section 7 extends this shape to agent work.
POST /v1/tenants/{tenant_id}/sync-jobs
Idempotency-Key: 7f8f... # scoped to tenant + operation
202 Accepted
{ "job_id": "job_01...", "status": "queued", "status_url": "/v1/sync-jobs/job_01..." }
Prefer opaque cursors on a deterministic sort key; evolve with contract tests, tolerant readers for additive fields, and telemetry-backed deprecation. “We will version later” is not a strategy.
2. Identity, tenant boundaries, and webhook admission
Authentication proves who is calling; authorization decides what that identity may do; tenant routing decides whose data the action can touch. Keep all three visible.
Delegated user OAuth
When actions must reflect a human’s permissions and consent. Encrypt refresh tokens, request narrow scopes, bind the connection to a tenant, treat revoked consent as normal.
Service account
For tenant-wide unattended sync. Prefer workload identity or asymmetric client auth over long-lived shared secrets; separate credentials by environment and tenant.
OIDC login
When the app needs an authenticated user session. An ID token describes authentication; it is not an API authorization token.
The IETF baseline — authorization code with PKCE, exact redirect matching, mix-up and CSRF protection, sender-constrained tokens — is RFC 9700, not the original OAuth 2.0 RFC.
Webhook admission pipeline
flowchart LR
PR["Provider webhook"] --> SZ["Raw-byte size and content-type limit"]
SZ --> SG["Signature and freshness check"]
SG --> TN["Tenant lookup by endpoint secret"]
TN --> IB["Durable inbox insert (unique delivery id)"]
IB --> AK["Fast 2xx acknowledgement"]
IB --> WR["Async worker"]
WR --> DE["Domain event via outbox"]
- Read the exact raw bytes under strict size and type limits; never parse and re-serialize first.
- Select the secret by authenticated endpoint or connection identifier — never by a tenant ID from the payload.
- Verify the signature in constant time (GitHub’s guidance is canonical).
- Enforce a bounded age on signed timestamps and record delivery IDs — signature validity does not prove freshness.
- Insert metadata and payload hash into an inbox with a unique key, acknowledge fast, process asynchronously.
Accept both secrets during a short audited rotation overlap; record the key version per delivery. Never log tokens, secrets, or unredacted payloads: a correlation ID is useful, a credential is not.
3. Delivery semantics are end-to-end properties
Brokers describe transport behavior; the business outcome also depends on producers, consumers, databases, and external side effects. Say exactly where duplication or loss can occur.
- At-most-once: no retry after uncertainty; work may be lost, duplicates avoided. Only when loss is cheaper than duplication.
- At-least-once: retry until acknowledged; handlers must make duplicates harmless.
- Effectively once: idempotency, uniqueness, and atomic transitions make the result occur once within a defined boundary.
Kafka supports idempotent production and transactions within its own model (delivery-semantics docs), but do not stretch that across an arbitrary email, payroll API, and database. Define the transaction boundary and the compensation path.
Inbox, outbox, and idempotent effects
BEGIN;
INSERT INTO processed_event(tenant_id, event_id)
VALUES (:tenant, :event) ON CONFLICT DO NOTHING; -- duplicate becomes a no-op
-- Continue only if one row was inserted.
UPSERT employee ...;
INSERT INTO outbox(event_id, aggregate_id, event_type, payload) ...;
COMMIT;
The inbox stops repeated consumption from re-applying a transition. The outbox stores a domain event in the same transaction as the domain change; a relay publishes committed rows. That closes the dual-write gap but not duplicate publication — consumers still deduplicate (AWS transactional outbox guide). For non-idempotent external effects, use a provider idempotency key or a local operation record, and resolve uncertain outcomes by querying the provider before retrying. A client timeout never proves failure.
Ordering, backpressure, and dead letters
Global ordering is expensive and rarely required. Partition by the smallest aggregate needing order — often tenant_id + employee_id — with an aggregate version so consumers reject stale updates. Backpressure is a correctness control: bound worker concurrency, honor rate-limit hints, back off exponentially with full jitter (AWS Builders’ Library), and stop admitting backfills before real-time changes starve. A dead-letter queue is quarantine, not a cemetery: keep error class, attempt history, and a redacted payload reference; replay revalidates authorization and schema, rate-limits release, and audits itself.
4. Streaming LLM responses end-to-end
GenAI adds a transport problem the classic connector never had: a response that streams for tens of seconds, its perceived quality dominated by time-to-first-token (Chapter 2). The job is moving that stream across every hop unbuffered — and defining what the client sees when something dies at token 500.
| Dimension | Server-sent events (SSE) | WebSockets |
|---|---|---|
| Direction | Server → client, plain HTTP | Full duplex |
| Infrastructure | Normal HTTP through L7 load balancers; just disable buffering | Upgrade handshake; every proxy and WAF must support it |
| Reconnection | Built in: auto-retry plus a Last-Event-ID cursor (WHATWG spec) | You design your own resume protocol |
| Best fit | One-way token streams — chat and copilots | Voice, collaboration, client events mid-stream |
Choose SSE unless the client must talk during the stream; “stop generation” works as a POST to a cancel endpoint. Disable proxy buffering and compression on the route, flush per event, heartbeat every 10–15 seconds so idle timeouts (60 seconds default on an ALB) do not sever quiet streams; on disconnect, cancel the provider stream — you pay for every generated token, read or not.
sequenceDiagram
participant C as "Client"
participant E as "Edge and load balancer"
participant S as "App service"
participant P as "Model provider"
C->>E: POST chat request with stream enabled
E->>S: forward with deadline
S->>P: open provider token stream
P-->>S: token deltas
S-->>E: SSE events with event ids
E-->>C: flushed chunks per event
P-->>S: provider error mid stream
S-->>C: typed terminal error event then clean close
- Mid-stream errors — the 200 left at token one. Emit typed in-band events (
delta,error,done) with a mandatory terminal event; an abrupt close without one is an uncertain outcome, not success. - Backpressure — a slow client cannot pause Bedrock or Vertex. Await socket writes, bound the per-connection buffer, coalesce deltas or cancel at the bound.
- Resumable streams — persist deltas to a durable log keyed by response ID, batched every 50–100 ms; reconnects with
Last-Event-IDreplay from the offset, then continue live. - Usage finality — record provider-reported token usage exactly once at finalization, even if the client vanished.
5. The LLM gateway: one front door for every model provider
Once several teams call several models, put a gateway between products and providers — the pattern popularized by LiteLLM-style proxies. Applications speak one internal dialect and reference model aliases; the gateway owns everything a provider swap should not break.
flowchart LR
APP["Product services"] --> GW["LLM gateway"]
GW --> ADM["Virtual key auth + budget check"]
ADM --> CA["Exact and semantic cache"]
CA -->|"hit"| RES["Serve result"]
CA -->|"miss"| RTR["Router with model aliases"]
RTR -->|"primary"| BR["Bedrock"]
RTR -->|"fallback"| VX["Vertex AI"]
RTR -->|"fallback"| OA["OpenAI-compatible endpoint"]
BR --> MET["Token metering + usage events"]
VX --> MET
OA --> MET
MET --> RES
Routing is configuration: aliases like fast and reasoning map to provider, model version, and region (data-residency pins live here), with canaries gated by the Chapter 6 evaluation harness. Retries stay in-provider — jittered backoff on 429/5xx inside a latency budget, hedged on time-to-first-token. Fallback crosses providers and is a product decision: tokenizers, context limits, tool schemas, and behavior differ, so only eval-approved pairs enter a chain, responses carry the served model, and a circuit breaker keyed on provider + region + model trips it automatically.
Metering records provider-reported tokens — never your own estimate — with tenant, principal, feature, model, cache flag, latency, and priced cost. Caching starts exact-match: tenant scope, model version, system-prompt hash, normalized parameters, prompt. Semantic caching (embedding similarity over a threshold) multiplies hit rate but risks stale answers and cross-tenant leakage: scope by tenant and ACL hash, exclude personalized or tool-using calls, eval-check served quality. Keep the gateway boring — stateless, horizontally scaled, counters in Redis — it sits on every request path. Bedrock’s cross-region inference and intelligent prompt routing cover slices natively (Bedrock docs); a gateway earns its place once you span providers or need tenant policy.
6. Rate limiting and quotas for token-denominated APIs
Request-per-minute limits fail because one LLM request can cost three orders of magnitude more than another — a 300-token lookup versus a 150k-token document analysis (example figures). Providers meter requests and tokens separately (OpenAI’s rate-limit guide; Bedrock and Vertex quotas are analogous); your platform must do the same per tenant.
The pattern is estimate–debit–reconcile: count prompt tokens exactly, estimate completion from max_tokens or a per-feature historical p95, debit the tenant’s bucket at admission, reconcile against provider-reported usage after. Unbounded max_tokens means unbounded debit — force a cap. Fairness matters more than raw limiting: per-tenant buckets draw weighted shares of provider quota, headroom is reserved for interactive traffic over batch, and a noisy tenant is throttled before the provider throttles everyone.
When rejecting, behave like a good provider: 429 with Retry-After, distinguish exhausted tenant quota from degraded platform capacity, and offer degrade paths — a cheaper alias, truncated context, or conversion to an async job.
7. Async jobs for long-running agent work
Agent runs (Chapter 5 owns the loop) are minutes long and unpredictable — the Section 1 job pattern with more states. Intake validates, persists, enqueues, returns 202; workers lease with a visibility timeout beyond the worst step or heartbeat to extend; each step checkpoints so a crash resumes rather than restarts; cancellation is observed between steps.
stateDiagram-v2
[*] --> queued
queued --> running : worker lease
running --> awaiting_approval : gated tool call
awaiting_approval --> running : human approves
awaiting_approval --> cancelled : rejected or expired
running --> succeeded
running --> partially_succeeded : some steps failed
running --> failed : retries exhausted
queued --> cancelled : cancel requested
running --> cancelled : observed at checkpoint
succeeded --> [*]
partially_succeeded --> [*]
failed --> [*]
cancelled --> [*]
Deliver status per consumer: polling with backoff and ETags as the baseline; outbound webhooks — you are now the Section 2 provider, owing signed payloads, retries, a DLQ, and redelivery; an SSE status stream for interactive UIs, reusing Section 4’s machinery. For approvals, waits, and compensation, use a durable orchestrator: Step Functions callback task tokens and Workflows callbacks model waiting for a human without a worker burning a lease.
AWS
- API Gatewayfront door, usage plans, WebSocket APIs
- Lambdaintake and workers; response streaming for SSE
- SQSwork queue: visibility timeout, redrive, DLQ
- Step Functionsdurable orchestration, callback task tokens
- EventBridgecompletion fan-out; API destinations for webhooks out
Google Cloud
- Apigeegateway with quota and spike-arrest policies
- Cloud Runstreaming services and job workers
- Pub/Subwork distribution, push or pull, dead-letter topics
- Cloud Tasksrate-controlled dispatch, per-queue throttles
- Workflowsdurable orchestration with callbacks
SQS or Pub/Sub for single-step fan-out; Step Functions or Workflows for branches, waits, and compensation; Cloud Tasks for per-queue rate control toward fragile downstreams — Section 3’s admission control, applied outbound.
8. Idempotent AI actions and multi-tenant cost attribution
An agent that sends emails, files tickets, or issues refunds makes Section 3’s idempotency adversarial: the model may phrase the same intent with different text on every retry, so keys must derive from the action, not the words — hash(tenant, job, step, tool, canonical_args).
The effect journal records intent → executing → done-with-result. On any retry — including an LLM retry replaying a step whose tool already ran — the journal answers instead of the tool: the agent sees the recorded result rather than sending a second email. Uncertain outcomes resolve by querying the provider or via provider idempotency keys (Stripe’s design); unrepeatable actions get reserve/confirm phases or a human gate from Figure 4. Classify tools by side-effect risk at registration, not in the prompt.
flowchart LR
UE["Usage event per call"] --> MT["Metering stream"]
MT --> EN["Enrich with versioned price sheet"]
EN --> AG["Hourly rollup per tenant and feature"]
AG --> BE["Budget engine"]
BE -->|"soft limit"| DG["Alert and degrade to cheaper alias"]
BE -->|"hard limit"| HL["Reject new work with 429"]
AG --> SB["Showback dashboards and invoice export"]
Attribution requires every AI-adjacent call to emit a usage event — model calls, embeddings, vector queries, GPU seconds — priced at enrichment from a versioned price sheet so reports survive price changes. Soft budgets alert and degrade the alias; hard budgets reject new work while in-flight streams finish. Reconcile with the invoice monthly: Bedrock application inference profiles and cost-allocation tags on AWS; labels plus the BigQuery billing export on Google Cloud. Chapter 10 covers wider FinOps.
9. Worked system: a resilient HRIS-to-AI knowledge connector
A hypothetical interview-practice system, not a claim about Purnendu’s experience. A customer wants employee directory and policy documents synchronized from an HRIS into an access-controlled AI assistant. Minutes of staleness are acceptable; cross-tenant or cross-group exposure is not.
flowchart TD
H["HRIS provider"] -->|"OAuth + webhooks"| API["Connector API"]
API --> IB["Inbox"]
IB --> Q["Queue"]
Q --> WK["Bounded workers"]
WK -->|"paginated reads"| H
WK --> PG["Canonical PostgreSQL"]
PG --> OB["Outbox relay"]
OB --> IX["ACL-aware indexer"]
SCH["Scheduler"] --> SY["Incremental sync"]
SY --> WK
SY --> RC["Reconciliation + audit report"]
Key records: connection (tenant, provider, scopes, encrypted credential reference, health), sync_job (cursor, high-water mark, counts), inbox_delivery, source_object (remote ID, version, payload hash, tombstone), employee, outbox_event, sync_anomaly. Tenant-owned keys begin with tenant_id; repositories require tenant context. Sync captures a start high-water mark, pages deterministically, and advances the checkpoint only after durable apply. Webhooks are latency hints; reconciliation compares IDs, counts, and sampled hashes to catch missed webhooks, permission loss, and drift. Deletions become tombstones and removal events.
| Failure | Immediate behavior | Repair |
|---|---|---|
| Access token expires | Single-flight refresh; pause that connection only | Rotate and retry within deadline; terminal failure marks reauthorization_required |
| 429 or provider outage | Honor retry guidance, jittered backoff, open circuit | Resume from committed cursor; surface lag per tenant |
| Worker dies after commit | Broker redelivers | Inbox/unique keys make the repeated page harmless |
| Schema adds an enum value | Preserve raw value; map to unknown; emit anomaly | Update adapter, replay quarantined records |
| Index write uncertain | Do not complete the outbox event | Retry with deterministic document ID; reconcile DB against index |
| Tenant disconnects | Revoke credentials, stop new work | Cancel at checkpoints; run retention/deletion workflow with evidence |
Python backend judgment
Use async def only for genuinely asynchronous I/O; offload CPU-heavy parsing and blocking SDKs off the event loop (FastAPI’s async guidance). Bound concurrency with a semaphore, pass deadlines through, close clients in lifespan hooks, treat cancellation as control flow.
async with asyncio.timeout(job.remaining_seconds()):
async with provider_slots: # protects provider and this process
page = await client.list_people(cursor=job.cursor)
await repository.apply_page_atomically(page, job_id=job.id)
await repository.commit_checkpoint(job.id, page.next_cursor)
# Never catch BaseException and swallow CancelledError.
Validate payloads into strict adapter models (Pydantic), then map to domain types. Unit-test mappings; integration-test transactions and redelivery; contract-test fixtures; end-to-end test a fake provider injecting 429s, timeouts, malformed pages, cursor loops, and duplicate webhooks.
10. Backend breadth checkpoint
Senior interviews probe whether depth in one stack generalizes. The bar: production-strong Python plus the ability to read, implement, and critically review one secondary stack.
- API surface — REST and GraphQL both need authentication, tenant scoping, pagination, stable error contracts, abuse protection; version additively with deprecation telemetry.
- Queues are products — Kafka: partitioned log with replay and offsets. SQS: visibility timeouts, standard/FIFO. Pub/Sub: acknowledged delivery, subscription retention. State ordering, duplication, retention, and DLQ policy per product — never “a message broker.”
- Python service depth — dependency injection for principals, tenant context, sessions; durable work in queues, not fire-and-forget tasks; tests for duplicates, timeouts, cancellation.
- Secondary stack — Node/TypeScript: event loop, promise cancellation, runtime validation. Go: contexts, channels, resource ownership. Same semantics, different spelling.
Interview playbook
Lead with the business invariant, then trace one request and one failure. A compact structure is BOUNDARY:
- B — Business truth: source of truth, freshness, deletion, conflicts, success metric.
- O — Ownership and identity: tenant, principal, scopes, data classification, audit actor.
- U — Uncertainty: timeouts, duplicates, partial failure, ordering, unknown outcomes.
- N — Normalized contract: canonical model, API/job state, event envelope, versioning.
- D — Durability: inbox/outbox, checkpoints, idempotency boundary, effect journal, reconciliation.
- A — Admission control: token-aware limits, budgets, backpressure, deadlines, circuit breakers.
- R — Recovery and rollout: replay, DLQ, resumable streams, provider fallback, canary tenants.
- Y — Yardsticks: sync lag, completion rate, duplicates suppressed, cost per tenant, tokens per feature.
Classic traps: “exactly once” without a boundary, email as immutable identity, tenant ID from an unsigned payload, retrying every error, DLQ as recovery. GenAI traps: request-count limits on a token API, SSE through a buffering proxy, no terminal-event protocol, a semantic cache keyed without tenant or model version, silent cross-provider fallback inside an agent loop, tools with no idempotency story. When coding, narrate cancellation, transaction scope, and how a test proves duplicate safety.
Question bank
Practise aloud. Each answer should state assumptions and defend one concrete boundary.
Q1How do you prevent a duplicated webhook from creating two employee records?
Strong answer outline
- Durably insert the provider delivery ID under a tenant-scoped unique constraint.
- Map by immutable provider object ID; upsert with source version or payload hash.
- Commit inbox status, domain change, and outbox event atomically.
Follow-up probes
- No delivery ID from the provider?
- First request timed out after commit?
You covered transport deduplication, business idempotency, and the uncertain-outcome case.
Q2Design the OAuth lifecycle for a tenant-wide HRIS connection.
Strong answer outline
- Authorization code with PKCE, exact redirects, state and issuer validation, narrow scopes.
- Bind connection to tenant and installer; encrypt tokens, record scope and version, serialize refresh.
- Handle revocation, re-consent, rotation, and offboarding without leaking credentials.
Follow-up probes
- When is a service account preferable?
- How do two workers avoid refresh races?
You covered grant, storage, refresh, loss of access, and tenant binding.
Q3When is a retry unsafe?
Strong answer outline
- Validation and auth errors are terminal; throttling and transient failures may retry.
- After a timed-out non-idempotent write, query by idempotency key before writing again.
- Deadline, capped attempts, full jitter, shared retry budget.
Follow-up probes
- Why can retries amplify an outage?
- Where does
Retry-Afterinfluence scheduling?
You treated timeout as uncertainty and bounded retries across layers.
Q4Where should tenant isolation be enforced in a connector?
Strong answer outline
- Derive tenant context from the authenticated connection, never solely the payload.
- Carry tenant through queue envelope, repository API, keys, caches, metrics, audit.
- Add database policy constraints and adversarial cross-tenant tests.
Follow-up probes
- Risks in a shared worker cache?
- How does a dedicated-tenant deployment change this?
You enforced at every hop and named a test that attempts a leak.
Q5What does cancellation mean for an async job?
Strong answer outline
- Persist
cancel_requested; workers observe it before pages or side effects. - Propagate cancellation, close clients, release leases, keep the last checkpoint.
- Expose
cancelledonly after cleanup; compensate anything in flight.
Follow-up probes
- Why not kill the worker process?
- Cancellation during a database commit?
You distinguished request, observation, atomic boundaries, and terminal state.
Q6SSE or WebSockets for LLM tokens — and what happens when the provider dies at token 500?
Strong answer outline
- SSE by default: plain HTTP, built-in reconnect with
Last-Event-ID; WebSockets only when the client must talk mid-stream. - Hop discipline: no buffering, per-event flush, heartbeats, bounded buffers, cancel the provider stream on disconnect.
- The 200 is committed: typed in-band events with a mandatory terminal event, a durable delta log for replay-then-resume, usage recorded once at finalization.
Follow-up probes
- Abrupt close with no terminal event?
- Cost of the delta log per token?
You named the committed-status trap and a concrete resume mechanism.
Q7Your platform calls Bedrock, Vertex, and OpenAI. Design retry and fallback at the gateway.
Strong answer outline
- 429/5xx retry in-provider with jittered backoff, hedged on time-to-first-token; invalid or filtered requests never retry.
- Fallback only along eval-approved pairs — tokenizers, context limits, tool schemas differ; record the served model.
- Circuit-break per provider + region + model; cap fallback spend; log routing decisions.
Follow-up probes
- Why is silent fallback dangerous for a tool-calling agent?
- How do you test the chain before an outage?
You treated fallback as an eval-gated product decision, not an availability trick.
Q8Why do request-per-minute limits fail for LLM APIs, and what replaces them?
Strong answer outline
- Cost scales with tokens; meter requests, tokens, and concurrent streams independently.
- Estimate–debit–reconcile: debit prompt plus estimated completion tokens; settle with provider-reported usage.
- Per-tenant buckets with weighted fair shares; 429 with
Retry-After; degrade paths.
Follow-up probes
max_tokensunset or enormous?- Gateway enforcement or provider limits — why both?
You named both meters, reconciliation, and tenant fairness.
Q9An agent sends emails and creates tickets. A step times out and retries. How is the email sent once?
Strong answer outline
- Key from tenant + job + step + canonical args — never model text; journal intent before executing.
- On retry the journal answers: done replays the recorded result; uncertain outcomes resolve via the provider before re-issuing.
- Classify tools by side-effect risk; unrepeatable actions get reserve/confirm or human gates.
Follow-up probes
- Model re-plans with slightly different arguments?
- Journal retention and concurrency semantics?
You separated LLM retries from tool retries and defined journal replay.
Q10Design multi-tenant cost attribution for a GenAI platform.
Strong answer outline
- Usage event per model, embedding, and vector call — tenant, feature, model, provider-reported tokens — priced at enrichment from a versioned price sheet.
- Aggregate to per-tenant showback; soft budgets degrade, hard budgets reject, at the gateway.
- Reconcile with the invoice: Bedrock inference profiles and tags; GCP labels plus BigQuery billing export.
Follow-up probes
- Charge cache-served responses?
- What breaks if prices are looked up at report time?
You covered capture, price versioning, enforcement, and invoice reconciliation.
Q11Expose a five-minute agent run through your public API: polling, webhooks, or streaming?
Strong answer outline
202plus a job resource with an explicit state machine; polling with backoff and ETags as baseline.- Outbound webhooks make you the provider: signed payloads, retries, DLQ, redelivery; SSE status for UIs.
- Durable orchestration: heartbeat leases, per-step checkpoints, safe-point cancellation, approval callback tokens.
Follow-up probes
- How do consumers deduplicate your webhooks?
- When is a synchronous variant acceptable?
You chose per consumer type and carried admission duties outbound.
Q12When is a semantic cache safe in front of an LLM, and how would you build the key?
Strong answer outline
- Exact-match first: tenant scope + model version + system-prompt hash + normalized parameters + prompt.
- Semantic lookup risks stale answers and cross-ACL leakage: scope by tenant and ACL hash; exclude personalized or tool-using calls.
- Eval-check cache-served quality; invalidate on model or system-prompt change.
Follow-up probes
- Cache streaming responses?
- What hit rate justifies the infrastructure?
Your key included tenant, model version, and ACL scope — risk framed before savings.
Proof artifact: LLM gateway and resilient connector laboratory
Build a small gateway-plus-connector system against fake providers. All targets are example acceptance thresholds, not claims of prior results.
- Gateway coreFastAPI proxy over two fake providers with different dialects: virtual keys, model aliases, per-tenant estimate–debit–reconcile buckets, usage events to a metering table.
- Streaming pathSSE with heartbeats, typed delta/error/done events, Redis delta log with
Last-Event-IDresume. Kill a provider mid-stream; demonstrate recovery. - Fallback and cacheJittered in-provider retries, an eval-gated cross-provider fallback pair with a circuit breaker, an exact-match cache keyed by tenant + model version + prompt hash.
- Async agent job202-based job API with the Figure 4 state machine, signed outbound webhooks with retries and DLQ, an idempotent tool-effect journal.
- Connector spineWebhook inbox, checkpointed sync against a fake HRIS injecting 429s, duplicates, and cursor loops, plus outbox-driven indexing and reconciliation.
- Cost reportPer-tenant, per-feature showback; a budget breach that degrades the alias, then hard-rejects.
Measure: p50/p95 time-to-first-token through the gateway, stream resume success rate, fallback activations and quality delta, token-estimate error after reconciliation, duplicate effects across 1,000 replayed webhooks and tool retries (lab target: zero), reconciliation drift, cost-report accuracy.
Inject failures: provider dies at token 500; slow client stalls a stream; both providers 429 at once; worker crashes between tool execution and journal write; a webhook replayed 100 times; budget exhausted mid-stream; unknown enum from the HRIS.
Present: a two-minute walkthrough, one client-to-provider trace, the resume demo, before/after reconciliation and cost reports, and a decision record covering SSE versus WebSockets, the fallback gate, and one rejected alternative.
Chapter review
A production integration assumes change, duplication, delay, partial failure, and revoked authority; a GenAI backend adds streams that outlive their status code, costs denominated in tokens, and agents whose retries can repeat side effects. The core is unchanged: explicit contracts, durable state, idempotent effects, bounded and metered work, auditable tenant-aware recovery.
Glossary
- Inbox / outbox
- Durable tables that deduplicate received messages and atomically stage messages to publish.
- High-water mark
- Durable boundary showing how far an incremental process has safely advanced.
- Reconciliation
- Comparing source and destination truth to repair drift event delivery missed.
- Last-Event-ID
- SSE’s built-in resume cursor: the client replays from the last event it saw.
- Terminal event
- Mandatory in-band end-of-stream marker; its absence signals an uncertain outcome.
- Model alias
- Indirection between a capability tier and a concrete provider model version.
- Estimate–debit–reconcile
- Token-bucket lifecycle: debit an estimate at admission, settle with provider-reported usage.
- Effect journal
- Durable record of tool-call intent and result that answers retries instead of re-executing.
- Showback
- Per-tenant, per-feature cost reporting from priced usage events.
- Retry budget
- Bound on extra attempts so recovery traffic cannot overwhelm a degraded dependency.
Mastery checklist
- I can define an idempotency key’s scope, retention, and concurrency — for API calls and agent tool effects.
- I can trace tenant identity from ingress through queue, database, cache, usage event, and billing tag.
- I can explain why outbox publication and end-to-end exactly-once are different claims.
- I can design an SSE path with heartbeats, in-band terminal events, and Last-Event-ID resume.
- I can defend an eval-gated cross-provider fallback chain and say where it must never trigger.
- I can design token-denominated rate limits with estimation, reconciliation, and tenant fairness.
- I can produce a per-tenant cost report that survives a price change and reconciles against the invoice.
Primary sources
Links checked 2026-08-04.
- IETF RFC 9700 — OAuth 2.0 Security Best Current Practice
- WHATWG HTML Standard — Server-sent events
- GitHub Docs — Validating webhook deliveries
- AWS Prescriptive Guidance — Transactional outbox
- AWS Builders’ Library — Timeouts, retries, backoff with jitter
- AWS Lambda — Response streaming
- Amazon API Gateway — WebSocket APIs
- Amazon Bedrock documentation
- AWS Step Functions documentation
- Google Cloud Run — WebSockets
- Generative AI on Vertex AI documentation
- Google Cloud Workflows and Cloud Tasks documentation
- Google Cloud Billing — BigQuery export
- LiteLLM documentation
- OpenAI — Rate limits guide
- Stripe API — Idempotent requests
- Apache Kafka — Delivery semantics
- FastAPI — Concurrency and async/await
CHAPTER 08 · CLOUD TRACK
The AWS GenAI Stack: Bedrock, SageMaker & Serverless AI
31 min read · 14 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- choose and defend a Bedrock inference mode — on-demand, cross-region inference profiles, provisioned throughput, or batch — with tokens-per-minute math rather than vibes;
- design an enterprise RAG system on AWS, including a justified vector-store choice among OpenSearch Serverless, Aurora pgvector, and Kendra GenAI index;
- compare Bedrock Agents, Bedrock AgentCore, and hand-rolled Step Functions orchestration for agentic workloads, and know when each is the wrong answer;
- layer Guardrails, IAM, PrivateLink, KMS, CloudTrail, and invocation logging into a security story a regulated-industry interviewer will accept;
- argue the Bedrock-managed versus SageMaker versus self-hosted-on-EKS decision with utilization and unit-economics reasoning;
- engineer cost deliberately: pricing dimensions, prompt caching, batch discounts, provisioned-throughput sizing, and per-tenant cost allocation; and
- answer the AWS solution-design scenarios that actually appear in senior GenAI and forward-deployed interviews.
1. The AWS GenAI stack as one mental model
Interviewers do not reward service-name recitation; they reward a coherent layering in which every AWS service has a job and an alternative. Hold this model: Amazon Bedrock is the managed model-consumption plane (inference, RAG, agents, safety, evaluation), SageMaker AI is the model-production plane (training, tuning, self-managed serving), and the serverless suite (Lambda, Step Functions, API Gateway, EventBridge, SQS) is the application plane that turns model calls into products. Security, observability, and FinOps primitives cut across all three.
Keep the GCP mirror in your head for portability questions — chapter 09 goes deep on the other side, so here you only need the mapping:
AWS
- Bedrockmanaged multi-vendor FM inference
- Bedrock Knowledge Basesmanaged RAG ingestion + retrieval
- Bedrock AgentCoreagent runtime, tools, memory, identity
- SageMaker AI + HyperPodtraining, tuning, self-managed serving
- OpenSearch Serverlessvector + hybrid search engine
- Lambda + Step Functionsapp logic and orchestration
Google Cloud
- Vertex AI Model Garden + Gemini APImanaged FM inference
- Vertex AI Search / RAG Enginemanaged grounding + retrieval
- Vertex AI Agent Enginemanaged agent runtime
- Vertex AI training + custom servingtuning and self-managed serving
- Vertex AI Vector Search / AlloyDBANN and pgvector-style retrieval
- Cloud Run + Workflowsapp logic and orchestration
2. Bedrock inference: Converse, throughput modes, and cross-region profiles
Bedrock exposes a per-account, per-region model catalog (Anthropic, Amazon Nova, Meta, Mistral, Cohere, DeepSeek, and others) with access enabled per model. The modern integration surface is the Converse and ConverseStream APIs: one request/response shape across vendors, with system prompts, multimodal content blocks, tool use, and usage metadata. Prefer Converse over the legacy per-model InvokeModel payloads in any design answer — it is what makes model routing and A/B swaps cheap, and it is the seam where an LLM gateway (chapter 07) plugs in.
The decision interviewers actually probe is how you buy tokens. Bedrock has three purchase modes plus a routing layer:
| Mode | What you get | What breaks if misused |
|---|---|---|
| On-demand | Per-token pricing, shared regional capacity, account-level quotas in requests/min and tokens/min | Throttling under bursts; no throughput guarantee for launch spikes |
| Cross-region inference profiles | One profile ID routes across a geography's regions for higher effective throughput and burst resilience (docs); required invocation path for many newer models | Data is processed anywhere inside the geography — you must clear that with compliance, not assume single-region processing |
| Provisioned throughput | Dedicated model units with committed tokens/min, no-commit hourly or 1/6-month terms (docs); required to serve most customized models | Paying for idle units; sizing from guesses instead of measured token telemetry |
| Batch inference | Async jobs over JSONL in S3 at roughly half the on-demand token price (docs) | Using it for anything latency-coupled; no SLA on completion time |
flowchart TD
A["New Bedrock inference workload"] --> B{"Interactive user traffic?"}
B -->|"no, latency-tolerant"| C["Batch inference: JSONL in S3, discounted tokens"]
B -->|"yes"| D{"Steady, predictable token volume?"}
D -->|"spiky or unknown"| E["On-demand via cross-region inference profile"]
D -->|"high and steady"| F["Provisioned throughput model units"]
E --> G["Engineer for throttles: retries with jitter, queue overflow"]
F --> H["Size units from measured tokens per minute, then commit"]
C --> J["Outputs feed evals and downstream stores"]
Two multipliers change the math on top of any mode. Prompt caching lets you mark cache checkpoints so a stable prefix (system prompt, tool schemas, long documents) is billed at a steeply discounted cache-read rate — on supported models the read discount is on the order of 90% versus fresh input tokens, and time-to-first-token drops because prefill is skipped. Intelligent prompt routing can send easy requests to a cheaper model within a family. Both are useless if your prompt layout churns the prefix on every call — put volatile content (user turn, retrieved chunks that change) after the stable blocks.
3. Knowledge Bases and the vector-store decision
Bedrock Knowledge Bases is managed RAG plumbing: connectors (S3, SharePoint, Confluence, Salesforce, web crawler), parsing (default text extraction or foundation-model parsing for tables and figures), chunking (fixed-size, hierarchical parent-child, semantic, none, or a custom Lambda transform), embedding, and writes into a vector store you choose. At query time you call Retrieve for chunks-plus-scores or RetrieveAndGenerate for a fully managed answer with citations. Chapter 04 covers retrieval science — chunking trade-offs, hybrid search, reranking — so here focus on the AWS-specific decisions: which store, which parsing mode, and when to abandon the managed path.
| Store | Strengths | Costs and caveats | Pick when |
|---|---|---|---|
| OpenSearch Serverless | Default KB choice; hybrid (BM25 + k-NN) search; scales to large corpora; no cluster ops | OCU-based billing has an always-on floor — idle dev collections still cost real money each month; capacity units are the tuning knob | Production RAG at meaningful scale, hybrid retrieval, teams without search-ops appetite |
| Aurora PostgreSQL + pgvector (Aurora docs) | Vectors co-located with relational data; SQL joins for metadata filtering; familiar ops; Serverless v2 scales down low | You own index choice (HNSW), tuning, and connection management; ANN at very large scale needs care | Corpus lives next to transactional data; strict metadata/ACL filtering via SQL; cost-sensitive small-to-mid scale |
| Kendra GenAI index | Managed semantic ranking with enterprise connectors and ACL-aware results; doubles as classic enterprise search; no embedding management | Highest per-unit cost of the three; less control over retrieval internals | Enterprise search + RAG on the same index; heavy connector/permission requirements |
| S3 Vectors | Vector storage in S3 at object-storage economics for massive, colder corpora | Newer service tier — validate latency and feature fit before defaulting to it | Very large archives where per-query latency tolerance is generous |
flowchart LR
subgraph ING["Ingestion path"]
S3D["S3 document bucket"] --> KB["Knowledge Base: parse, chunk, embed"]
KB --> VDB["OpenSearch Serverless vector index"]
end
subgraph SRV["Serving path"]
CL["Client app"] --> GW["API Gateway + Lambda"]
GW --> ORC["Orchestrator: Retrieve then Converse"]
ORC --> VDB
ORC --> FM["Bedrock model (Converse API)"]
FM --> GRD["Guardrail: grounding check + PII mask"]
GRD --> GW
end
FM --> LOGS["Invocation logging to S3 + CloudWatch"]
The senior move is knowing when to leave RetrieveAndGenerate: keep it for internal tools and fast pilots; switch to Retrieve plus your own Converse call the moment you need custom reranking, multi-index federation, query rewriting, or response contracts the managed generator cannot express. That split — managed ingestion, custom generation — is the most common production posture and cites well in interviews (the underlying pattern is the original RAG formulation, Lewis et al. 2020).
4. Bedrock Agents and AgentCore: the agent platform decision
AWS now has two agent stories, and interviewers check whether you know the difference. Bedrock Agents (the original) is a Bedrock-native, prompt-template-driven orchestrator: you define action groups (function or OpenAPI schemas backed by Lambda), attach knowledge bases and guardrails, optionally enable code interpretation and memory, and Bedrock runs the reason-act loop — including return-of-control when you want the client to execute an action. It is fast to stand up and tightly coupled to Bedrock models and its own orchestration style.
Bedrock AgentCore is the 2025-generation answer to a different question: "I already built my agent in LangGraph/Strands/CrewAI with whatever model I want — now give me production infrastructure." It is a set of composable, framework-agnostic services rather than an orchestrator:
- Runtime — serverless execution with per-session isolation and long-running sessions (hours, not API-gateway seconds), so tool-using agents don't inherit request/response time limits.
- Gateway — turns existing APIs and Lambda functions into MCP-compatible tools with auth handled, instead of hand-writing tool adapters per agent.
- Memory — managed short-term session memory and long-term extracted memory shared across sessions.
- Identity — inbound caller auth plus outbound OAuth to third-party services, so agents act with scoped, auditable credentials rather than a god-mode service account.
- Built-in tools — managed Code Interpreter and Browser sandboxes, isolated per session.
- Observability — OpenTelemetry traces of every step into CloudWatch, which is what makes agent debugging tractable.
flowchart TD
U["Caller: app or workflow"] --> IDN["AgentCore Identity: inbound auth"]
IDN --> RT["AgentCore Runtime: isolated long-running session"]
RT --> MEM["AgentCore Memory: session + long-term"]
RT --> GWY["AgentCore Gateway: APIs and Lambda as MCP tools"]
GWY --> T1["Lambda tool: order lookup"]
GWY --> T2["Internal REST API"]
RT --> CIN["Code Interpreter sandbox"]
RT --> BRW["Browser tool"]
RT --> FM["Bedrock models via Converse"]
RT --> OBS["OTEL traces to CloudWatch"]
The third option is no agent service at all: a Step Functions state machine calling Bedrock at each step. That is the right answer more often than vendors admit — when the "agent" is really a deterministic workflow with one or two LLM steps, a state machine gives you replayability, per-step retries, and an audit trail for free, with none of the autonomy risk. Chapter 05 covers agent design patterns themselves; your AWS-specific claim is the placement decision: deterministic flow → Step Functions; dynamic tool choice with production infra needs → AgentCore; Bedrock-native quick build → Bedrock Agents.
5. Guardrails, model customization, and Bedrock evaluations
Bedrock Guardrails is a policy layer evaluated on input and/or output, attachable to direct invocations, agents, and knowledge bases — or callable standalone via the ApplyGuardrail API against any model, including ones outside Bedrock. Know the policy types and, more importantly, what each does and does not catch:
| Policy | Mechanism | Honest limitation |
|---|---|---|
| Content filters | Classifier tiers for hate, insults, sexual, violence, misconduct, prompt-attack | Statistical — tune thresholds against your own red-team set, not defaults |
| Denied topics | Natural-language topic definitions blocked on input/output | Paraphrase-sensitive; needs eval coverage, not one-line definitions |
| Sensitive information | PII entity detection with mask or block, plus custom regexes | Masking output PII does not fix a retrieval layer that leaked the document |
| Contextual grounding checks | Scores grounding (is the answer supported by source?) and relevance against supplied context | Threshold-based hallucination screen, not proof of correctness; complements — not replaces — chapter 06 evals |
Model customization on Bedrock (docs) spans fine-tuning on labeled pairs, distillation (a larger teacher generates training data to specialize a cheaper student), and custom model import for open weights you tuned elsewhere. The operational fact candidates miss: most customized models must be served on provisioned throughput or dedicated model-copy capacity — so a fine-tune that saves 20% on tokens can lose the business case to an always-on capacity bill. Run the utilization math before recommending tuning; chapters 03 covers when adaptation beats prompting at all.
Prompt + RAG first
Zero capacity commitment, instantly reversible, benefits from every base-model upgrade. Exhaust this before any tuning conversation.
Distillation
When a frontier model nails the task but unit economics demand a smaller model at scale. Teacher-generated data plus eval gates; serve the student where utilization justifies dedicated capacity.
Fine-tuning
For stable, high-volume, narrow tasks with real labeled data — formatting contracts, domain classification. Budget the provisioned-throughput floor into the ROI.
Custom model import
You tuned open weights on SageMaker or elsewhere and want Bedrock's API surface and guardrails over your own artifact.
Close the loop with Bedrock Evaluations: automatic metric jobs, LLM-as-a-judge with your prompt datasets, human-workforce evals, and RAG-specific evaluations that score retrieval and citation quality against a knowledge base. In interviews, position these as the AWS-native execution of the evaluation discipline from chapter 06 — the discipline is portable, the job runner is vendor-specific.
6. SageMaker for GenAI — and the honest self-hosting comparison
SageMaker AI is where you go when Bedrock's catalog or control surface is insufficient. The GenAI-relevant subset: JumpStart for one-click deploy/fine-tune of curated open models; managed training jobs (with spot capacity for interruptible work); real-time endpoints running Large Model Inference (LMI) containers — DJL-Serving images that wrap vLLM/TensorRT-LLM backends with continuous batching, so the chapter-02 serving optimizations arrive pre-packaged; asynchronous inference for large-payload, long-running requests with S3 in/out, an internal queue, and scale-to-zero; and HyperPod for large distributed training — resilient clusters with automated faulty-node replacement and checkpoint-resume, sold on improving goodput (useful training time over wall-clock) for multi-week jobs where a single unhandled hardware fault can cost days.
The interview classic is "Bedrock or self-hosted?" Answer it as a utilization and control question, not a loyalty question:
| Dimension | Bedrock (managed) | SageMaker LMI endpoint | vLLM on EKS (self-hosted) |
|---|---|---|---|
| Model choice | Catalog + custom import | Any open weights | Any weights, any runtime, day-zero releases |
| Ops burden | None on serving; quotas to manage | Instance/container choices; managed autoscaling | Full: GPU procurement, drivers, schedulers, upgrades, on-call |
| Unit economics | Per-token; excellent at low/spiky utilization | Per-instance-hour; wins at sustained moderate load | Per-GPU-hour; wins only at high sustained utilization with a capable team |
| Latency control | Limited knobs | Container/instance tuning | Everything: batching policy, quantization, speculative decoding |
| Compliance surface | AWS-attested service posture | Your containers in your VPC | Maximal control, maximal audit responsibility |
The senior nuance: this is not a one-time decision. Healthy platforms start on Bedrock for speed, instrument token telemetry from day one, and revisit placement per-workload once volumes stabilize — often landing on a hybrid where one high-volume, narrow task moves to a tuned open model on LMI/EKS while everything else stays managed.
7. The serverless application layer: streaming, orchestration, and queues
Most GenAI system-design failures on AWS happen in the application layer, not the model layer. Three patterns cover nearly every interview scenario.
Streaming chat. API Gateway REST integrations buffer responses and historically capped integrations near 29 seconds (now raisable for regional REST APIs, at a throttling trade-off) — both properties are wrong for token streams. The standard pattern is Lambda response streaming behind a function URL (fronted by CloudFront for auth, WAF, and TLS domain control), forwarding ConverseStream deltas as SSE. WebSockets via API Gateway or AppSync remain the fallback for bidirectional or fan-out cases — chapter 07 covers the protocol-level details.
sequenceDiagram
participant C as "Client"
participant F as "CloudFront + function URL"
participant L as "Lambda with response streaming"
participant B as "Bedrock ConverseStream"
C->>F: POST chat turn
F->>L: forward request
L->>B: ConverseStream call with tools
B-->>L: content deltas and tool-use events
L-->>C: SSE chunks as they arrive
B-->>L: stop reason plus usage counts
L-->>C: terminal event with usage
Workflow orchestration. Step Functions has optimized Bedrock integrations (synchronous InvokeModel and run-to-completion batch-job steps), per-state retry/backoff/catch, and a distributed map mode that fans out to very high concurrency over S3 objects. Standard workflows give exactly-once-style, auditable state transitions for long processes; Express workflows suit high-rate, short orchestration. This is the backbone for document pipelines and deterministic "agentic" flows.
Event-driven decoupling. EventBridge routes domain events; SQS absorbs bursts in front of workers with per-queue DLQs; every LLM-calling consumer gets bounded concurrency so a traffic spike degrades into queue depth rather than model-quota exhaustion. Chapter 07's idempotency and DLQ discipline applies verbatim — model calls are expensive side effects that must not be replayed blindly.
flowchart LR
IN["Documents arrive in S3"] --> EVB["EventBridge rule"]
EVB --> Q["SQS queue with DLQ"]
Q --> SFN["Step Functions distributed map"]
SFN --> PRS["Parse: Textract or custom Lambda"]
PRS --> EXT["Bedrock extraction: batch job or on-demand"]
EXT --> VAL["Validate: schema + guardrail + confidence"]
VAL -->|"pass"| OUT["S3 curated zone + DynamoDB metadata"]
VAL -->|"fail"| REV["Human review queue"]
OUT --> BI["Athena and QuickSight"]
In the batch pipeline, the highest-leverage details are the validation stage (schema-check every model output; route low-confidence extractions to human review rather than averaging them into the lake) and the choice between Bedrock batch inference for large nightly backfills versus on-demand calls inside the map for streaming arrivals. Textract (docs) still beats LLM parsing on cost and determinism for structured forms; use FM parsing where layout understanding genuinely requires it.
8. Security architecture and the Bedrock privacy posture
Regulated-industry interviews are won here. The Bedrock data-privacy posture, per the data protection documentation: prompts and outputs are not used to train base models and are not shared with model providers; inference for a region is processed in that region (or within the profile's geography when you opt into cross-region inference profiles); content is encrypted in transit and at rest. State those four clauses precisely — hand-waving "AWS says it's private" is a junior tell.
flowchart LR
subgraph VPC1["Customer VPC private subnets"]
APP["App or Lambda in VPC"]
end
APP -->|"PrivateLink, no public internet"| VPE["Interface VPC endpoint bedrock-runtime"]
VPE --> BRT["Bedrock runtime API"]
BRT --> GDR["Guardrail policy applied"]
BRT --> MIL["Invocation logs: KMS-encrypted S3 + CloudWatch"]
BRT --> TRL["CloudTrail audit trail"]
APP -.-> ROLE["IAM role: InvokeModel on pinned model ARNs"]
- IAM least privilege — scope
bedrock:InvokeModel/InvokeModelWithResponseStreamto specific model and inference-profile ARNs; separate roles for ingestion, serving, and evaluation; deny wildcard model access in SCPs for regulated accounts. - Network isolation — interface VPC endpoints for
bedrock-runtimeand agent runtimes keep invocation traffic off the public internet; endpoint policies restrict which principals and models the endpoint will serve. - Encryption — customer-managed KMS keys for custom models, knowledge bases, agent sessions, and log destinations; key policy = another audit boundary.
- Audit — CloudTrail for control-plane and API activity; model invocation logging for full request/response bodies to S3/CloudWatch — enabling it is a deliberate compliance decision because prompts often contain the sensitive data itself.
- Guardrails as policy, not the whole defense — prompt-injection resistance also requires tool-permission scoping and retrieval ACLs (chapters 05 and 11).
9. Observability and cost engineering
Bedrock emits CloudWatch metrics per model — invocation counts, latency, input/output token counts, and throttles (monitoring docs). Alarm on throttle rate and p95/p99 InvocationLatency, and trend token counts per feature because tokens are your bill. Invocation logging plus trace IDs from your gateway gives request-level forensics; AgentCore adds step-level OTEL traces for agents. Deeper LLMOps practice — drift, quality regression, incident runbooks — is chapter 11; here, know which AWS surface emits which signal.
Cost engineering is a pricing-dimension inventory plus three levers (Bedrock pricing):
Provisioned-throughput sizing deserves a worked shape (illustrative numbers): if telemetry shows a steady floor of, say, 200k input + 40k output tokens/min during business hours, and one model unit sustains a documented tokens/min ceiling for your model, you buy units to cover the floor and let on-demand (or an inference profile) absorb the spikes above it. Committing to peak instead of floor is the classic overspend; committing before you have per-feature token telemetry is the classic premature optimization. Application inference profiles — invocation profiles you create and tag per workload or tenant — are how spend shows up in Cost Explorer attributable to a team, which chapter 07's per-tenant metering then reconciles.
Fold this into the Well-Architected Generative AI Lens vocabulary when asked "how do you know this is production-ready?": model selection as a reversible decision, safety controls at every trust boundary, evaluation gates before promotion, cost visibility per workload, and operational readiness (quota plans, failover, incident paths). Naming the lens and then demonstrating two of its questions beats reciting all six pillars.
10. Interview scenarios: AWS solution-design drills
These three scenarios cover most senior AWS GenAI loops. Practice narrating each in under four minutes with a drawn diagram.
Scenario 1 — "A bank wants a contact-center assistant grounded in policy documents."
Strong answer skeleton: requirements first (residency, PII exposure, auditability, latency), then Figure 2's shape: S3 + Knowledge Base with hierarchical chunking, OpenSearch Serverless, Retrieve + Converse with a pinned model version, Guardrails with PII masking and contextual grounding, PrivateLink invocation path, invocation logging with CMK, and an eval gate (chapter 06) before any model or prompt change ships. The differentiator is naming what you would refuse: no cross-region inference profile until compliance clears the geography clause; no RetrieveAndGenerate if the bank requires a custom citation contract.
Scenario 2 — "The team has a LangGraph agent prototype; make it production-grade on AWS."
Strong answer skeleton: keep the framework, deploy on AgentCore Runtime for session isolation and long-running executions; move ad-hoc tool code behind Gateway as MCP tools with Identity handling inbound caller auth and outbound OAuth; add Memory instead of a homegrown session store; wire OTEL traces to CloudWatch; put Guardrails on model I/O; add an offline eval harness of recorded trajectories before each release. Flag the alternative honestly: if the graph is actually static, compile it into Step Functions and delete the autonomy.
Scenario 3 — "Process ten million archived contracts and extract obligations monthly."
Strong answer skeleton: Figure 5's pipeline with Bedrock batch inference as the extraction engine (the ~50% discount at this volume is decisive), Step Functions distributed map for orchestration, Textract for the structurally simple pages, JSON-schema validation with confidence thresholds routing to human review, DLQs with replay tooling, and per-run cost reporting via tagged inference profiles. Quantify: estimate tokens/document × corpus size, show the batch-versus-on-demand delta, and state the completion-time trade-off since batch has no latency SLA.
Interview playbook
Answer framework for AWS GenAI questions: (1) restate the workload in capability terms — latency class, volume shape, data sensitivity, autonomy level; (2) place it on the three-plane model (Bedrock consumption / SageMaker production / serverless application); (3) pick services with one named alternative each and a reason; (4) attach the cross-cutting story — IAM, network path, logging, cost attribution; (5) close with the first three production metrics you would watch.
- Senior signals — tokens/min math for throughput decisions; knowing custom models usually need provisioned capacity; the cross-region-profile geography caveat; treating Guardrails as one layer of defense-in-depth; cost per task, not cost per call.
- Common traps — proposing agents for deterministic workflows; defaulting to fine-tuning before RAG and prompting are exhausted; "OpenSearch because it's the default" without the cost floor; API Gateway in front of a token stream; enabling invocation logging without a PII story.
- Stay in your lane — retrieval science lives in chapter 04, agent patterns in 05, evals in 06, streaming protocols in 07, GCP equivalents in 09, LLMOps in 11. Reference them; do not re-derive them mid-answer.
Question bank
Q1You have three workloads: a spiky customer chatbot, a steady internal summarizer at 500k tokens/min, and a nightly re-processing job. How do you buy Bedrock inference for each?
Strong answer outline
- Chatbot: on-demand via a cross-region inference profile for burst headroom; retries with jitter plus a queue for overflow.
- Summarizer: measure the sustained floor, buy provisioned throughput units to cover it, let on-demand absorb the excess; commit to 1/6-month terms only after weeks of telemetry.
- Nightly job: batch inference from JSONL in S3 at the discounted rate; no latency SLA, so schedule with slack.
- Cross-cutting: prompt caching on the shared system prompt in all three.
Follow-up probes
- What changes if the summarizer uses a fine-tuned model? (Custom models generally require provisioned capacity anyway.)
- How do you detect that a provisioned commitment is now oversized?
Did you size from measured tokens/min, name the batch discount, and mention the throttle-handling ladder for on-demand? All three, or the answer reads as pricing-page recital.
Q2Design enterprise RAG on AWS. When do you use Knowledge Bases end-to-end, and when do you break out?
Strong answer outline
- Managed KB for ingestion: connectors, FM parsing for complex layouts, hierarchical or semantic chunking, sync scheduling.
RetrieveAndGeneratefor pilots and internal tools — fastest credible baseline with citations.- Break out to
Retrieve+ your own Converse call for custom reranking, query rewriting, multi-index routing, or strict response contracts. - Guardrails contextual grounding on the generation step; eval harness scoring retrieval and answer quality separately (chapter 06).
Follow-up probes
- How do you enforce document-level ACLs at retrieval time?
- What breaks when the corpus grows 100×?
You should articulate the managed-ingestion/custom-generation split as the default production posture, with a concrete trigger for leaving the fully managed path.
Q3OpenSearch Serverless, Aurora pgvector, or Kendra GenAI index — how do you choose the vector store for a Bedrock Knowledge Base?
Strong answer outline
- OpenSearch Serverless: hybrid search and scale with zero cluster ops, but an always-on OCU cost floor — wrong for tiny corpora and idle dev stacks.
- Aurora pgvector: vectors beside relational data, SQL metadata/ACL filtering, lowest incremental cost when Aurora already exists; you own HNSW tuning.
- Kendra GenAI index: managed relevance plus enterprise connectors and permission-aware results; pay a premium to skip embedding ops.
- Decide on: corpus size, hybrid-search need, existing ops skills, ACL model, and monthly cost floor — in that order.
Follow-up probes
- Where does S3 Vectors fit for a 500M-chunk archive?
- When would you run two stores deliberately?
Strong answers include at least one cost-floor observation and one ops-ownership observation; store choice is never purely a recall-quality argument.
Q4What do cross-region inference profiles actually do, and when would a regulator object?
Strong answer outline
- A profile ID that routes invocations across a set of regions within a geography for higher effective throughput and burst resilience; the required invocation path for many newer models.
- Data is processed in any region of that geography — in-transit encrypted, logs stay in the source region — but "processed only in eu-central-1" is no longer a true statement.
- Regulator conflict: residency commitments pinned to a single country/region; answer is single-region on-demand or provisioned capacity, accepting lower burst headroom.
- Check the documented region list per profile before promising anything.
Follow-up probes
- How do profiles interact with per-tenant cost attribution?
- What is your fallback when a single-region quota is exhausted and profiles are off the table?
You must state the geography-scope caveat unprompted; it is the entire point of the question.
Q5Bedrock Agents, AgentCore, or Step Functions calling Bedrock — how do you place an "agentic" workload?
Strong answer outline
- First test: is the flow actually dynamic? If steps are enumerable, Step Functions — replayable, auditable, per-step retries, no autonomy risk.
- Bedrock Agents for Bedrock-native quick builds: action groups on Lambda, KB attachment, return-of-control.
- AgentCore when you bring your own framework/model and need production infra: Runtime isolation and long sessions, Gateway for MCP tools, Identity for scoped credentials, Memory, OTEL observability.
- Whichever you pick: guardrails on I/O, tool permission scoping, trajectory evals before release.
Follow-up probes
- How does AgentCore Identity change the security review versus a shared service role?
- What breaks when an agent session must run for two hours?
The deterministic-flow test must come first. Recommending an agent platform before asking whether agency is needed is the trap.
Q6Walk through Guardrails policy types. What do contextual grounding checks catch, and what do they miss?
Strong answer outline
- Inventory: content filters (including prompt-attack), denied topics, word filters, PII detection with mask/block, contextual grounding and relevance scoring; attachable to invocations, agents, KBs, or any model via
ApplyGuardrail. - Grounding checks score whether the answer is supported by supplied context and relevant to the query — a threshold screen against hallucinated claims in RAG.
- They miss: correct-looking answers from wrong retrieval, factual errors within grounded text, multi-hop reasoning failures, and anything in a modality or language the scorer handles poorly.
- Position: one runtime layer inside defense-in-depth — retrieval ACLs, tool scoping, offline evals, and red-teaming still required.
Follow-up probes
- How do you tune thresholds without exploding false-block rates?
- Where does the guardrail run in a streaming response?
Give at least two concrete misses for grounding checks. "It stops hallucinations" is a failing answer at senior level.
Q7The product team wants to fine-tune on Bedrock to cut costs. Argue the full decision, including serving implications.
Strong answer outline
- Order of operations: prompting + RAG first, then distillation or fine-tuning only for stable, high-volume, narrow tasks with eval-proven gaps (chapter 03).
- Serving reality: customized models typically require provisioned or dedicated model-copy capacity — an always-on floor that can erase per-token savings at low utilization.
- Math sketch: tokens/day × per-token saving versus capacity-hours × unit price; include retraining cadence and eval-gate costs.
- Distillation alternative: teacher-generated data to specialize a cheaper student; custom model import if tuning happens on SageMaker.
Follow-up probes
- What eval evidence would greenlight the tune?
- How do you roll back a bad custom model in production?
The provisioned-capacity floor must appear in your cost argument. Without it you have answered a different, easier question.
Q8Bedrock, a SageMaker LMI endpoint, or vLLM on EKS — build the decision framework with numbers.
Strong answer outline
- Axes: model availability (catalog vs open weights vs day-zero), utilization shape, latency-control needs, team ops capacity, compliance surface.
- Economics: per-token beats per-GPU-hour at low/spiky utilization; the crossover appears only at high sustained utilization on open weights — show the break-even structure with example numbers, labeled as examples.
- SageMaker LMI as the middle path: managed instances, continuous-batching containers, your VPC, no GPU-cluster ops.
- Recommend hybrid-by-workload with telemetry-triggered revisits, not a single global answer.
Follow-up probes
- What telemetry proves the crossover has been reached?
- What hidden costs does self-hosting add beyond GPU hours?
Your answer needs a break-even formula shape and the honest admission that most teams overestimate their sustained utilization.
Q9Design the security architecture for Bedrock in a healthcare account. Be specific.
Strong answer outline
- Privacy posture stated precisely: no training on customer content, no sharing with model providers, in-region (or in-geography) processing, encryption in transit/at rest.
- IAM: invoke permissions pinned to model/profile ARNs; separate ingestion, serving, and eval roles; SCP denies wildcard model access.
- Network: interface VPC endpoints with endpoint policies; no public egress from serving subnets.
- KMS CMKs on KBs, custom models, and log destinations; CloudTrail everywhere.
- Invocation logging as a deliberate decision: prompts contain PHI, so pair with CMK, retention, access controls, or pre-log redaction.
Follow-up probes
- Who can read the invocation logs, and how do you prove that to an auditor?
- How do Guardrails PII filters interact with clinically necessary PHI in prompts?
The invocation-logging-contains-PHI observation is the discriminator; most candidates present logging as pure upside.
Q10Your Bedrock bill doubled last quarter. Take me through a 40% reduction without hurting quality.
Strong answer outline
- Attribute first: application inference profiles + cost allocation tags to find which workload/tenant grew; tokens per task, not per call.
- Prompt caching on stable prefixes (system prompts, tool schemas) — reorder prompts so volatile content comes last.
- Route: cheaper models for classify/extract tiers, frontier models only where evals prove the gap; shorten outputs with max-token and format contracts.
- Move offline work to batch inference; right-size or drop underused provisioned commitments.
- Gate every change with the eval suite so "cheaper" cannot silently mean "worse".
Follow-up probes
- Which of these ships in week one versus quarter one?
- How do you stop the regression from recurring?
Attribution before optimization, and eval gates on every lever — miss either and the answer is a cost-cutting listicle.
Q11Stream tokens to a browser through AWS serverless. What breaks with API Gateway, and what do you build instead?
Strong answer outline
- API Gateway REST buffers responses and has integration-timeout constraints — both hostile to long token streams.
- Pattern: Lambda response streaming behind a function URL, fronted by CloudFront for TLS, WAF, and auth; forward ConverseStream deltas as SSE.
- WebSockets (API Gateway) or AppSync for bidirectional needs, fan-out, or strict corporate proxy environments.
- Operational details: heartbeats, terminal events carrying usage, reconnect/resume semantics (chapter 07).
Follow-up probes
- Where do you apply output guardrails in a streaming path?
- How do you authenticate a function URL properly?
Name the buffering problem explicitly and give the CloudFront-fronted function-URL pattern; "just use WebSockets" without trade-offs is a shallow pass.
Q12Compare SageMaker async inference, Bedrock batch inference, and SQS + Lambda for long-running GenAI work.
Strong answer outline
- SageMaker async: your own model on a managed endpoint, big payloads via S3, internal queue, scale-to-zero — per-request async against a self-managed model.
- Bedrock batch: catalog models over large JSONL corpora at a discount — bulk offline, no completion SLA.
- SQS + Lambda (or Step Functions): async orchestration of any API-based model with your own retry, DLQ, and idempotency discipline — most flexible, most engineering.
- Choose by: whose model, payload/corpus shape, latency tolerance, and who owns the retry semantics.
Follow-up probes
- Where do duplicate model invocations come from in each design?
- How does the DLQ replay path revalidate before re-invoking?
You should place all three on the "whose model × latency tolerance" grid without conflating batch (corpus) with async (request).
Q13When does HyperPod beat standard SageMaker training jobs, and what does it actually buy you?
Strong answer outline
- Standard training jobs: ephemeral, managed, ideal for fine-tunes and experiments measured in hours; spot-friendly.
- HyperPod: persistent resilient clusters for multi-week distributed training — automated faulty-node detection/replacement and checkpoint-resume protect goodput, where a single unhandled hardware fault can cost days.
- Also relevant: cluster reuse across runs, Slurm/EKS orchestration options, and task governance for sharing capacity across teams.
- Framing: it is infrastructure for training-as-a-program, not for a one-off tune — most application teams never need it.
Follow-up probes
- What checkpoint cadence balances goodput against storage cost?
- Where does PEFT (chapter 03) remove the need for any of this?
The goodput argument — useful training time over wall-clock — is the senior framing; naming it beats listing features.
Q14Design a multi-tenant GenAI platform on AWS: isolation, fairness, and per-tenant cost.
Strong answer outline
- Identity and data isolation: tenant-scoped IAM/session context end-to-end; retrieval filtered by tenant ACLs at the store (pgvector SQL filters or per-tenant indexes).
- Fairness: per-tenant token budgets and rate limits at the gateway; bounded per-tenant queue concurrency so one tenant's burst degrades into their own queue depth.
- Cost: application inference profiles or tags per tenant; reconcile gateway-metered tokens against the AWS bill monthly.
- Blast radius: separate guardrail configs and eval baselines per tenant tier; noisy-neighbor alarms on throttle share.
Follow-up probes
- Pooled versus siloed vector indexes — where is the crossover?
- What changes for a tenant demanding single-region processing?
Cost attribution and fairness controls must be distinct mechanisms in your answer; conflating them signals you have not run a shared platform.
Proof artifact: a costed, secured, breakable RAG stack on AWS
Build one small system that produces evidence for five chapter themes at once: inference-mode choice, managed RAG, guardrails, security posture, and cost attribution. Keep the corpus tiny (50–100 public documents) so the whole lab stays in low tens of dollars — and tear down the vector store after, because that is where the idle cost lives.
- DeployS3 corpus → Knowledge Base → OpenSearch Serverless (or pgvector to compare); serving Lambda using Retrieve + ConverseStream behind a CloudFront-fronted function URL; guardrail with PII masking and contextual grounding attached.
- HardenPin IAM invoke permissions to exact model ARNs; add a bedrock-runtime VPC endpoint; enable invocation logging to a CMK-encrypted bucket; capture the CloudTrail evidence.
- MeasureDrive 200 scripted queries; record p50/p95 time-to-first-token, grounding-score distribution, tokens per answer, and cost per answered question from a tagged application inference profile.
- Break it — throttlingBurst well past your request quota; show the failure mode without retries, then with jittered retries plus an SQS shock absorber; graph both.
- Break it — groundingInject a contradictory "poison" document into the corpus; show the grounding-check score drop and the blocked/flagged response; discuss what it did not catch.
- CompareRe-run the query set through batch inference and through a cached-prefix variant; produce a one-page cost table: on-demand vs cached vs batch per 1k answers (label all figures as your measurements, dated).
What to present in an interview: the architecture diagram, the two failure-injection graphs, and the cost table. Three artifacts, each one sentence of setup — "I measured the cache discount myself" is worth more than any certification line on a résumé.
Chapter review
You now hold the AWS GenAI stack as three planes — Bedrock for consumption, SageMaker for production, serverless for application — crossed by security, observability, and cost. The recurring senior pattern: buy tokens deliberately (Figure 1), keep managed ingestion but own generation when contracts demand it (Figure 2), give agents infrastructure only after proving they need agency (Figure 3), stream without buffering hops (Figure 4), industrialize offline work with batch and state machines (Figure 5), and make the private, audited invocation path your default drawing (Figure 6).
- Converse API
- Bedrock's uniform multi-vendor inference interface with tools, multimodality, and streaming.
- Inference profile
- Routing identity for invocations; cross-region profiles trade single-region processing for throughput, application profiles carry cost tags.
- Provisioned throughput
- Dedicated model units with committed tokens/min; the serving floor for most customized models.
- Batch inference
- Async JSONL-over-S3 jobs at a deep token discount with no completion SLA.
- Knowledge Base
- Managed RAG ingestion and retrieval: connectors, parsing, chunking, embedding, vector-store writes, Retrieve/RetrieveAndGenerate.
- AgentCore
- Framework-agnostic agent infrastructure: Runtime, Gateway, Memory, Identity, built-in tools, OTEL observability.
- Contextual grounding check
- Guardrails policy scoring answer support and relevance against supplied context — a screen, not a proof.
- LMI container
- SageMaker large-model-inference image wrapping continuous-batching backends like vLLM.
- HyperPod
- Resilient persistent training clusters engineered for goodput on multi-week distributed jobs.
- Model invocation logging
- Optional full request/response capture to S3/CloudWatch — an audit asset and a PII liability simultaneously.
- I can choose among on-demand, inference profiles, provisioned throughput, and batch with tokens-per-minute reasoning, and state the cross-region geography caveat.
- I can defend a vector-store choice with a cost floor, an ops-ownership argument, and an ACL story.
- I can place a workload across Step Functions, Bedrock Agents, and AgentCore — and justify refusing agency.
- I can name all Guardrails policy types and two failure modes of contextual grounding checks.
- I can state the Bedrock privacy posture in four precise clauses and design the PrivateLink + KMS + CloudTrail invocation path.
- I can sketch the Bedrock-versus-self-hosted break-even and say what telemetry would trigger a revisit.
- I can cut a Bedrock bill with caching, routing, batch, and right-sized commitments — attribution first, eval gates always.
- I can narrate all three reference architectures from memory in under four minutes each.
Primary sources
Links checked 2026-08-04.
- Amazon Bedrock — User Guide
- Amazon Bedrock — Converse API
- Amazon Bedrock — Provisioned throughput
- Amazon Bedrock — Batch inference
- Amazon Bedrock — Cross-region inference
- Amazon Bedrock — Prompt caching
- Amazon Bedrock — Knowledge Bases
- Amazon Bedrock — Agents
- Amazon Bedrock AgentCore
- Amazon Bedrock — Guardrails
- Amazon Bedrock — Custom models
- Amazon Bedrock — Model evaluation
- Amazon Bedrock — Data protection
- Amazon Bedrock — Interface VPC endpoints
- Amazon Bedrock — Model invocation logging
- Amazon Bedrock — Monitoring
- Amazon Bedrock — Pricing
- Amazon SageMaker AI — Developer Guide
- SageMaker JumpStart
- SageMaker — Large model inference containers
- SageMaker — Asynchronous inference
- SageMaker HyperPod
- OpenSearch Serverless — Vector search collections
- Amazon Aurora — User Guide
- Amazon Kendra — Developer Guide
- Amazon Textract — Developer Guide
- AWS Lambda — Response streaming
- AWS Step Functions — Developer Guide
- AWS CloudTrail — User Guide
- AWS KMS — Developer Guide
- AWS Well-Architected — Generative AI Lens
- Generative AI on Vertex AI — documentation (for the GCP mapping)
- Lewis et al., 2020 — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv:2005.11401)
CHAPTER 09 · CLOUD TRACK
The GCP GenAI Stack: Vertex AI, Gemini & Agent Builder
37 min read · 14 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- choose between the consumer Gemini API and Vertex AI Gemini on auth, quota, and enterprise-control grounds — and explain Dynamic Shared Quota versus Provisioned Throughput with tokens-per-second reasoning;
- design enterprise RAG on GCP with a justified pick among Grounding with Google Search, Vertex AI Search, RAG Engine, and a fully custom Vector Search build;
- defend a GCP vector-store decision — Vertex AI Vector Search (ScaNN), AlloyDB pgvector, or BigQuery vector search — with latency-class and cost-floor arguments;
- place agent workloads across ADK, Agent Engine, and self-managed runtimes, and say precisely what A2A adds and when it is premature;
- argue Gemini supervised fine-tuning versus adapters on open models, including the serving-economics asymmetry against AWS;
- use BigQuery as a GenAI data platform — ML.GENERATE_EMBEDDING, vector indexes, and row-wise generation — and know its latency limits;
- layer IAM, VPC Service Controls, CMEK, residency, and audit logging into a security answer, and cut a Vertex bill with caching, batch, and routing.
1. The GCP GenAI stack as one mental model
Google Cloud's GenAI story is more centralized than AWS's: almost everything routes through Vertex AI — model consumption (Gemini plus Model Garden), tuning, vector search, agents, and evaluation live under one API surface and one IAM model. The second pillar is BigQuery, which has quietly become a GenAI platform in its own right: embeddings, vector search, and LLM calls as SQL over governed data. The third is the compute substrate — GKE, Cloud Run, and TPUs — for anything you serve yourself. Hold this layering and you can place any interview question in seconds.
Since senior loops routinely ask "why GCP over AWS here?", keep the service-level mirror ready. Chapter 08 owns the AWS depth; this map is the translation table:
AWS
- Bedrock Conversemulti-vendor managed inference
- Bedrock Knowledge Basesmanaged RAG ingestion + retrieval
- Bedrock Agents / AgentCoreagent orchestration + runtime infra
- OpenSearch Serverless / Aurora pgvectorvector stores
- SageMaker LMI / EKS + vLLMself-managed open-model serving
- PrivateLink + KMS + CloudTrailprivate path, keys, audit
- Redshift MLin-warehouse ML (thin GenAI story)
Google Cloud
- Vertex AI generateContentGemini + Model Garden inference
- Vertex AI Search / RAG Enginemanaged grounding + retrieval
- ADK + Agent Engineagent framework + managed runtime
- Vertex AI Vector Search / AlloyDB / BigQueryvector stores
- GKE + vLLM / Cloud Run GPUs / Vertex endpointsself-managed serving
- VPC-SC + PSC + CMEK + audit logsperimeter, private path, keys, audit
- BigQuery ML + AI functionsin-warehouse GenAI at full strength
2. Gemini inference: two API surfaces, quotas, caching, and batch
The first discriminating question in any GCP loop: which Gemini API are you calling? The Gemini Developer API (AI Studio) authenticates with API keys, is optimized for velocity and free-tier experimentation, and offers few enterprise controls. Vertex AI Gemini serves the same models through aiplatform.googleapis.com with IAM/service-account auth, regional endpoints, VPC Service Controls compatibility, CMEK on stored artifacts, audit logging, and the data-governance commitments enterprises require. The unified Google Gen AI SDK targets both surfaces with a flag flip, so the senior recommendation is: prototype anywhere, but production traffic with corporate data goes through Vertex — the migration is a config change, not a rewrite. The core methods are generateContent and streamGenerateContent (SSE), with system instructions, tool/function declarations, structured output, and safety settings in one request shape across Gemini versions.
How you buy tokens is the second discriminator. Vertex replaced most fixed per-model rate quotas for Gemini with Dynamic Shared Quota (DSQ): pay-as-you-go requests draw from a shared regional capacity pool with no per-project guarantee — you get elasticity but must engineer for 429s. When you need contractual throughput, Provisioned Throughput sells committed capacity in generative AI scale units (GSUs) on weekly-to-multi-month terms per model, with overage spilling to DSQ by default. This is the exact analog of the Bedrock on-demand-versus-provisioned decision from chapter 08, with one twist worth naming: the global endpoint raises availability by routing to any region — mirroring AWS cross-region inference profiles, and carrying the same residency caveat.
flowchart TD
A["New Gemini workload on Vertex AI"] --> B{"Latency-coupled user traffic?"}
B -->|"no, offline"| C["Batch prediction: BigQuery or GCS input at a deep discount"]
B -->|"yes"| D{"Hard throughput SLO or launch spike?"}
D -->|"spiky, exploratory"| E["Pay-as-you-go on Dynamic Shared Quota"]
D -->|"steady floor, hard SLO"| F["Provisioned Throughput in GSUs"]
E --> G["Engineer for 429s: retries with jitter, regional failover"]
F --> H["Size GSUs from measured tokens per second, spill bursts to DSQ"]
C --> J["Outputs land in BigQuery for evals and joins"]
Two multipliers change the math on any purchase mode. Context caching comes in two forms: implicit caching, on by default for current Gemini models, which discounts input tokens automatically when your request shares a prefix with recent traffic; and explicit caching, where you create a cached-content object with a TTL and pay a per-token-hour storage fee in exchange for a steep discount (on the order of 75% off standard input pricing per the pricing page, as of the checked date) on every hit. Explicit caching wins when a large stable block — tool schemas, a policy manual, a video — is reused heavily inside the TTL; it loses when hit rates are low and storage-hours dominate. Batch prediction reads from a BigQuery table or GCS JSONL, runs with no latency SLA, and prices at roughly half of interactive rates — the BigQuery-native input/output is a genuine differentiator over AWS's S3-only batch path for analytics-adjacent workloads.
sequenceDiagram
participant C as "Client"
participant R as "Cloud Run service"
participant V as "Vertex AI Gemini endpoint"
C->>R: POST chat turn
R->>V: streamGenerateContent with cachedContent reference
V-->>R: token deltas over server-sent events
R-->>C: forwarded SSE chunks as they arrive
V-->>R: usageMetadata with cached token counts
R-->>C: terminal event with usage and finish reason
Finally, Model Garden is the catalog: Gemini first-party; partner models as-a-service — notably Anthropic Claude on Vertex, billed through GCP with the same enterprise controls; and open models (Gemma, Llama, Mistral, DeepSeek, Qwen) that you deploy to Vertex endpoints or export to your own GKE/Cloud Run serving. The senior distinction is model-as-a-service versus self-deployed: MaaS models bill per token with zero capacity management; self-deployed open models bill per accelerator-hour on endpoints you size — that difference drives section 8's economics.
3. Grounding: Google Search, Vertex AI Search, and RAG Engine
GCP offers a graduated ladder of grounding options, and interviewers test whether you can place a use case on the right rung instead of hand-building RAG by reflex. The grounding overview frames three managed rungs before DIY:
| Option | What it is | Pick when | Watch out for |
|---|---|---|---|
| Grounding with Google Search | One request flag; Gemini retrieves from the live web and returns cited, grounded answers | Freshness on public facts: news, products, competitors, regulations | Priced per grounded request beyond a free tier; results are the public web — no corporate data; citation display requirements apply |
| Vertex AI Search | Turnkey enterprise retrieval app: connectors (GCS, Drive, SharePoint, Confluence, sites), parsing, chunking, hybrid retrieval + Google-grade ranking, ACL-aware results | Enterprise search + RAG over heterogeneous corpora with document permissions; fastest credible production baseline | Less control over chunking/embedding internals; per-query and per-index pricing needs modeling at scale |
| RAG Engine | Managed RAG framework: corpora and files API, configurable chunking/embedding, pluggable vector backends (managed store, Vector Search, Pinecone, Weaviate) | You want programmable control of the pipeline without owning infrastructure — the middle rung | Newer surface; validate connector and scale fit before promising it in a design |
| DIY on Vector Search | Own everything: parsing, chunking, embeddings, index, reranking, generation | Custom retrieval science (chapter 04), strict contracts, extreme scale | You now own relevance, ops, and evals end to end |
The reference architecture below is the one to draw for "enterprise RAG on GCP." The load-bearing choices: Vertex AI Search for ingestion and ACL-aware retrieval (managed), your own generation call for contract control (custom), Model Armor screening on the way out, and logging to BigQuery so evaluation (chapter 06) runs as SQL over real traffic.
flowchart LR
subgraph ING["Ingestion path"]
SRC["GCS, Drive, SharePoint sources"] --> VAS["Vertex AI Search: parse, chunk, embed, rank"]
end
subgraph SRV["Serving path"]
CL["Client app"] --> RUN["Cloud Run API with IAM auth"]
RUN --> VAS
RUN --> GEM["Gemini generateContent with retrieved chunks"]
GEM --> MA["Model Armor: injection and safety screen"]
MA --> RUN
end
GEM --> LOG["Request-response logging to BigQuery"]
RUN --> OBS["Cloud Trace and Cloud Logging"]
The break-out logic mirrors chapter 08's Knowledge Bases discussion: stay fully managed (Vertex AI Search answer generation) for pilots and internal tools; split to managed retrieval, custom generation when you need response contracts, custom reranking, or multi-source federation; drop to RAG Engine or DIY only when the managed ranker demonstrably fails your eval set. Saying "I would benchmark Vertex AI Search's ranking against my custom pipeline before building anything" is a stronger senior move than defaulting to either extreme — Google's ranking stack is genuinely hard to beat on heterogeneous enterprise corpora.
4. The vector-store decision: Vector Search, AlloyDB, BigQuery
Vertex AI Vector Search is the productization of Google's ScaNN research (Guo et al., anisotropic vector quantization) — the same ANN family behind Google Search and YouTube retrieval. Know its shape: a tree-AH index (with a brute-force option for ground-truthing recall), deployed to index endpoints with dedicated serving replicas, streaming index updates for near-real-time upserts versus cheaper batch rebuilds, namespace/numeric restricts for filtered ANN, and autoscaling. Its trade profile: excellent recall-QPS-latency at tens of millions to billions of vectors, but an always-on serving-node cost floor — the same "idle dev index still bills" caveat as OpenSearch Serverless on AWS.
| Store | Strengths | Costs and caveats | Pick when |
|---|---|---|---|
| Vertex AI Vector Search | ScaNN-grade recall/latency at very large scale; streaming upserts; filtered ANN; managed autoscaling | Always-on endpoint replicas; separate system from your relational data; index-build costs on batch updates | Large corpora (tens of millions+), strict low-latency ANN, high QPS |
| AlloyDB AI + pgvector | Vectors beside relational rows; SQL joins and ACL filters; AlloyDB's ScaNN index option accelerates pgvector well beyond stock HNSW at scale | You own index choice and tuning; a full Postgres fleet to run; scale ceiling below dedicated ANN services | Corpus lives next to transactional data; metadata/ACL filtering in SQL; small-to-mid scale with existing Postgres skills |
| BigQuery vector search | Embeddings and VECTOR_SEARCH where the data already lives; IVF and ScaNN-based TreeAH indexes; zero new infrastructure; governed by BQ IAM and row-level security | Analytical latency class — seconds, not milliseconds; slot/on-demand query economics; not for chat-path retrieval | Batch semantic joins, offline RAG evals, entity resolution, analytics enrichment |
The interview-ready heuristic: latency class first, data gravity second, ops ownership third. Chat-path retrieval under ~100 ms at high QPS → Vector Search. Retrieval that must join user entitlements and transactional state → AlloyDB pgvector. Anything batch or analytical → BigQuery, and moving those workloads out of a serving store is often the cheapest optimization you can name. Chapter 04 owns retrieval science (chunking, hybrid search, reranking); your GCP-specific value is this placement argument plus the streaming-versus-batch index-update trade: streaming updates cost more per write but close the freshness gap for use cases like product catalogs, where a nightly batch rebuild silently serves stale inventory all day.
5. The Agent Builder ecosystem: ADK, Agent Engine, and A2A
Google's agent stack cleanly separates framework from runtime, and interviewers reward candidates who use that separation. The Agent Development Kit (ADK) is the open-source framework (Python and Java): LLM agents, deterministic workflow agents (sequential/parallel/loop), multi-agent hierarchies with delegation, tools (function tools, built-in Google Search and Vertex AI Search tools, OpenAPI tools, MCP tool support), callbacks for guardrails, and a local dev UI with built-in evaluation. ADK code is deployable anywhere a container runs. Vertex AI Agent Engine is the managed runtime: sessions, a managed Memory Bank for long-term memory, scaling, identity, VPC-SC compatibility, and tracing — the place you deploy an ADK (or LangGraph, or CrewAI) agent when you stop wanting to own that infrastructure. It is the direct counterpart of Bedrock AgentCore Runtime from chapter 08.
- ADK — agent logic as code: planners, workflow agents, tool declarations, callbacks; framework-agnostic deploy target.
- Agent Engine — managed sessions, Memory Bank, scaling, and runtime isolation; bring ADK or another framework.
- A2A protocol — an open, Linux Foundation-governed protocol (spec) for agent-to-agent discovery and task exchange across vendors and runtimes via agent cards.
- Tools & extensions — built-in Search/RAG tools, code execution, function calling, MCP servers, and Apigee/Application Integration connectors for enterprise systems.
- Observability — Agent Engine emits traces to Cloud Trace and logs to Cloud Logging; ADK's eval harness runs trajectory tests pre-deploy.
flowchart TD
U["Caller: app or workflow"] --> AE["Agent Engine runtime: sessions and scaling"]
AE --> ADK["ADK agent: planner plus workflow sub-agents"]
ADK --> GEM["Gemini on Vertex AI"]
ADK --> T1["Built-in tool: Vertex AI Search grounding"]
ADK --> T2["Function tools on Cloud Run"]
ADK --> T3["MCP tools and Apigee connectors"]
AE --> MB["Memory Bank: long-term user memory"]
ADK --> PA["Peer agent via A2A agent card"]
AE --> OBS["Cloud Trace spans and Cloud Logging"]
The placement decision mirrors AWS but with GCP's accents. If the flow is deterministic, Workflows or a plain Cloud Run service calling Gemini beats any agent framework — same "refuse the agency" test as chapter 08, and chapter 05 covers when agency is genuinely warranted. If you need an agent, ADK-on-Agent-Engine is the default GCP answer because session state, memory, and tracing arrive managed. Self-hosting an agent loop on Cloud Run remains right when you need custom runtime behavior or already operate that infra. On A2A: position it as the inter-organizational and cross-runtime seam — valuable when agents from different teams or vendors must interoperate; premature when one team owns all agents in one runtime, where in-process delegation via ADK sub-agents is simpler and faster. Knowing that boundary — MCP standardizes agent-to-tool, A2A standardizes agent-to-agent — is a reliable senior discriminator in 2026 loops.
6. Tuning on Vertex: Gemini SFT and the serving-economics asymmetry
Supervised fine-tuning for Gemini is the managed adaptation path: JSONL prompt-completion datasets, adapter-based training under the hood, versioned tuned models, and — the fact that wins interviews — tuned Gemini models serve on the shared endpoint at the same per-token price as the base model. Contrast chapter 08: Bedrock custom models generally require provisioned throughput or dedicated model copies, an always-on capacity floor that can erase per-token savings. On GCP, the marginal serving cost of a tune is approximately zero, which moves the break-even for fine-tuning meaningfully earlier. If a candidate can articulate that asymmetry with the utilization math, they have demonstrated real multi-cloud judgment rather than feature recitation.
Prompt + RAG first
Same discipline as everywhere: exhaust prompting, context engineering, and grounding (chapter 03) before any tuning conversation. Tuning locks you to a model version and adds an eval-and-retrain cadence.
Gemini SFT
Stable, narrow, high-volume tasks — formatting contracts, domain classification, style enforcement — with hundreds-to-thousands of quality labeled pairs. No serving-capacity penalty on Vertex; budget the tuning-job cost and the eval gates.
Distillation
Use a frontier model to generate training data for a cheaper model (Gemini Flash tiers or an open model). The chapter 06 eval harness is the gatekeeper; the win is unit economics at scale.
PEFT on open models
LoRA/QLoRA adapters on Gemma/Llama via Model Garden training or your own GKE jobs, served on endpoints you size (section 8). Choose when you need weight ownership, on-prem portability, or task performance no API model reaches.
Scope discipline for the interview: chapter 03 owns when and how to adapt (LoRA math, data curation, catastrophic-forgetting risks); your GCP-specific claims are the offering matrix — SFT is the supported managed path for current Gemini text models, adapter training for open models runs as Vertex custom jobs or Model Garden recipes — and the serving-economics argument above. If asked about RLHF-style preference tuning on Vertex, the honest answer is that managed preference tuning has come and gone from the catalog; verify current model support in the docs rather than asserting from memory — saying exactly that earns more trust than a confident guess.
7. BigQuery as an AI data platform
BigQuery's GenAI surface (docs) turns the warehouse into a batch GenAI runtime: ML.GENERATE_EMBEDDING calls a remote Vertex embedding model over millions of rows; CREATE VECTOR INDEX builds IVF or TreeAH (ScaNN-family) indexes; the VECTOR_SEARCH table function does semantic joins in SQL; and AI.GENERATE-family functions (successors to ML.GENERATE_TEXT) run row-wise Gemini generation with structured outputs — all governed by BigQuery IAM, row-level security, and lineage, with zero data movement. Classic BigQuery ML still handles tabular models beside it. There is no AWS equivalent of comparable depth — Redshift ML is far thinner — which makes this a legitimate, defensible "why GCP" argument in cloud-choice questions.
flowchart LR
SRC["Raw tables in BigQuery"] --> EMB["ML.GENERATE_EMBEDDING via remote Vertex model"]
EMB --> IDX["Vector index: IVF or TreeAH"]
IDX --> VS["VECTOR_SEARCH semantic joins"]
SRC --> GEN["AI.GENERATE row-wise enrichment with Gemini"]
GEN --> CUR["Curated, enriched tables"]
VS --> APP["Similarity features for apps and models"]
CUR --> LKR["Looker dashboards and activation"]
Use cases that belong here: product-catalog enrichment (classify, normalize, describe millions of SKUs), support-ticket triage and clustering, entity resolution via embedding similarity, offline evaluation of RAG systems against logged traffic, and semantic deduplication. The boundary to state crisply: BigQuery is a batch and analytical latency class. Row-wise generation over big tables runs through Vertex quota (reserve Provisioned Throughput or run during batch windows for large jobs), and VECTOR_SEARCH answers in seconds, not the tens of milliseconds a chat path needs. The senior architecture is complementary: embeddings computed and evaluated in BigQuery, then synced to Vector Search or AlloyDB for online serving — one embedding lineage, two latency classes.
8. Serving open models: GKE + vLLM, Cloud Run GPUs, Vertex endpoints
When the model is open-weights, GCP gives you three serving postures, and the decision is utilization economics plus ops appetite — the same framework as chapter 08's Bedrock-versus-EKS argument, with chapter 02 owning the serving internals (continuous batching, paged KV cache, quantization).
| Posture | What you get | Economics | Pick when |
|---|---|---|---|
| Vertex endpoint (Model Garden deploy) | One-click deploy with prebuilt vLLM/TGI containers; managed autoscaling; Vertex API surface, IAM, monitoring | Per accelerator-hour while deployed; no scale-to-zero on dedicated endpoints | You want open weights behind the same Vertex plane as Gemini with minimal ops |
| GKE + vLLM | Full control: GPU classes (A3/A4) or TPUs, GKE Inference Gateway with prefix-cache-aware routing, custom autoscaling on batch-depth metrics | Per GPU/TPU-hour; wins only at high sustained utilization with a team to run it; committed-use discounts apply | High steady volume, day-zero models, custom runtimes, latency engineering |
| Cloud Run GPUs | Serverless L4 GPUs, per-second billing, scale-to-zero, fast cold-ish starts for small models | Pay only while serving; the cheapest posture for spiky or low-duty-cycle traffic on 7–27B-class models | Bursty internal tools, small fine-tuned models, prototypes that must not idle-bill |
Cloud Run GPUs are the genuinely differentiated option to name: AWS has no serverless-GPU-container equivalent with scale-to-zero in the same shape, so "spiky Gemma-class workload → Cloud Run GPU" is a crisp GCP-specific answer. The TPU card matters too — vLLM has TPU support, and TPU capacity is sometimes easier to obtain than H100-class GPUs — but present it honestly: it pays off at sustained scale with a team willing to benchmark, not as a default. For everything else, the chapter 08 break-even logic transfers verbatim: per-token MaaS beats per-hour self-hosting until utilization is provably high, and most teams overestimate their sustained utilization.
9. Security architecture: IAM, VPC Service Controls, CMEK, residency
The GCP security story has one concept AWS answers differently, and leading with it wins regulated-industry interviews: VPC Service Controls puts a data-exfiltration perimeter around API services themselves. Inside a perimeter, aiplatform.googleapis.com can only be called from authorized networks/identities, and — the crucial half — data cannot flow out to non-perimeter projects even by a credentialed insider or a leaked service-account key. PrivateLink on AWS gives you a private network path; VPC-SC gives you a policy boundary on the service plane. Pair it with Private Service Connect for private connectivity and you have both.
flowchart LR
subgraph PER["VPC Service Controls perimeter"]
APP["App in private VPC"] --> PSC["Private Service Connect endpoint"]
PSC --> VAPI["Vertex AI regional endpoint"]
VAPI --> KMS["CMEK via Cloud KMS on tuned models, indexes, caches"]
end
APP -.-> SA["Least-privilege service account via Workload Identity"]
VAPI --> AUD["Admin Activity plus opt-in Data Access audit logs"]
VAPI --> RESID["Regional processing for residency commitments"]
- IAM and service accounts — scope
roles/aiplatform.userper workload; workloads authenticate via Workload Identity Federation, never exported keys; separate service accounts for ingestion, serving, and tuning so blast radius is per-function. - CMEK — customer-managed keys on tuned models, Vector Search indexes, context caches, and datasets; key revocation is your kill switch and the key policy is a second audit boundary.
- Data governance — per the Gen AI data-governance docs: customer prompts and outputs are not used to train foundation models without permission; state that precisely, then name the operational caveats you would verify — caching behavior and any abuse-monitoring retention, with zero-retention configurations available for stricter regimes.
- Residency — regional endpoints keep ML processing in-region; the global endpoint and cross-region features trade that away for availability. Same clause structure as AWS cross-region inference profiles: compliance signs off first.
- Audit — Cloud Audit Logs: Admin Activity is always on; Data Access logs for Vertex are opt-in and can capture request content — the same "audit asset, PII liability" tension as Bedrock invocation logging, so pair enabling them with CMEK buckets, retention rules, and access reviews.
- Model Armor — a model-independent screening service for prompt injection, jailbreaks, sensitive-data leakage, and unsafe content on both prompts and responses; GCP's counterpart to Bedrock Guardrails, and like it, one layer of defense-in-depth, not the whole story (chapters 05 and 11).
10. Observability and cost engineering on Vertex
Vertex AI publishes per-model metrics to Cloud Monitoring — invocation counts, latencies, token throughput, and error/throttle rates — and application-level telemetry flows through Cloud Logging and Cloud Trace (Agent Engine emits OpenTelemetry spans natively). The GCP-specific habit worth naming: route request-response logs and token usage into BigQuery, because your eval harness, cost attribution, and drift analysis then become SQL over one table instead of three tools. Alarm on 429 rate (DSQ pressure), p95 time-to-first-token, and tokens-per-request drift — the last one catches prompt regressions and context-stuffing bugs before the invoice does. Deeper LLMOps discipline lives in chapter 11; here, know which surface emits which signal.
Cost engineering on Vertex is an inventory of pricing dimensions (pricing page) plus levers in a fixed order. First, model routing: Gemini Flash tiers are多 an order of magnitude cheaper than Pro — route by task difficulty and escalate on failure (chapter 07's gateway pattern). Second, context caching: explicit caches for heavy shared prefixes, with the storage-hour term in your break-even; implicit caching rewards stable prompt layouts for free. Third, batch everything latency-tolerant for the ~50% discount, with BigQuery-native I/O. Fourth, Provisioned Throughput sized to the measured floor, not the peak — GSU commitments are per-model and per-term, so commit late and let DSQ absorb spikes. Watch the long-context pricing tier: above the per-model threshold, input tokens price higher, so a lazy "stuff the whole corpus in context" design can double unit cost silently. For self-hosted serving, standard committed use discounts on GPUs/TPUs apply — a different ledger from token spend, and conflating the two in an interview reads as never having owned a bill.
11. Interview scenarios: GCP solution-design drills
Three scenarios cover most senior GCP GenAI loops. Practice narrating each in under four minutes with a drawn diagram, and close every one with the first-week production metrics you would watch.
Scenario 1 — "A hospital network wants a clinician assistant over internal protocols; data cannot leave their boundary."
Strong answer skeleton: requirements first — residency, PHI exposure, auditability, then Figure 3's shape hardened by Figure 6: Vertex AI Search over GCS/Drive protocol sources with ACL-aware retrieval, custom generation via regional Gemini endpoints (no global endpoint until compliance clears it), the whole project set inside a VPC Service Controls perimeter with PSC access, CMEK on every stored artifact, Data Access logs enabled into a locked, CMEK-encrypted sink, Model Armor on both directions, and chapter 06 eval gates before any prompt or model change. Name the refusals: no Grounding with Google Search (public-web egress), no consumer Gemini API anywhere in the path, no Data Access logging without a PHI-retention review.
Scenario 2 — "We have three teams building agents; make it a platform, not chaos."
Strong answer skeleton: Figure 4 as the target: ADK as the shared framework (with its eval harness as the pre-deploy gate), Agent Engine as the common runtime giving sessions, Memory Bank, and Cloud Trace uniformly; tools exposed through a governed catalog — Apigee/MCP for enterprise APIs so tool auth is centralized, not per-agent; A2A only at the seams where teams' agents must interoperate as black boxes. Per-team service accounts and budgets; token telemetry to BigQuery for per-agent cost. The differentiator: state the deterministic-flow test first — any "agent" whose steps are enumerable gets compiled into Workflows and removed from the platform's risk surface.
Scenario 3 — "A retailer wants semantic search and AI-enriched product data for 20M SKUs."
Strong answer skeleton: Figure 5 for the offline half — AI.GENERATE enrichment and ML.GENERATE_EMBEDDING in BigQuery as scheduled batch (batch-priced tokens, PT reservation if windows are tight), evaluated in place with SQL over sampled outputs; then sync embeddings to Vertex AI Vector Search with streaming updates for the online half, because inventory freshness is the business requirement. Quantify the token math per SKU as an estimate you would refine, name the two latency classes explicitly, and give the cost levers: Flash-tier models for enrichment, TreeAH index economics versus serving-replica count, and re-embedding cadence tied to catalog churn, not calendar habit.
Interview playbook
Answer framework for GCP GenAI questions: (1) restate the workload in capability terms — latency class, volume shape, data sensitivity, freshness; (2) place it on the three-pillar model (Vertex AI plane / BigQuery data plane / GKE-Cloud Run compute plane); (3) pick services with one named alternative each and the reason; (4) attach the cross-cutting story — service accounts, VPC-SC, CMEK, audit logs, cost attribution; (5) close with the first three production metrics you would watch.
- Senior signals — DSQ versus GSU capacity reasoning; explicit-versus-implicit caching with the storage-hour term; tuned-Gemini shared-serving economics versus Bedrock; the two-latency-class embedding architecture (BigQuery offline, Vector Search online); VPC-SC as a perimeter, not a network path.
- Common traps — treating DSQ as "no limits"; proposing BigQuery vector search on a chat path; Grounding with Google Search in a data-egress-restricted design; conflating Agent Engine with Gemini Enterprise; quoting prices from memory instead of naming the dimension and checking the page.
- Stay in your lane — retrieval science is chapter 04, agent patterns chapter 05, evals chapter 06, gateway/streaming chapter 07, AWS equivalents chapter 08, LLMOps chapter 11. Reference them; do not re-derive them mid-answer.
Question bank
Q1When would you use the Gemini Developer API versus Vertex AI Gemini, and what actually changes when you switch?
Strong answer outline
- Developer API: API-key auth, free-tier velocity, AI Studio prototyping — right for experiments and consumer-grade apps without corporate data.
- Vertex AI: IAM/service-account auth, regional endpoints, VPC-SC and CMEK compatibility, audit logging, enterprise data-governance terms, DSQ/Provisioned Throughput purchasing.
- The unified Gen AI SDK makes the switch a configuration change — so prototype fast, then promote to Vertex before real data flows.
- Decision rule: the moment corporate data, compliance scope, or throughput guarantees enter, Vertex — the controls do not exist on the other surface.
Follow-up probes
- What breaks in your security review if a team ships to production on API keys?
- How do quotas differ between the two surfaces?
You must name at least three concrete enterprise controls (not "it's more secure") and the SDK portability fact. Bonus: knowing both serve the same underlying models.
Q2Explain Dynamic Shared Quota. How do you guarantee throughput for a product launch on Vertex?
Strong answer outline
- DSQ: pay-as-you-go Gemini requests draw from a shared regional pool — no fixed per-project TPM, no guarantee; 429s appear under regional pressure.
- Guarantees come from Provisioned Throughput: GSUs committed per model and term, sized from measured tokens/sec, with overage spilling to DSQ.
- Launch plan: load-test to get tokens/sec at peak, buy GSUs for the confident floor, keep DSQ spill with retries/jitter and a queue for bursts, consider the global endpoint if residency allows.
- Post-launch: watch 429 rate and GSU utilization; shrink or grow the commitment at term boundaries.
Follow-up probes
- How does this differ from Bedrock's quota model?
- What if the launch spike is 10× the floor for one hour a day?
Sized-from-telemetry GSUs plus the spill-to-DSQ behavior are mandatory. "Vertex autoscales for you" is a failing answer.
Q3Walk through context caching on Vertex — implicit versus explicit — and when caching loses money.
Strong answer outline
- Implicit: automatic prefix-matching discount on current Gemini models; free to enable, rewards stable prompt layouts (volatile content last).
- Explicit: a cachedContent object with a TTL; you pay per-token-hour storage and get a steep per-hit discount on cached input tokens.
- Break-even: discount × hits must exceed storage cost over the TTL — high-traffic shared prefixes (tool schemas, policy docs, long videos) win; low-QPS or per-user-unique contexts lose.
- Also a latency lever: cache hits skip prefill, cutting time-to-first-token on long contexts (chapter 02 mechanics).
Follow-up probes
- How do you observe your cache-hit rate in production?
- Why does putting the user question first in the prompt destroy implicit caching?
The storage-hour term must appear in your break-even, and you should connect caching to prompt layout discipline, not just pricing.
Q4Grounding with Google Search, Vertex AI Search, RAG Engine, or DIY — how do you choose for a given use case?
Strong answer outline
- Google Search grounding: freshness on public facts; per-grounded-request pricing; never for private data, and a data-egress question in locked-down environments.
- Vertex AI Search: turnkey enterprise retrieval with connectors, ACL-aware results, and Google-grade ranking — the default production baseline for heterogeneous corpora.
- RAG Engine: programmable pipeline control (chunking, embedding, backend choice) without owning infrastructure — the middle rung.
- DIY on Vector Search: custom retrieval science, strict contracts, extreme scale; you own relevance and evals. Benchmark before descending rungs.
Follow-up probes
- A design needs both fresh web facts and internal policy — how do you combine rungs safely?
- What eval evidence justifies leaving Vertex AI Search for a custom pipeline?
You should present it as a ladder with descent criteria, not a feature list — and mention ACL-aware retrieval, which is what enterprises actually buy.
Q5Vertex AI Vector Search versus AlloyDB pgvector versus BigQuery vector search — build the decision framework.
Strong answer outline
- Latency class first: chat-path ANN at high QPS → Vector Search; transactional joins with entitlements → AlloyDB; batch/analytical → BigQuery.
- Data gravity second: keep vectors where their source rows and ACLs live unless scale forces a dedicated store.
- Cost shape: Vector Search has an always-on serving-replica floor; AlloyDB is a Postgres fleet you already may run; BigQuery bills per query/slot with zero standing vector infra.
- Name the hybrid: embed and evaluate in BigQuery, serve online from Vector Search or AlloyDB — one lineage, two latency classes.
Follow-up probes
- Where does the AlloyDB ScaNN index change the pgvector scale ceiling?
- At what corpus size does Vector Search stop being overkill?
Latency-class-first ordering and at least one cost-floor observation are required; a pure recall-quality argument misses the point.
Q6What is ScaNN, and what do streaming versus batch index updates mean operationally in Vector Search?
Strong answer outline
- ScaNN: Google's ANN method using anisotropic vector quantization — quantization loss weighted toward directions that affect inner-product ranking — published and benchmarked; Vector Search productizes it as tree-AH indexes.
- Brute-force index option exists for recall ground-truthing — use it to calibrate recall@k before tuning ANN parameters.
- Batch updates: cheaper full/partial rebuilds on a cadence — fine for slowly changing corpora, but stale between rebuilds.
- Streaming updates: near-real-time upserts at higher write cost — mandatory when freshness is a product requirement (inventory, tickets); the choice is business-driven, not technical taste.
Follow-up probes
- How do namespace restricts interact with recall?
- How would you detect and handle index staleness in production?
You should connect the update-mode choice to a concrete freshness requirement and mention recall calibration against brute force — that is what operating an ANN index actually looks like.
Q7A team built a LangGraph prototype. Do you move them to ADK, Agent Engine, both, or neither?
Strong answer outline
- Separate framework from runtime: Agent Engine hosts LangGraph fine — sessions, Memory Bank, tracing, and VPC-SC arrive without a rewrite.
- Rewriting to ADK is justified by platform standardization (shared eval harness, tool catalog, sub-agent patterns), not by capability necessity.
- First apply the deterministic-flow test: if the graph is static, compile it into Workflows or a plain service and delete the autonomy (chapter 05).
- Whatever runs: tool auth via a governed layer, trajectory evals pre-release, Cloud Trace wired from day one.
Follow-up probes
- What does Agent Engine give you that Cloud Run plus a session store does not?
- Where would A2A enter this picture, and where is it premature?
The framework/runtime separation must be explicit, and the deterministic test must come before any platform recommendation.
Q8What problem does A2A solve that MCP does not, and when would you refuse to adopt it?
Strong answer outline
- MCP standardizes agent-to-tool: a model invoking capabilities with structured schemas. A2A standardizes agent-to-agent: discovery via agent cards, task lifecycle, and message exchange between opaque peers.
- A2A earns its cost at organizational seams — different teams, vendors, or runtimes whose agents must interoperate without sharing internals.
- Refuse it when one team owns all agents in one runtime: in-process delegation (ADK sub-agents) is simpler, faster, and easier to trace.
- Governance note: A2A is Linux Foundation-governed with multi-vendor backing — a real standardization bet, but adoption maturity varies; pilot at one seam first.
Follow-up probes
- How do you authenticate and authorize a peer agent you did not build?
- What observability do you lose when a sub-task crosses an A2A boundary?
The tool-versus-agent seam distinction must be crisp, and you need one concrete refusal condition — enthusiasm without a boundary reads junior.
Q9Argue for or against fine-tuning Gemini for a high-volume formatting task, including serving economics — and contrast AWS.
Strong answer outline
- Order of operations: prompting + few-shot + structured output first; SFT only if the eval gap persists on a stable, narrow task (chapter 03).
- Gemini SFT mechanics: JSONL pairs, adapter-based managed tuning, versioned tuned model.
- Economics: tuned Gemini serves on shared infrastructure at base per-token rates — near-zero marginal serving cost, so break-even arrives at modest volume.
- Contrast: Bedrock custom models generally need provisioned or dedicated capacity — an always-on floor. The same tune can be economical on Vertex and uneconomical on Bedrock at identical traffic.
Follow-up probes
- What eval evidence gates the tune, and what is the rollback plan?
- What happens to your tuned model when the base model version is deprecated?
The shared-serving asymmetry is the entire point; miss it and the answer is generic chapter-03 material. Version-deprecation awareness is the bonus signal.
Q10When does GenAI belong inside BigQuery, and where is the hard boundary?
Strong answer outline
- Belongs: batch enrichment and classification over governed tables, ML.GENERATE_EMBEDDING at scale, VECTOR_SEARCH semantic joins, offline RAG evals over logged traffic — governance and lineage for free, zero data movement.
- Boundary: analytical latency class — seconds, not chat-path milliseconds; row-wise generation runs through Vertex quota and needs batch windows or PT for big jobs.
- The composite pattern: embed and evaluate in BigQuery, sync to Vector Search/AlloyDB for online serving.
- Cloud-choice note: Redshift ML has no comparable depth — this is a legitimate structural argument for GCP in analytics-heavy shops.
Follow-up probes
- How do you control cost when an analyst can trigger a million Gemini calls with one query?
- How does row-level security interact with AI functions?
You must state the latency boundary unprompted and name the guardrail problem of SQL-triggered LLM spend — both are operating-experience tells.
Q11Serve a fine-tuned 27B open model on GCP: Vertex endpoint, GKE + vLLM, or Cloud Run GPU?
Strong answer outline
- Traffic shape first: spiky/low duty cycle → Cloud Run GPU (scale-to-zero, per-second billing — a genuinely GCP-differentiated posture); steady high volume → GKE + vLLM with CUDs; middle ground with minimal ops → Vertex endpoint from Model Garden.
- Vertex endpoints bill per accelerator-hour while deployed — no scale-to-zero — so idle endpoints are the classic waste.
- GKE adds the levers: Inference Gateway prefix-aware routing, custom autoscaling on batching metrics, TPU option via vLLM — worth it only with a team to run it.
- Show the crossover as a calculation: utilization × per-hour cost versus per-token MaaS alternatives, with the honest prior that teams overestimate utilization.
Follow-up probes
- What cold-start behavior do you accept on Cloud Run GPUs and how do you mitigate it?
- When do TPUs beat GPUs for this model class?
Traffic-shape-first reasoning with the scale-to-zero distinction is required; naming all three postures without a decision rule is a catalog recital.
Q12What does VPC Service Controls give a GenAI platform that IAM and private networking do not?
Strong answer outline
- IAM answers "who may call"; private networking answers "over what path"; VPC-SC answers "where may data flow" — a service-plane perimeter that blocks exfiltration to non-perimeter projects even with valid credentials.
- Threat model: leaked service-account keys, malicious insiders, and misconfigured tools copying data to attacker-controlled projects — IAM alone stops none of these once credentials are valid.
- GenAI specifics: put Vertex AI, GCS corpora, BigQuery logs, and KMS inside one perimeter; use ingress/egress rules for the narrow, audited exceptions.
- Operational honesty: perimeters break naive integrations (SaaS webhooks, cross-project service calls) — plan dry-run mode and exception governance from the start.
Follow-up probes
- How does Grounding with Google Search interact with a strict perimeter?
- What is the AWS-equivalent conversation, and what does it lack?
The three-question framing (who/path/where) and one concrete stolen-credential scenario are the pass bar; mentioning dry-run rollout is the senior bonus.
Q13A CISO asks: "Is Google training on our prompts? Where is our data processed and who can see it?" Answer precisely.
Strong answer outline
- Training: per the Vertex Gen AI data-governance documentation, customer prompts and outputs are not used to train foundation models without permission — cite the doc, not vibes.
- Processing location: regional endpoints keep ML processing in-region; global endpoint and cross-region features change that — a deliberate opt-in with compliance sign-off.
- Retention nuances: name what you would verify — caching behavior, abuse-monitoring retention, and zero-retention configuration options for stricter regimes.
- Visibility: CMEK on stored artifacts, Admin Activity logs always on, Data Access logs opt-in (and themselves a PII surface), Access Transparency for provider-side access.
Follow-up probes
- What changes in this answer for the consumer Gemini API?
- Which of these claims would you re-verify before a contract signature, and where?
Four precise clauses with the opt-in caveats beat any confident generality. Volunteering "here is what I would re-verify in the docs" is a trust-builder, not a weakness.
Q14Your Vertex bill doubled month-over-month with flat traffic. Diagnose and fix it.
Strong answer outline
- Attribute first: usage logs in BigQuery sliced by feature/tenant/model — find whether tokens-per-request, model mix, or a new dimension (grounding calls, cache storage, long-context tier) moved.
- Usual suspects: a prompt change bloating context (long-context pricing tier crossed), implicit-cache hit rate destroyed by a prompt-layout change, Pro traffic that should be Flash, an explicit cache with storage-hours but no hits, or an idle self-hosted endpoint.
- Fixes in order: restore cache-friendly layout, route Flash-first with escalation, move offline work to batch, right-size or cancel GSU/endpoint commitments.
- Prevention: tokens-per-request alerting, per-feature budgets, and cost-per-task (not per-call) as the tracked KPI.
Follow-up probes
- How would you catch the long-context tier crossing before the invoice?
- Which of these levers risks quality regressions, and how do you gate them?
Attribution before levers, and at least four distinct pricing dimensions named. Jumping straight to "use a cheaper model" fails the diagnosis half.
Proof artifact: a two-latency-class RAG stack on GCP, costed and broken
Build one small system that produces evidence for five chapter themes: API-surface choice, grounding-ladder judgment, the two-latency-class embedding architecture, security posture, and cache/batch economics. Keep the corpus small (50–100 public documents) so the lab stays in low tens of dollars — and tear down Vector Search endpoints afterward, because that is where the idle cost lives.
- DeployLoad documents into GCS → BigQuery; compute embeddings with ML.GENERATE_EMBEDDING; build a TreeAH index in BigQuery for offline search AND sync the same vectors to a Vertex AI Vector Search streaming index; serve via Cloud Run calling streamGenerateContent with retrieved chunks.
- HardenDedicated service account with least-privilege roles; CMEK on the GCS bucket and index; enable Data Access audit logs into a locked sink; document what a VPC-SC perimeter would add (dry-run it if you have an org).
- MeasureDrive 200 scripted queries against both retrieval paths; record p50/p95 latency per path, recall overlap between BigQuery VECTOR_SEARCH and Vector Search results, tokens per answer, and cost per answered question from usage logs in BigQuery.
- Break it — quotaBurst pay-as-you-go traffic until 429s appear; show the failure without retries, then with exponential backoff and a queue; graph both. Write one paragraph on what GSU floor you would buy from the measured tokens/sec.
- Break it — stalenessUpdate 10 documents; show the streaming index reflecting changes in near-real-time while the batch-built BigQuery index serves stale results until rebuild; connect this to the freshness argument of section 4.
- Compare economicsRe-run the query set three ways: cold, with an explicit context cache on the system prompt + policy block, and as a batch prediction job. Produce a one-page table of cost per 1k answers for each mode, dated and labeled as your own measurements.
What to present in an interview: the architecture diagram with both latency classes, the staleness demonstration, and the three-mode cost table. "I measured the cache discount and the batch discount myself, on this date, and here is where each wins" is the sentence that separates you from the certification crowd.
Chapter review
You now hold GCP GenAI as three pillars — Vertex AI as the single model/agent plane, BigQuery as the governed data-and-batch-GenAI plane, GKE/Cloud Run as the self-managed compute plane — crossed by a security model whose signature move is the VPC-SC perimeter. The recurring senior patterns: buy Gemini capacity deliberately against DSQ's non-guarantee (Figure 1), stream with caching discipline (Figure 2), descend the grounding ladder only on eval evidence (Figure 3), give agents a managed runtime and refuse unneeded agency (Figure 4), run batch GenAI where the data lives (Figure 5), and draw the perimeter-protected invocation path by default (Figure 6).
- Dynamic Shared Quota
- Pay-as-you-go Gemini capacity from a shared regional pool — elastic, but no per-project guarantee; 429s are a design input.
- Provisioned Throughput / GSU
- Committed Gemini capacity purchased in generative AI scale units per model and term; overage spills to DSQ.
- Context caching
- Implicit prefix-hit discounts by default; explicit cachedContent objects trade a per-token-hour storage fee for steep per-hit discounts.
- Model Garden
- Vertex catalog spanning Gemini, partner models-as-a-service (including Claude), and deployable open models.
- Vertex AI Search
- Turnkey enterprise retrieval: connectors, parsing, ranking, ACL-aware results; the default managed RAG baseline.
- RAG Engine
- Managed, programmable RAG framework with pluggable vector backends — the middle rung between turnkey and DIY.
- ScaNN / tree-AH
- Google's anisotropic-quantization ANN lineage behind Vertex AI Vector Search, AlloyDB's ScaNN index, and BigQuery's TreeAH.
- ADK
- Open-source Agent Development Kit: agents, tools, callbacks, multi-agent hierarchies, built-in evals; deploys anywhere.
- Agent Engine
- Managed agent runtime on Vertex: sessions, Memory Bank, scaling, tracing; hosts ADK and other frameworks.
- A2A
- Open, Linux Foundation-governed agent-to-agent protocol — discovery via agent cards and task exchange across runtimes.
- VPC Service Controls
- Service-plane perimeter blocking data exfiltration to non-perimeter projects even with valid credentials.
- Model Armor
- Model-independent prompt/response screening for injection, jailbreaks, and sensitive data — GCP's guardrails layer.
- I can explain Gemini API versus Vertex AI Gemini with three concrete enterprise controls and the SDK portability fact.
- I can reason about DSQ versus GSU commitments from measured tokens/sec, and design for 429s.
- I can compute a context-caching break-even including the storage-hour term, and protect implicit caching with prompt layout.
- I can place a use case on the grounding ladder and state the descent criteria.
- I can defend a vector-store choice by latency class, data gravity, and cost floor — and name the two-latency-class hybrid.
- I can separate ADK from Agent Engine from A2A, and refuse agency for deterministic flows.
- I can state the tuned-Gemini shared-serving asymmetry against Bedrock with its break-even consequence.
- I can deliver the four-clause data-governance answer with its opt-in caveats, and draw the VPC-SC invocation path.
- I can cut a Vertex bill with attribution-first diagnosis across five pricing dimensions.
Primary sources
Links checked 2026-08-04.
- Generative AI on Vertex AI — documentation
- Gemini Developer API — documentation
- Vertex AI Generative AI — quotas and Dynamic Shared Quota
- Vertex AI — Provisioned Throughput
- Vertex AI — Context caching
- Vertex AI — Batch prediction for Gemini
- Vertex AI Model Garden
- Vertex AI — Grounding overview
- Vertex AI Search — introduction
- Vertex AI RAG Engine — overview
- Vertex AI Vector Search — overview
- Guo et al. — Accelerating Large-Scale Inference with Anisotropic Vector Quantization (ScaNN, arXiv:1908.10396)
- AlloyDB AI — documentation
- BigQuery — vector search introduction
- BigQuery — generative AI overview
- BigQuery ML — introduction
- Agent Development Kit (ADK) — documentation
- Vertex AI Agent Engine — overview
- A2A protocol — specification repository
- Vertex AI — Gemini model tuning overview
- GKE — AI/ML orchestration documentation
- Cloud Run — GPU configuration
- VPC Service Controls — overview
- Cloud KMS — customer-managed encryption keys
- Generative AI on Vertex AI — data governance
- Cloud Audit Logs — documentation
- Model Armor — overview
- Cloud Logging — documentation
- Cloud Monitoring — documentation
- Cloud Trace — documentation
- Vertex AI — Generative AI pricing
- Google Cloud — committed use discounts
- Amazon Bedrock — User Guide (for the AWS mapping)
CHAPTER 10 · PRIORITY 1
Data Systems, Cloud & Platform Engineering
22 min read · 14 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- design a PostgreSQL schema and index strategy from access patterns, consistency needs, and tenant boundaries;
- read an execution plan, distinguish estimates from measurements, and improve a slow query with evidence;
- build a replayable ingestion pipeline for malformed documents with checkpoints, validation, and lineage;
- package and operate an AI service using safe container images, Kubernetes probes, resources, autoscaling, and rollout controls;
- explain networking, identity, infrastructure-as-code, recovery, and cost as one platform design; and
- size and defend a reference document-processing platform without pretending illustrative estimates are production facts.
1. PostgreSQL: begin with invariants and access paths
A senior data answer starts with what must remain true under concurrency. Tables, indexes, and transactions are mechanisms for those invariants—not independent checklist items.
Model stable facts, preserve uncertain input
Normalize entities that have independent identity and lifecycle: tenant, source, document, document version, processing run, chunk, and access grant. Use foreign keys and unique constraints for truths the database can enforce. Keep the immutable source object in object storage and a raw metadata reference or carefully bounded JSONB column for provider-specific fields. Promoting every uncertain field into a column creates migration churn; placing every stable relationship in JSON discards relational guarantees and makes query behavior harder to predict.
document(
tenant_id, document_id, source_id, external_id,
current_version_id, lifecycle_state, created_at, updated_at,
UNIQUE (tenant_id, source_id, external_id)
)
document_version(
tenant_id, version_id, document_id, content_hash,
object_uri, parser_version, source_modified_at, status,
UNIQUE (tenant_id, document_id, content_hash)
)
Put tenant_id in ownership and uniqueness keys, not only in a nullable filter. Row-level security can add defense in depth: when enabled, normal access must be allowed by a policy, as the current PostgreSQL row-security documentation explains. Still test connection-pool session state, roles that bypass RLS, background jobs, migrations, and administrative access. RLS is not a substitute for explicit tenant-aware application APIs.
Index for a concrete query
| Access pattern | Candidate | Trade-off to mention |
|---|---|---|
| Tenant’s recent failed runs | B-tree on (tenant_id, status, created_at DESC), perhaps partial on failures | Writes and storage increase; column order must match predicates and ordering. |
| Lookup by source object | Unique B-tree on (tenant_id, source_id, external_id) | Enforces deduplication as well as speeding lookup. |
| Containment in selected JSON metadata | GIN on the queried JSONB path/operator class | A wide generic GIN index can be large and write-expensive. |
| Lexical document search | Generated tsvector plus GIN | Language configuration and ranking must match the corpus. |
| Vector nearest neighbors | pgvector exact or approximate index, filtered by tenant/ACL strategy | Recall, build time, memory, filtering, and update behavior must be measured. |
“Add an index” is not a diagnosis. Capture the representative query and parameters, table/index sizes, data distribution, concurrency, cache state, and latency percentiles. Use EXPLAIN (ANALYZE, BUFFERS, FORMAT JSON) in a safe environment: ANALYZE executes the query, including writes. Compare estimated versus actual rows at each node, loops, scan type, join algorithm, sorts/spills, heap fetches, and shared reads/hits. PostgreSQL’s official guide emphasizes that a plan is a tree and estimates depend on statistics: Using EXPLAIN.
Transactions, locks, and pools
PostgreSQL defaults to Read Committed, where each statement sees a snapshot at statement start. Repeatable Read provides a stable transaction view; Serializable detects executions that cannot be ordered safely and requires the application to retry serialization failures. The exact behavior is documented in Transaction Isolation. Choose by invariant, keep transactions short, update resources in a consistent order, inspect lock waits, and make retries idempotent. A deadlock victim is expected safety behavior, not evidence the database is broken.
Connection pools protect a finite database resource. Size them from database capacity across all replicas/workers, not from request count. Long model calls must not hold an open transaction or idle connection. Use a pool acquisition timeout, statement timeout, transaction timeout, and metrics for active, idle, waiting, and age. Pool exhaustion can look like a slow query even when no query has started.
2. Build ingestion as a replayable ledger
Messy enterprise data means malformed files, duplicate exports, password-protected PDFs, changing parsers, conflicting encodings, and tables whose visual layout carries meaning. A trustworthy pipeline never makes the only copy of the input its latest derived output.
Stages and contracts
flowchart LR
D["Discover sources"] --> L["Land immutable bytes"]
L --> V["Validate and quarantine"]
V --> E["Extract and OCR"]
E --> N["Normalize schema"]
N --> C["Chunk and enrich"]
C --> I["Index generation"]
I --> P["Publish manifest"]
P --> R["Serve retrieval"]
R -->|"replay request"| L
- Discover: enumerate source objects with a stable source ID, version/etag, modified time, and permission snapshot. Do not assume directory listing order is a checkpoint.
- Land: stream bytes to immutable object storage while computing size and cryptographic content hash. Enforce file and decompression limits before expensive parsing.
- Validate: detect actual media type, malware policy result, encryption, structural damage, and tenant/source authorization. Quarantine with a reason; do not silently skip.
- Extract: select a versioned parser by media type. Preserve page, sheet, cell, bounding box, and OCR confidence so answers can cite the source.
- Normalize: produce a canonical document representation. Keep extraction warnings and unknowns; do not fabricate clean text from missing data.
- Publish: atomically point the document at a complete index generation only after validation. Old generation remains available for rollback.
ETL transforms before loading into the analytical destination and is useful when strict sanitization or a stable target schema is required. ELT lands raw data first and transforms in a capable data platform, improving reprocessing and exploratory flexibility. A robust document pipeline often combines them: immutable landing is ELT-like; security validation before broad availability is ETL-like.
Batch, stream, checkpoints, and deduplication
Streaming lowers freshness but increases state and operational complexity. Batch improves throughput and gives clean processing boundaries. Use events for fast discovery and a scheduled inventory for correctness. A checkpoint must name a durable completed boundary: source cursor plus tie-breaker, manifest ID, input partition, and pipeline version. Commit it only after all outputs for that boundary are durable.
Deduplicate at several levels: source object ID/version prevents repeat discovery, content hash detects identical bytes across exports, and deterministic derived IDs prevent duplicate chunks. Whether two tenants may share physical bytes is a security and encryption decision; logical documents and access controls remain separate. Do not use fuzzy text similarity as the only identity rule.
Lineage and data-quality gates
For every indexed chunk, answer: which tenant, source object, source version, byte hash, parser/OCR version, normalization version, page/region, chunker configuration, embedding model/version, processing run, and access-control snapshot produced it? That lineage makes a targeted replay possible when a parser defect affects only scanned PDFs.
Quality gates should distinguish hard failures from warnings: empty extraction, implausible page counts, unreadable percentage, duplicated pages, corrupted tables, missing mandatory columns, OCR confidence distribution, and access-control mismatch. Route ambiguous cases to human review with the original rendering and extracted representation side by side.
3. Containers and Kubernetes: make runtime intent explicit
Kubernetes can restart and route around failures only when the application exposes truthful health and resource behavior. Deployment YAML cannot compensate for an endpoint that lies.
Build a small, reproducible image
Use a multi-stage build so compilers and build caches do not enter the runtime image; Docker’s official guide shows how stages selectively copy artifacts: multi-stage builds. Pin base images by an intentional version or digest, run as a non-root user, use a read-only filesystem where possible, exclude secrets and build context, emit an SBOM/signature in the supply-chain workflow, and scan both dependencies and the final image. Rebuild images for patches; do not mutate running containers.
# Illustrative structure; pin reviewed versions/digests in a real build.
FROM python:3.13-slim AS build
WORKDIR /build
COPY pyproject.toml uv.lock ./
RUN ... build a locked wheelhouse ...
FROM python:3.13-slim AS runtime
RUN useradd --system --uid 10001 app
COPY --from=build /build/wheels /wheels
RUN ... install only locked runtime wheels ...
USER 10001
CMD ["python", "-m", "service"]
Three probes, three questions
- Startup: has this slow-starting process initialized enough for other probes to begin?
- Readiness: should this pod receive new traffic now? Overload or loss of a mandatory local capability may make it unready.
- Liveness: is the process irrecoverably stuck such that restart is likely to help?
Kubernetes suppresses readiness/liveness until a configured startup probe succeeds, and a failed readiness probe removes the pod from service endpoints. Its documentation also warns that bad liveness probes can cause cascading restarts under load: probe guidance. Do not make liveness depend on every remote model provider; restarting healthy pods during a provider outage increases damage.
Resources, scaling, and rollout
CPU requests influence scheduling and CPU limits can throttle; memory limits can end in OOM termination. Start from load-test profiles for API, parser, and worker separately. Observe working set, CPU throttling, garbage collection, queue age, and latency. Keep headroom for bursts and node disruption. A HorizontalPodAutoscaler adjusts replicas from observed metrics, but scaling on CPU alone can fail for I/O-bound queue workers; queue age or outstanding work per ready worker is often closer to user pain. See Kubernetes autoscaling concepts.
A rolling deployment needs enough surge capacity, readiness that reflects warm-up, termination grace, request/lease draining, and a disruption budget. Canary releases route a small, observable cohort to the new parser/model/service and compare errors, latency, cost, and quality before expansion. Blue-green gives a clean environment switch and fast rollback but doubles capacity during transition and does not magically reverse database migrations. Prefer expand/migrate/contract schemas and backward-compatible readers.
4. Cloud platform design: identity, network, delivery, and recovery
A reference architecture should explain traffic and trust, not display a cloud catalog. Trace both the data plane and the control plane.
One request from DNS to data
DNS resolves the service name; TLS authenticates the endpoint and encrypts transport; a load balancer or ingress applies routing and connection policy; the service authenticates the caller; workload identity authorizes narrowly scoped calls to database, object storage, queue, secret manager, and model provider. Network policy and private endpoints reduce paths but do not replace identity checks. Timeouts must descend: client deadline > ingress > service subcalls, leaving time to return a useful error.
Prefer short-lived workload identity to static cloud keys. Separate deployer identity from runtime identity. A worker reading source objects need not mutate infrastructure; an API serving job status need not decrypt every source credential. Centralize secrets in a managed store, rotate them, and record access without printing values.
Infrastructure as code and CI/CD
Terraform describes desired resources and records bindings in state. State can contain sensitive information and must be shared and locked safely; HashiCorp recommends remote state for team use and warns against insecure version-control storage: Terraform state documentation. Pin provider/module versions, review a saved plan, use separate state blast radii, run policy/security tests, and avoid broad production credentials in pull-request jobs.
- Build once; produce immutable image digest, tests, SBOM, vulnerability result, and provenance.
- Plan infrastructure and schema changes; require review for destructive or privilege-expanding operations.
- Deploy to a representative pre-production environment; run smoke, contract, migration, and rollback tests.
- Canary by environment, tenant cohort, or traffic; automatically pause on user-facing SLO and correctness signals.
- Promote the same artifact, then verify and retain evidence. Rollback or roll forward through an exercised procedure.
Public cloud favors managed-service velocity and elastic capacity; private or customer-hosted deployments may be required for data residency, network control, or procurement policy, but increase version skew, upgrade, capacity, and support burdens. Hybrid design needs an explicit connectivity failure mode and a support boundary.
Backups are not recovery
Define recovery point objective (maximum acceptable data loss) and recovery time objective (maximum acceptable restoration time) by component. Test database point-in-time recovery, object versioning, queue redrive, index rebuild from canonical data, infrastructure recreation, identity/secrets restoration, and DNS failover. Record actual drill times. A replica can copy corruption; a backup can be unusable; an index may be cheaper and safer to rebuild than back up.
5. Worked system: multi-tenant document processing platform
This hypothetical example turns the prior decisions into a system-design narrative. The numbers are illustrative sizing assumptions, not production results.
Estimate before selecting capacity
Assume 100 tenants, 50,000 documents each, four pages per document: 20 million pages initially. If 0.5% change daily, steady state is about 100,000 pages/day, but an onboarding backfill may be 50 times the average. If OCR consumes an example 1.5 CPU-seconds/page, the daily steady-state CPU work is about 42 CPU-hours; a 10-hour processing objective needs roughly 4.2 continuously busy cores before concurrency inefficiency, retries, and headroom. Benchmark the real corpus before committing.
Store original bytes and manifests in object storage, canonical metadata/checkpoints in PostgreSQL, and tasks in a durable queue. Separate fetch, virus/format validation, OCR/extraction, normalization/chunking, embedding, and index-publish workers so each scales and retries independently. Use deterministic artifact keys and an index generation pointer. Interactive query traffic runs in a different deployment and resource pool from ingestion.
Capacity and cost model
| Driver | Simple estimate | Control |
|---|---|---|
| Raw storage | source bytes + versions + retention | lifecycle tiers, deletion policy, dedupe only where isolation permits |
| OCR/parse compute | pages × seconds/page × retry factor | format routing, bounded retries, spot/preemptible only with checkpoints |
| Embeddings | changed chunks × tokens/chunk × provider price | content hashes, incremental updates, batch API where suitable |
| Vector index | vectors × dimensions × bytes plus graph/index overhead | measure compression/recall, retention, tenant placement |
| Egress | cross-region/cloud bytes | co-locate stages, compress, model residency deliberately |
Protect fairness with per-tenant quotas and weighted scheduling. Large backfills consume a separate budget and can pause when interactive latency or database saturation rises. Operational dashboards join queue age, stage throughput, success/warning/quarantine rates, CPU/memory, database waits, provider usage, cost per usable page, and freshness by tenant.
Deploy in one region first if requirements allow, using multi-zone managed services. Keep raw and canonical data sufficient to rebuild derived indexes. For regional disaster recovery, choose active-passive unless the recovery objective justifies active-active data consistency and operational complexity. Explain how tenant data residency changes placement and how the control plane routes a tenant to the correct deployment stamp.
Syllabus checkpoint: database, Kubernetes, Linux, and cloud breadth
PostgreSQL beyond one slow query
Schema design starts from invariants, update patterns, and query grain; normalization reduces contradictory facts, while deliberate denormalization needs an ownership and refresh rule. PostgreSQL full-text search uses document/query representations, dictionaries, ranking, and indexes and can complement pgvector for hybrid retrieval. Test migrations forward and backward against production-like data, and test backups by restoring them to a clean environment.
Configuration and delivery on Kubernetes
ConfigMaps hold non-secret configuration; Secrets are transport/storage objects whose encryption, RBAC, rotation, and workload delivery still need design. Prefer workload identity to long-lived cloud keys. Helm packages parameterized Kubernetes resources; GitOps reconciles declared state through reviewed changes. Neither replaces readiness tests, safe database migration order, canary or blue-green analysis, termination/draining, nor a verified rollback.
Linux and networking diagnosis
Trace a request through DNS resolution, TCP connection, TLS handshake, proxy/load balancer, service routing, application, and downstream dependency. On Linux, inspect process state, sockets, CPU, memory, disk and inode pressure, file descriptors, cgroups, logs, and permissions before changing configuration. Distinguish connection timeout, refusal, reset, TLS verification, and application timeout; they imply different layers and owners.
AWS and GCP mapping
Map requirements to durable primitives before provider names: identity/IAM, network boundary, compute, object storage, queue/event service, managed PostgreSQL, Kubernetes/serverless, secrets/KMS, monitoring, and audit. AWS and GCP differ in service mechanics and defaults, so validate the chosen managed service’s quotas, availability model, backup/restore, private networking, egress, and pricing. Terraform should produce reviewed, repeatable state with remote locking, least-privilege credentials, drift detection, and a recovery plan for state—not merely create resources.
Interview playbook
Use QUERY → PIPELINE → PLATFORM:
- Query: identify invariants, access patterns, scale, distribution, and transaction boundary; sketch keys before indexes.
- Pipeline: show immutable input, versioned stages, idempotent IDs, checkpoint, quarantine, lineage, and replay.
- Platform: trace network and identity, stateful dependencies, probes/resources, scaling metric, rollout/rollback, recovery, and cost.
For a slow-query question, ask for plan and data distribution before proposing an index. For Kubernetes, do not confuse liveness with readiness or autoscaling with capacity. For cloud architecture, name RPO/RTO and trust boundaries. State illustrative numbers as assumptions, show the arithmetic, and explain what benchmark would replace them.
Common traps include using JSONB for every field, running EXPLAIN ANALYZE on a dangerous production write, holding a database transaction across OCR/model calls, committing a checkpoint before outputs, making liveness depend on a remote provider, setting resources by guesswork, scaling workers until the database fails, and claiming a backup without a restore test.
Question bank
Answer with one concrete workload and measurable verification.
Q1How do you diagnose a PostgreSQL query that became slow for only one tenant?
Strong answer outline
- Capture normalized query, tenant-safe parameters, latency distribution, plan, table/index sizes, locks, and pool wait.
- Compare estimated/actual rows and data skew; inspect scans, loops, sorts/spills, and buffers.
- Test query/index/statistics changes against small and large tenants, including write cost and regression.
Follow-up probes
- Why might the generic plan be poor?
- How can extended statistics help?
You diagnosed estimation, execution, and waiting—not merely “add an index.”
Q2When would you use JSONB instead of normalized columns?
Strong answer outline
- Use JSONB for sparse/provider-specific or preserved raw metadata with evolving shape.
- Use columns/tables for identity, constraints, joins, frequently filtered fields, and independent lifecycle.
- Index only demonstrated JSON paths/operators and validate document size/update cost.
Follow-up probes
- How do you migrate a JSON field into a column?
- What does a GIN index cost?
You balanced schema agility with integrity and query predictability.
Q3Choose an isolation level for claiming queue jobs from PostgreSQL.
Strong answer outline
- Define invariant: one active lease per job while abandoned leases can be reclaimed.
- Use a short transaction with row locking such as
FOR UPDATE SKIP LOCKED, atomically setting lease owner/expiry. - Make job effects idempotent; handle lease expiry and deadlocks/serialization errors with bounded retry.
Follow-up probes
- Why not hold the transaction while processing?
- What if a worker outlives its lease?
You protected the invariant without a long transaction and addressed fencing/duplicate work.
Q4How can row-level security still fail to protect tenants?
Strong answer outline
- Policies may be absent/wrong, table owners or privileged roles may bypass them, or pool session context may leak.
- Use least-privilege runtime roles, transaction-local tenant context, composite constraints, and explicit repository filters.
- Test cross-tenant reads/writes, background/admin paths, migrations, and cache/index isolation.
Follow-up probes
- How do you test a connection pool?
- Can RLS protect object storage?
You treated RLS as defense in depth and covered non-database paths.
Q5Design a checkpoint for a paginated document source.
Strong answer outline
- Persist source/version, cursor or ordered high-water key with tie-breaker, run/version, and completed manifest.
- Apply a page and derived task creation atomically before advancing the checkpoint.
- On resume, overlap when source semantics are weak and rely on deterministic IDs; reconcile full inventory periodically.
Follow-up probes
- What if the cursor expires?
- How are deletions discovered?
Your checkpoint denotes durable output, not “last item fetched.”
Q6How do you process a malformed 5 GB archive safely?
Strong answer outline
- Stream with request/object size, entry count, path, compression-ratio, nesting, and total-expanded-byte limits.
- Validate media type, isolate parsing with CPU/memory/time budgets, and never trust archive paths.
- Quarantine immutable input and structured reason; do not partially publish derived content.
Follow-up probes
- How do you avoid a zip bomb?
- What is safe to expose to an operator?
You bounded resource use, contained parsing, and preserved evidence without publishing unsafe output.
Q7Batch or streaming ingestion for enterprise documents?
Strong answer outline
- Derive from freshness, volume/burst, source capabilities, ordering, and recovery objectives.
- Use events for low-latency discovery and micro-batches/work queues for efficient processing.
- Retain scheduled inventory/reconciliation because streams can be delayed, duplicated, or missed.
Follow-up probes
- When is a daily batch enough?
- How does backpressure change freshness?
You offered a hybrid correctness path and quantified the latency/complexity trade-off.
Q8What lineage is required to remove output from a defective parser version?
Strong answer outline
- Map every chunk/index record to tenant, source/version/hash, page/region, parser and normalization versions, run, and ACL snapshot.
- Query affected artifacts, rerun only their immutable inputs with a fixed pipeline, and publish a new generation.
- Compare quality and counts, switch the pointer, retain rollback, then retire defective artifacts.
Follow-up probes
- How does an embedding-model change differ?
- What if the original is deleted by retention policy?
Your lineage supports bounded impact analysis, replay, comparison, and rollback.
Q9Design readiness and liveness for an AI API.
Strong answer outline
- Startup protects initialization; readiness checks local ability to admit work and critical warmed state.
- Liveness detects unrecoverable process deadlock, not availability of every model provider.
- Keep probes cheap with separate budgets; test overload, provider outage, shutdown, and cold start.
Follow-up probes
- Should database loss make the pod unready?
- What creates a restart storm?
You connected each probe to the controller action and cascading-failure risk.
Q10How do you choose CPU and memory requests and limits?
Strong answer outline
- Measure representative load by workload class, including peaks, initialization, and parser/model behavior.
- Set requests for reliable scheduling and limits with awareness of CPU throttling and memory OOM behavior.
- Observe saturation/throttling/OOM/latency, reserve disruption headroom, and tune iteratively.
Follow-up probes
- Why separate API and OCR workers?
- What happens when every pod uses its full request?
You used evidence, differentiated resource semantics, and planned cluster capacity.
Q11What metric should autoscale an ingestion worker?
Strong answer outline
- Use user-aligned backlog age or outstanding weighted work per ready worker, not only CPU.
- Account for stage cost, downstream database/provider capacity, startup time, and maximum safe concurrency.
- Set scale-down stabilization and test bursts, poison jobs, and dependency degradation.
Follow-up probes
- How do long and short jobs distort queue depth?
- Why can scaling worsen an outage?
You selected a causal metric and capped scaling at system—not cluster—capacity.
Q12How do you roll out a parser plus database schema change?
Strong answer outline
- Expand schema compatibly; deploy readers/writers that handle old and new; backfill with checkpoints.
- Canary parser by document cohort into a new generation and compare quality, errors, latency, and cost.
- Switch publication pointer, retain rollback, then contract schema only after old code/artifacts are gone.
Follow-up probes
- What cannot be rolled back?
- How do you validate OCR quality automatically?
You separated code, data, and derived-index rollback units.
Q13How would you secure Terraform state and production delivery?
Strong answer outline
- Use encrypted remote backend, access control, locking/versioning, backups, audit, and separated state blast radii.
- Use short-lived CI identity, pinned providers/modules, reviewed saved plans, and policy checks.
- Avoid secret values where possible, restrict state readers, and test state/recovery procedures.
Follow-up probes
- Why can a “sensitive” output still be in state?
- How do concurrent applies fail?
You recognized state as sensitive operational data, not a harmless build artifact.
Q14Give a cost estimate for a document pipeline with incomplete information.
Strong answer outline
- Declare ranges for documents/pages, churn, format mix, retention, OCR seconds, chunks/tokens, vector dimensions, and traffic geography.
- Calculate storage, compute, model/embedding, database/index, egress, observability, and redundancy separately.
- Show sensitivity and peak capacity, label assumptions, then propose a corpus benchmark and billing telemetry to replace them.
Follow-up probes
- Which variable dominates?
- How do enterprise isolation requirements change cost?
Your estimate is auditable, range-based, and tied to a measurement plan.
Proof artifact: operable document platform
Build a local or low-cost reference deployment. Any thresholds are example lab objectives, not claims about past work.
- Generate a synthetic tenant-safe corpus containing clean text PDFs, scans, malformed PDFs, CSV/Excel edge cases, duplicates, deletes, and deliberate cross-tenant IDs.
- Implement immutable landing, manifest/checkpoint tables, versioned extraction, quarantine, deterministic chunk IDs, lineage, and a mock index-generation switch.
- Create one deliberately slow PostgreSQL workload. Save schema, data generator, query,
EXPLAIN (ANALYZE, BUFFERS, FORMAT JSON), change hypothesis, new plan, and load-test comparison. - Containerize stages with a multi-stage non-root image. Deploy API and workers to a local Kubernetes cluster with startup/readiness/liveness, requests/limits, HPA or event-based scaling, graceful shutdown, and a rollback command.
- Describe infrastructure in Terraform for a disposable environment or use a safe mock plan; keep state outside version control and document identity boundaries.
Measure: usable pages per minute, p95 stage latency, queue oldest age, quarantine/warning rate, checkpoint recovery time, duplicate suppression, database plan rows/buffers/time, CPU throttling, memory peak/OOM, rollout error rate, estimated cost per 1,000 usable pages, and restore/rebuild time.
Inject failures: kill a worker between artifact write and checkpoint, corrupt a file, force OCR timeout, exhaust the database pool, introduce one tenant with a huge backfill, fail readiness, OOM a parser under a limit, interrupt a rollout, and rebuild the index from canonical state. Prove that a poison document does not block its partition and a tenant cannot read another tenant’s manifest.
Present: a scale worksheet, data/lineage diagram, before/after plan with reasoning, Kubernetes manifest excerpt, one failure timeline, recovery evidence, cost sensitivity chart, and an architecture decision record for batch versus streaming or shared versus deployment-stamp isolation.
Chapter review
An operable AI data platform preserves raw truth, makes transformations versioned and replayable, enforces database invariants, exposes honest runtime health, and treats delivery, recovery, and cost as design inputs.
Glossary
- Execution plan
- The planner’s tree of scan, join, sort, and aggregation operations, with estimated or measured work.
- Isolation level
- The visibility and anomaly guarantees a transaction receives under concurrency.
- Lineage
- The trace from a derived artifact back through versions, transformations, and source input.
- Manifest
- A durable inventory of inputs and outputs for one processing boundary or generation.
- Quarantine
- An isolated state for unsafe or invalid input that preserves evidence and prevents publication.
- Readiness
- Whether a workload should receive new traffic now; distinct from whether its process should restart.
- RPO / RTO
- Maximum acceptable data loss and maximum acceptable restoration time.
- Workload identity
- A short-lived identity assigned to running software for authorized service access.
Mastery checklist
- I can derive schema keys and indexes from invariants and access patterns.
- I can interpret estimated versus actual rows, loops, buffers, waits, and pool pressure.
- I can resume an ingestion run without duplicate publication and trace every chunk to source.
- I can explain probe controller actions and demonstrate graceful shutdown.
- I can choose a scaling signal while protecting downstream capacity and tenant fairness.
- I can separate application, schema, artifact-generation, and infrastructure rollback.
- I can state RPO/RTO and show a tested restore or rebuild path.
- I can produce a cost range with assumptions and sensitivity rather than a false-precision total.
Primary sources
Links checked 2026-08-04.
- PostgreSQL 18 — Using EXPLAIN, Transaction Isolation, and Row Security Policies
- Docker Docs — Multi-stage builds
- Kubernetes — Liveness, Readiness, and Startup Probes, resource management, and autoscaling workloads
- HashiCorp Terraform — State and state storage and locking
- Google Cloud Well-Architected Framework
CHAPTER 11 · PRIORITY 1
LLMOps, Reliability, Observability & Security
21 min read · 14 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- define user-centered SLIs, SLOs, and an error-budget policy for an AI application;
- design logs, metrics, and traces that preserve correlation without leaking sensitive data or exploding cardinality;
- combine deadlines, retries, backpressure, bulkheads, load shedding, graceful degradation, and recovery without creating retry storms;
- threat-model prompt injection, tool abuse, retrieval poisoning, data exfiltration, and cross-tenant access as system risks;
- run a disciplined incident from detection through mitigation, evidence-based root cause, corrective action, and learning; and
- produce a failure drill and security artifact that demonstrates production ownership rather than theoretical awareness.
1. Reliability is a user-visible contract
“The pods were up” is not a product outcome. A user needs an authorized, sufficiently correct answer or a truthful, recoverable failure within an acceptable time. Define reliability at that boundary.
SLI, SLO, SLA, and error budget
- A service-level indicator (SLI) is a measured ratio or distribution, such as valid successful assistant requests divided by eligible requests.
- A service-level objective (SLO) is a target over a window, such as an illustrative 99.5% of eligible requests producing a valid response within a specified latency threshold over 28 days.
- A service-level agreement (SLA) is a business/legal commitment and may use different definitions or consequences.
- An error budget is the allowed unreliability: for a 99.5% example SLO, 0.5% of eligible events. It becomes useful only when a written policy changes rollout and engineering decisions.
Google’s SRE guidance emphasizes user-centered targets and organizational backing for error-budget consequences: The Art of SLOs. Do not set an SLO from aspiration alone. Examine user tolerance, dependency capability, cost, historical performance, and what action the team will take when the budget burns.
Define eligibility and “good” precisely
| User journey | Candidate SLI | Important exclusions or dimensions |
|---|---|---|
| Interactive answer | valid, policy-compliant response within latency threshold / eligible requests | Separate user cancellations and invalid auth; slice by tenant tier, region, model route, and request class. |
| Tool action | confirmed correct terminal outcome / accepted actions | Do not count “model emitted a tool call” as success; distinguish denied, cancelled, compensated, and uncertain. |
| Document freshness | documents published within freshness target / changed eligible documents | Slice by source, format, tenant, and quarantine reason. |
| Retrieval quality | evaluated queries meeting grounded-answer threshold / sampled eligible queries | Quality labels arrive slowly; stratify and report uncertainty rather than hiding it in availability. |
Use request-based SLIs for interactive traffic and window/backlog-age SLIs for pipelines. For latency, a histogram or event distribution preserves tail behavior; an average can stay healthy while one customer cohort suffers. Define the measurement point and denominator so client disconnects, policy denials, and dependency timeouts cannot be reclassified opportunistically.
Burn rate turns a monthly target into an alert
Burn rate is observed bad-event rate divided by allowed bad-event rate. A service consuming one day’s budget each day burns at 1×. Multi-window alerts combine a short window that detects fast incidents with a longer window that filters transient noise. Page on budget-threatening user impact; ticket on slower trends; dashboard everything else. Exact thresholds depend on the SLO and response model.
2. Observability: connect symptoms to causes
Monitoring asks known questions; observability lets an engineer investigate unanticipated states from system outputs. Instrument a coherent event model, not three disconnected vendors.
Logs, metrics, traces, and context
- Metrics aggregate rates, errors, durations, saturation, queue age, token/cost use, and quality samples cheaply enough for dashboards and alerts.
- Traces show causality and time across ingress, retrieval, model calls, tools, databases, queues, and policy checks.
- Logs explain discrete state transitions and diagnostics with structured fields.
- Baggage/context carries selected correlation metadata across process boundaries; it must be size-bounded and must not carry secrets or raw PII.
OpenTelemetry currently defines traces, metrics, logs, and baggage as supported signals: OpenTelemetry signals. Use one resource identity (service.name, version, environment, region), propagate W3C trace context through HTTP and message metadata, and retain application-level IDs such as job or conversation ID in a privacy-safe form.
Trace an AI request without logging the world
flowchart LR
Q["User request"] --> A["Auth and quota"]
A --> R["Retrieval path"]
R --> S["Search and ACL filter"]
S --> M["Model generation"]
M --> P["Policy and output checks"]
P --> T["Tool proposal"]
T --> H["Approval and execution"]
H --> X["Response and citations"]
Record duration, status/error class, model/provider route, token counts, cache outcome, retrieval counts/scores, tool name and outcome, approval decision, validation result, and policy version. Default to hashes, classifications, counts, and references rather than prompts, retrieved text, model output, secrets, or tool arguments. Provide an explicitly authorized, short-retention diagnostic mode for rare cases, with access logs and redaction.
For queued work, the producer span ends before the consumer starts; propagate context in the message and consider a span link when work is batched or fan-outs merge. Trace the retry attempt separately while keeping a shared logical operation ID. Otherwise a three-attempt dependency call looks like one slow span and hides amplification.
Cardinality, sampling, and useful dashboards
Metric label values must remain bounded. model_route or normalized error_class may be useful; user_id, prompt text, document ID, URL, or exception message can create unbounded series and cost. Keep high-cardinality correlation in traces/logs under access control. Use histograms for latency rather than client-side percentile labels; review the official Prometheus histogram guidance for aggregation trade-offs.
Head sampling decides before a trace finishes and is cheap but may miss rare failures. Tail sampling can retain errors, high latency, or important cohorts after observing the trace but requires buffering and collector capacity. OpenTelemetry documents these choices in Sampling. Preserve enough unbiased baseline traffic to estimate rates; an errors-only trace store cannot reveal how unusual an error path is.
A service dashboard should follow user journey → dependencies → resources: SLO and burn, traffic, error classes, latency, quality/freshness, queue age, model/tool outcomes, saturation, and deployment annotations. An alert must say what user promise is threatened, the affected scope, likely first checks, and runbook; if no one should act now, it is not a page.
3. Resilience is a coordinated control loop
Timeouts, retries, breakers, queues, and fallbacks interact. Configure them as a budgeted system; independent defaults often turn one slow provider into a fleet-wide outage.
Deadlines first, then retry
A deadline is the total time the caller is willing to wait. Each downstream timeout must fit inside the remaining deadline and leave time for cleanup or a useful response. Connect, request, stream-idle, and pool-acquisition timeouts protect different waits. Propagate cancellation so abandoned work does not continue spending tokens and database capacity.
Retry only errors likely to improve on another attempt, only when the operation is safe or idempotent, with exponential backoff and full jitter, a capped attempt count, and a shared retry budget. Honor provider rate-limit guidance. If three service layers each retry three times, the lowest dependency may see up to 27 attempts for one user request; choose one owner for retries or coordinate them.
Protection patterns and their limits
| Control | Protects against | Failure when misused |
|---|---|---|
| Circuit breaker | Repeated calls to a dependency known to be failing | Global breaker hides healthy regions/tenants; probe storms occur in half-open state. |
| Bulkhead | One dependency or tenant consuming all shared resources | Partitions are too small or unused capacity cannot be borrowed safely. |
| Backpressure | Producers outpacing consumers | Unbounded queues merely move the outage and increase stale work. |
| Load shedding | Overload threatening core traffic | Random shedding harms critical traffic; clients retry immediately. |
| Rate limit/quota | Abuse, runaway cost, and noisy neighbors | A single global quota blocks unrelated tenants; rejected work has no retry guidance. |
| Fallback | Dependency or model route unavailable | Fallback is untested, lower quality, policy-incompatible, or doubles traffic. |
Graceful degradation should preserve truth. Examples: answer from a verified cache with a visible age; switch from an agentic write flow to read-only retrieval; queue a document update and show delayed status; return a cited search result instead of generating; or fail closed for a high-risk tool. A smaller model is not automatically safe: re-run policy and quality gates, disclose capability differences where relevant, and ensure it supports the required region and data terms.
Capacity, disaster recovery, and dependency isolation
Load testing finds throughput and latency under expected mix; stress testing finds the failure boundary; soak testing exposes leaks and slow degradation. Include model latency distributions, streaming connections, long documents, retries, tenant bursts, and cold caches. Capacity plans reserve headroom for failover: if one zone fails, remaining capacity must carry critical load without triggering autoscaling too late.
Set RPO/RTO per state. Conversation records may need point-in-time database recovery; derived vectors may be rebuilt; queued tool actions may require reconciliation before replay; provider credentials may need separate secured recovery. Exercise failover and restore, including the route back to primary. “Multi-region” without conflict, identity, secret, data-residency, and failback design is a diagram, not a recovery plan.
4. AI security: constrain authority, not just text
An LLM processes instructions and untrusted data in the same medium. Prompt rules are useful behavior guidance, but authorization must be enforced by deterministic systems outside the model.
Threat-model the complete flow
List assets (tenant documents, credentials, tool authority, model inputs/outputs, audit evidence), actors (user, tenant admin, insider, compromised document/source, provider, operator), trust boundaries, entry points, and abuse outcomes. Then trace data and authority through retrieval, memory, prompt construction, model, tool broker, downstream API, and logs.
| Threat | Preventive controls | Detective/recovery controls |
|---|---|---|
| Indirect prompt injection in a retrieved document | Treat retrieved text as data; isolate instructions; least-privilege tools; deterministic authorization; approval for consequential actions | Adversarial evals, tool-policy denials, canary documents, trace/audit review |
| Excessive agency/tool abuse | Narrow typed tools, user-context credentials, allowlisted parameters, budgets, sandbox, preview/approval | Rate/anomaly alerts, immutable action log, revocation, compensation workflow |
| Retrieval poisoning | Authenticated ingestion, provenance, version review, publisher trust, ACL at query and fetch | Quality/security scans, lineage lookup, generation rollback and targeted purge |
| Cross-tenant data exfiltration | Tenant-derived identity, database/index/object isolation, cache-key scoping, output mediation | Cross-tenant tests, access audit, canary tokens, incident deletion workflow |
| Sensitive output/log leakage | Minimize collection, redact/tokenize, output DLP/policy, retention, provider data controls | Access monitoring, deletion verification, sampled privacy review |
The OWASP 2025 LLM guidance identifies prompt injection and excessive agency as distinct but related risks. Its excessive-agency mitigations emphasize minimizing tool functionality, permissions, and autonomy and enforcing downstream authorization: OWASP Excessive Agency. Do not claim input filtering “solves” prompt injection; reduce the impact when the model is influenced.
Tool broker pattern
flowchart LR
MP["Model proposal"] --> SV["Schema validation"]
SV --> AU["Policy and user auth"]
AU --> RK["Risk and limit check"]
RK --> AP["Human approval when required"]
AP --> EX["Scoped execution"]
EX --> OV["Output validation"]
OV --> AR["Audit receipt"]
The model proposes; the broker decides. Bind each action to authenticated tenant/user, purpose, resource scope, idempotency key, monetary or row limit, and expiry. Resolve opaque resource IDs server-side instead of letting the model supply arbitrary URLs or SQL. High-impact actions receive a human-readable preview based on validated parameters, not free-form model prose. Approval must bind to the exact action digest so parameters cannot change afterward.
Enterprise privacy and evidence
Classify data before choosing model/provider and region. Document data-processing purpose, storage/retention, training/use terms, subprocessors, encryption, residency, deletion, incident handling, and access. Minimize prompts, redact when compatible with the task, and separate customer content from operational telemetry. Rotate secrets and encryption keys through an exercised procedure; never make raw secrets model context.
Audit records should answer who or what principal acted, for which tenant, under which policy/model/tool versions, on which resource, with what approval, outcome, and correlation ID. Avoid storing the sensitive payload when a hash/reference and separately controlled evidence store suffice. Retention and legal requirements vary by customer and jurisdiction; do not turn awareness of GDPR, SOC 2, or ISO 27001 into a claim of compliance. NIST’s Generative AI Profile organizes voluntary risk work across govern, map, measure, and manage: NIST AI 600-1.
5. Worked incident: provider latency becomes a retry storm
This is a hypothetical interview scenario. Example times and measurements illustrate how to tell an incident story; they are not Purnendu’s experience or results.
Detection and mitigation
At 10:02 UTC, fast-burn alerts fire for interactive-answer latency and valid-success SLOs. Queue age and model-call attempts rise, while application CPU remains moderate and provider first-attempt latency rises. A deployment marker shows no internal release. Traces reveal that gateway, orchestration service, and SDK each retry the same timeout, amplifying attempts.
- Declare incident, assign incident commander, operations lead, communications lead, and scribe; preserve a shared timeline.
- Disable lower-level retries through dynamic configuration, reduce per-request attempt budget, open a route-scoped circuit, and shed noncritical batch traffic.
- Route eligible low-risk requests to a previously evaluated fallback; fail closed for unsupported tools and expose a clear retryable status.
- Protect recovery by limiting client retry guidance, monitoring fallback capacity/quality, and keeping one controlled probe path to the primary.
By the illustrative 10:18, burn rate falls; by 10:40, the primary is stable but traffic is restored in steps. The team verifies SLO, queue drain, error mix, fallback quality, and cost before resolving. Customer communication states observed impact and current mitigation without speculating about root cause.
Root cause versus trigger
The provider slowdown is the trigger. The internal root cause of severity is uncoordinated retries without a propagated deadline or shared budget, plus a fallback path whose capacity alarm was missing. Contributing conditions include a timeout longer than the upstream request budget, an alert on CPU rather than attempt amplification, and a runbook that did not identify retry owners.
Corrective actions should have owners and verification: one retry layer; deadline propagation test; attempt-count metric; dependency-specific bulkhead; fallback load/quality drill; client retry contract; and a chaos scenario in release qualification. Avoid “be more careful.” A blameless postmortem holds the system and decisions accountable while creating conditions for truthful reporting.
Syllabus checkpoint: operations and enterprise identity
From structured logs to an on-call decision
Structured logs capture discrete, queryable events with stable fields; metrics summarize rates and distributions; traces connect causal work across retrieval, model, tool, queue, and database boundaries. OpenTelemetry provides shared context and export, Prometheus stores/scrapes metrics, and Grafana commonly visualizes and alerts across data sources. Tool choice is secondary to cardinality control, redaction, sampling, retention, and a trace ID that connects the user-visible failure to evidence.
An on-call system needs severity definitions, ownership, escalation, runbooks, safe mitigations, communication cadence, and post-incident follow-through. Alerts should describe a user symptom and an action, not every internal anomaly. Test alert delivery and runbooks during failure drills; an unexercised pager path is not a control.
OAuth/OIDC, IAM/RBAC, and key rotation
OAuth delegates access; OIDC adds an identity layer and ID-token semantics. IAM defines principals and permissions across the platform, while application RBAC maps verified identity to domain roles—often with attribute checks for tenant, resource, or risk. Keep authorization server-side and test denied paths. Key rotation needs overlapping validity, versioned key identifiers, atomic rollout, detection of stale consumers, revocation for compromise, and an audit trail; “replace the secret” is not an operational plan.
Interview playbook
Use the PROMISE framework:
- P — Promise: user journey, SLI, SLO, eligibility, window, and policy.
- R — Risks: dependency, overload, data, security, privacy, and operator failure modes.
- O — Observability: correlation, metrics/logs/traces, sampling, redaction, dashboards, alerts.
- M — Mitigation: deadlines, isolation, backpressure, load shedding, safe fallback, and fail-closed cases.
- I — Incident: roles, timeline, evidence, communication, and recovery verification.
- S — Security: assets, trust boundaries, least authority, deterministic mediation, audit, and retention.
- E — Exercise: load/chaos/security tests, restore drills, measured outcomes, and corrective ownership.
Common traps are defining availability at the pod, paging on every error, using raw user IDs as metric labels, storing full prompts by default, stacking retries, treating fallback as merely a cheaper model, claiming prompt injection is prevented by a system prompt, trusting the model to authorize its own tools, listing compliance acronyms as controls, or ending an incident at mitigation without root cause and verification.
Question bank
Practise answers that join reliability, observability, and security rather than treating them as separate teams.
Q1Define an SLO for a retrieval-augmented assistant.
Strong answer outline
- Name the journey and eligible events; define good as authorized, valid, sufficiently grounded, and within a latency target.
- Separate immediate serving SLI from delayed sampled quality SLI; slice by tenant tier, request class, and route.
- Choose target/window from user tolerance and capability, then attach an error-budget policy.
Follow-up probes
- Are policy denials errors?
- How do you measure answer quality online?
Your numerator, denominator, measurement point, slices, and action policy are unambiguous.
Q2Why use multi-window burn-rate alerts instead of a 5% error-rate alarm?
Strong answer outline
- Burn normalizes observed bad rate to the SLO’s allowed rate and connects alerts to budget risk.
- A short window catches fast incidents; a longer window confirms sustained impact and reduces noise.
- Page only actionable threats, with scope/runbook; use slower alerts for gradual consumption.
Follow-up probes
- What happens with low traffic?
- How do maintenance windows affect eligibility?
You tied alerting to user promise and response urgency, not arbitrary percentages.
Q3What telemetry would you capture for one agent tool call?
Strong answer outline
- Trace proposal, schema validation, authorization/policy, approval, execution, output validation, and compensation.
- Record bounded tool name/version, risk class, outcome/error class, duration, attempt, tenant-safe correlation, and token/cost counts.
- Keep arguments/results out by default; use controlled redacted evidence references with retention/access audit.
Follow-up probes
- How do retries appear?
- What belongs in an audit log versus a trace?
You can reconstruct authority and outcome without creating a sensitive shadow dataset.
Q4How do you control metric cardinality in a multi-tenant service?
Strong answer outline
- Allow only bounded enumerations such as route and normalized error class; prohibit prompt, document, user, trace, and raw tenant IDs.
- Use logs/traces for high-cardinality investigation and aggregate selected tenant cohorts or top-impact reports outside core metrics.
- Enforce label allowlists/tests and monitor active series/cost.
Follow-up probes
- How do you debug one tenant?
- Why are exception messages unsafe labels?
You preserve drill-down through correlated signals while bounding the metric dimension space.
Q5Head or tail trace sampling for rare model timeouts?
Strong answer outline
- Head sampling is simple and predictable but decides before knowing the outcome.
- Tail sampling can retain errors/high latency after completion but needs collector buffering, capacity, and a decision wait.
- Use tail rules for rare failures plus an unbiased baseline; monitor dropped telemetry and protect sensitive attributes.
Follow-up probes
- How do distributed services make one decision?
- Can sampled traces calculate the true error rate?
You explain operational cost, statistical bias, and the role of metrics.
Q6How do timeouts and retries become a retry storm?
Strong answer outline
- Independent layers exceed the user deadline and multiply attempts while the dependency is already saturated.
- Propagate one deadline, assign retry ownership, bound attempts/backoff/jitter, use budgets and idempotency.
- Add circuit/bulkhead/admission controls and measure attempts per logical request.
Follow-up probes
- What should happen to queued retries after the deadline?
- How do clients receive retry guidance?
You quantify amplification and coordinate—not merely tune—retry behavior.
Q7Design graceful degradation when the primary model provider fails.
Strong answer outline
- Classify requests by capability/risk; use an evaluated compatible fallback only for eligible classes.
- Reapply policy, authorization, residency, quality, latency, and cost gates; limit fallback capacity and prevent oscillation.
- Fail closed or queue unsupported tool actions, communicate state, and canary restoration.
Follow-up probes
- What if fallback output format differs?
- How do you prevent doubled spend?
Your fallback has a contract, capacity plan, security review, and recovery path.
Q8How would you protect a shared queue from a noisy tenant?
Strong answer outline
- Apply authenticated per-tenant admission quota, weighted fair scheduling, concurrency caps, and maximum queue age.
- Separate critical interactive and batch pools; bound downstream calls and storage.
- Expose tenant lag/rejections with safe retry guidance and tune quotas from contractual capacity.
Follow-up probes
- How can unused capacity be borrowed?
- What is the failure mode of per-tenant queues?
You enforce fairness at admission, scheduling, execution, and dependencies.
Q9Threat-model indirect prompt injection in enterprise RAG.
Strong answer outline
- Model a malicious/compromised document crossing ingestion and prompt boundaries to influence tool or output behavior.
- Preserve provenance and ACLs; treat retrieved text as untrusted data; mediate typed least-privilege tools with deterministic auth and approvals.
- Run adversarial evals, monitor policy/tool denials, audit actions, and support targeted purge/rollback.
Follow-up probes
- Why is input sanitization insufficient?
- How does retrieval poisoning differ?
You reduce blast radius even when model influence succeeds.
Q10How should a human approval gate be secured?
Strong answer outline
- Render a deterministic preview from validated typed parameters, actor, target, limits, and expected effect.
- Bind approval to an action digest, approver identity, tenant, policy version, expiry, and one-time nonce.
- Reauthorize and revalidate at execution; log outcome and support cancellation/compensation.
Follow-up probes
- What if the resource changes after approval?
- Can the model approve its own request?
Your approval cannot be reused or silently altered and does not replace execution-time authorization.
Q11How do you prove a tenant cannot leak through caches and vector search?
Strong answer outline
- Derive tenant/security filter from authenticated context and include it in database/index queries and cache namespace.
- Fetch-authorize returned source documents, avoid trusting model citations, and isolate administrative paths.
- Run property/adversarial tests with identical IDs, crafted filters, stale ACLs, cache collisions, and tenant deletion.
Follow-up probes
- Pre-filter or post-filter vector results?
- How do ACL changes invalidate caches?
You enforce before retrieval, at fetch, and in cache invalidation, with concrete negative tests.
Q12What makes an alert actionable?
Strong answer outline
- It maps to threatened user impact/SLO with affected scope and urgency.
- It has a responder, likely first checks, relevant dashboards/traces, and a safe runbook action.
- It is tested, deduplicated/inhibited appropriately, and reviewed after incidents for precision and recall.
Follow-up probes
- Page on CPU at 80%?
- How do you detect silent quality regression?
You distinguish pages, tickets, and dashboards by required human action.
Q13Tell an incident story when an external provider was the trigger.
Strong answer outline
- Quantify user impact and detection; explain roles, timeline, containment, and communications.
- Use telemetry to distinguish external trigger from internal severity multipliers such as retries or missing isolation.
- Verify recovery and name owned, testable prevention actions and what the team learned.
Follow-up probes
- What did you believe initially that was wrong?
- How did you avoid unsafe fallback?
You demonstrate judgment and system learning without blaming the dependency or inventing metrics.
Q14How do you decide what AI telemetry may be retained?
Strong answer outline
- Start from purpose, classification, tenant/customer requirements, provider terms, jurisdiction, and minimum necessary fields.
- Prefer derived counts/classes/hashes; segregate sensitive evidence with encryption, access audit, short retention, and deletion propagation.
- Document approval, sampling, diagnostic exceptions, and test that logs/traces do not capture secrets or cross tenants.
Follow-up probes
- What if debugging needs a prompt?
- How do legal holds affect deletion?
You balance operational need with explicit governance and do not make unsupported legal claims.
Proof artifact: production ownership drill
Instrument a small RAG-plus-tool service or the chapter 06 platform. Use synthetic data. Any numeric thresholds are example lab objectives.
- Define two user journeys with SLI equations, eligibility, example SLOs, error-budget policy, fast/slow burn alerts, and runbooks.
- Add OpenTelemetry traces across API, queue, retrieval, model, policy, approval, and tool; export metrics and structured redacted logs. Document attribute allowlists and retention.
- Implement deadlines, one retry owner, retry budget, bounded queues, per-tenant bulkhead/quota, circuit breaker, and one evaluated degraded mode.
- Create a data-flow threat model and abuse cases for indirect injection, excessive agency, cross-tenant retrieval, secret leakage, and malicious tool output. Add typed tool mediation and an approval digest.
- Write an incident template with roles, timeline, impact updates, mitigation decision log, root-cause tree, and corrective-action verification.
Measure: SLI and burn rate, attempts per logical request, p50/p95/p99 latency, queue oldest age, circuit state, shed/degraded requests, fallback quality/cost, trace/log drop rate, active metric series, policy denials, cross-tenant test failures, and recovery time. Keep labels bounded.
Inject failures: disable the model provider; add 2-second latency; return 429s; fill a queue with one tenant; kill the telemetry collector; corrupt a retrieved document with tool instructions; try arbitrary tool parameters; expire approval; leak a fake secret and verify redaction; restore state; and canary recovery. Capture what the user sees and whether the error budget stops release.
Present: SLO sheet, dashboard and alert screenshot, one end-to-end trace, redaction/cardinality tests, threat-model diagram, tool-policy test, incident timeline, root-cause tree, corrective action with owner/test, and a three-minute live dependency-failure drill.
Chapter review
Production ownership connects a measurable promise to telemetry, bounded failure controls, deterministic authority, practiced response, and verified recovery. Reliability and security both reduce uncontrolled blast radius.
Glossary
- Burn rate
- Observed bad-event rate divided by the rate allowed by an SLO.
- Bulkhead
- Resource isolation that prevents one workload, dependency, or tenant from exhausting all capacity.
- Deadline
- The total remaining time within which an operation remains useful to its caller.
- Excessive agency
- Risk created by giving an AI system more functionality, permission, or autonomy than its task requires.
- High cardinality
- A label/attribute dimension with many or unbounded distinct values.
- Prompt injection
- Untrusted content influencing model behavior contrary to the application’s intended instruction hierarchy.
- SLI / SLO
- A measured service indicator and its target over a defined population and time window.
- Tail sampling
- Selecting traces after enough of their outcome is known to apply error or latency policies.
Mastery checklist
- I can write an SLI equation with an honest denominator and cohort slices.
- I can connect error-budget burn to a release and incident policy.
- I can trace an AI/tool request without retaining unnecessary sensitive content.
- I can bound labels, sampling, retries, queues, concurrency, and fallback capacity.
- I can explain which actions degrade, queue, reject, or fail closed.
- I can enforce tool authorization outside the model and bind human approval to an exact action.
- I can distinguish an incident trigger, root cause, contributors, and verified corrections.
- I have exercised restore, dependency failure, cross-tenant access, and prompt-injection scenarios.
Primary sources
Links checked 2026-08-04.
- Google SRE — The Art of SLOs and Production Services Best Practices
- OpenTelemetry — Signals and Sampling
- Prometheus — Histograms and summaries
- OWASP Top 10 for Large Language Model Applications and LLM06:2025 Excessive Agency
- NIST AI 600-1 — Generative Artificial Intelligence Profile
- IETF RFC 9700 — OAuth 2.0 Security Best Current Practice
CHAPTER 12 · ARCHITECTURE
GenAI System Design & Forward Deployed Engineering
22 min read · 14 interview drillsLearning objectives
By the end of this chapter, you should be able to:
- lead technical discovery that converts an ambiguous customer request into outcomes, constraints, risks, and a definition of done;
- estimate load and data scale, define APIs and records, and trace critical flows before selecting products;
- design and compare enterprise RAG, multi-tenant assistant, vector search, agent, document AI, LLM gateway, connector, and evaluation platforms;
- make tenant isolation, failure recovery, observability, security, cost, migration, and rollback first-class architecture decisions;
- plan a Forward Deployed Engineering engagement from workshop through pilot, rollout, operational handoff, and measurable value; and
- communicate rejected alternatives and unsafe requirements clearly to engineers, security reviewers, and executives.
1. Discovery before diagrams
A Forward Deployed Engineer operates where customer process, messy data, security policy, product capability, and production engineering meet. The first deliverable is a shared problem definition. Drawing components too early hardens guesses into architecture.
Move from request to decision
A customer may ask, “Build an AI assistant for all our policies.” Clarify the job: Who asks which questions? What decision or task follows an answer? Which source is authoritative? Is a citation mandatory? How stale may content be? What must the system refuse? Is the assistant read-only, or may it take actions? What existing workflow, cost, or risk should change?
| Discovery lens | Questions that change architecture | Evidence to request |
|---|---|---|
| Outcome | Which user/business behavior improves? What is baseline and target? | Current workflow, sampled cases, handling time/quality measure, owner |
| Users and authority | Employees, agents, managers, external users? Read or write? Approval? | Role matrix, identity provider, sample entitlements, escalation process |
| Data | Sources, formats, volume, churn, language, ACLs, retention, residency? | Representative redacted corpus, inventory, permission model, deletion rules |
| Quality/risk | What is an unacceptable answer or action? Human review? Audit? | Gold cases, incidents, policy documents, risk classification |
| Operations | Latency, availability, recovery, deployment environment, support hours? | SLOs, network diagram, runbooks, change windows, procurement constraints |
| Adoption | Who changes process, trains users, approves rollout, and owns steady state? | Stakeholder map, rollout cohorts, communications/training plan, RACI |
Write assumptions as tests
Maintain an assumption log with owner, evidence, risk if wrong, and validation date. “Documents contain reliable ACL metadata” becomes: sample 500 representative objects across systems; compare source entitlements with retrieved results; require zero unauthorized returns before pilot. “The model is accurate” becomes an evaluation dataset segmented by task and risk, an acceptance threshold, and human adjudication for disagreement.
Define done at four levels:
- Functional: supported journeys and explicit non-goals.
- Quality and safety: groundedness/task success, policy compliance, access control, and human escalation.
- Operational: latency, availability/freshness, recovery, observability, runbook, support owner.
- Value/adoption: eligible users, sustained usage, process outcome, and a measurement design that avoids misleading attribution.
2. A repeatable system-design method
Strong system design is structured uncertainty reduction. Use the same sequence in a 45-minute interview and a customer architecture workshop, changing only depth.
Requirements, estimates, contracts, flows
- Frame: users, core use cases, non-goals, system boundary, source of truth, and success.
- Quantify: tenants/users, requests per second, concurrency, objects/bytes/vectors, churn, payload size, fan-out, growth, peak factor, latency, SLO, RPO/RTO, and budget range.
- Define contracts: external APIs/events, job state machines, core records, identity and authorization context, idempotency and versioning.
- Trace flows: one normal read/write and the highest-risk asynchronous flow, including checkpoints and audit.
- Choose architecture: components only after their responsibility is clear; identify data/control planes and trust boundaries.
- Stress: overload, dependency loss, duplicate/out-of-order events, schema/model migration, tenant leak, regional failure, human/operator error.
- Operate and evolve: SLIs, capacity/cost, deployment, canary, rollback, migration, ownership, and rejected alternatives.
Estimate with ranges, not theatre
Suppose an illustrative design has 50,000 users, 10% active in a peak hour, and six assistant turns per active user: about 30,000 requests/hour or 8.3 requests/second average during that hour. Apply an example 3× burst factor: roughly 25 requests/second. If each request retrieves 20 candidates and sends an average 8,000 total tokens through the model, provider throughput and cost—not API CPU—may dominate. Show the arithmetic, label every assumption, and state which load test or billing sample will replace it.
Capacity is multi-dimensional: requests/second, simultaneous streams, tokens/minute, queue age, database connections, index working set, OCR CPU, provider quotas, and human approval throughput. A system can have idle CPU while blocked on tokens/minute or a saturated connection pool.
Make interfaces carry correctness
POST /v1/assistant/turns
Authorization: Bearer ...
Idempotency-Key: ...
{
"conversation_id": "...",
"message": "...",
"requested_tools": ["policy_lookup"],
"client_context": {"locale": "en-IN"}
}
The server derives tenant, user, roles, policy and quotas from auth.
The client never supplies a trusted tenant_id or unrestricted tool URL.
Separate synchronous admission from long-running work. Version events and prompts/policies. Give every action and artifact a stable ID, lifecycle state, actor, tenant, provenance, and timestamps. Prefer deterministic downstream authorization over asking the model whether access is allowed.
3. Eight practice architectures and their hard parts
The syllabus’s systems share primitives, but each has a different correctness center. In an interview, spend time where failure is uniquely expensive.
| Design | Correctness center | Decisions worth defending |
|---|---|---|
| Enterprise RAG platform | Authorized, attributable, fresh retrieval and evaluated answers | Ingestion/index generations, hybrid retrieval/reranking, ACL enforcement, citations, eval segments, rollout |
| Multi-tenant AI assistant | No cross-tenant state or authority; fair capacity | Shared versus stamp isolation, memory lifecycle, tenant cache/index keys, quotas, audit, residency |
| Large-scale vector search | Recall/latency under filtering, updates, and migration | Exact versus approximate, shard/replica key, index parameters, hot tenants, dual-read/write migration, backfill |
| Agent platform | Bounded, resumable, authorized action | Typed tools, checkpoint state machine, budgets, human approval, sandbox, idempotency, compensation, trace |
| Document AI pipeline | Replayable, traceable extraction with visible uncertainty | Immutable raw input, OCR routing, quality gates, lineage, human review, generation publish, deletion |
| LLM gateway | Policy-consistent routing and auditable provider use | Auth/quotas, capability registry, deadlines/retries, semantic versus exact cache, fallback, residency, token/cost telemetry |
| Enterprise connector | Durable synchronization under duplicate/missed events | OAuth lifecycle, webhooks, inbox/outbox, checkpoints, backpressure, reconciliation, schema evolution |
| Evaluation platform | Comparable, reproducible evidence that catches segment regressions | Dataset/version lineage, offline/online metrics, judge calibration, experiment assignment, release gates, dashboard uncertainty |
Isolation is a spectrum
A fully shared deployment is cost-efficient and operationally simple but relies heavily on correct logical isolation and fairness. Per-tenant infrastructure improves blast-radius and configuration isolation but costs more and creates fleet-management/version-skew work. Deployment stamps group selected tenants and provide a scalable middle ground. Microsoft’s current multitenancy guidance frames the choice as trade-offs among isolation, cost, scale, performance, complexity, and manageability: architectural approaches for multitenancy.
Decide separately for compute, database, object storage, vector index, queues, encryption keys, network, and model route. A regulated tenant might have a dedicated data plane while sharing a global control plane. The control plane provisions connections, policies, tenants, and deployments; the data plane serves tenant traffic. Compromise or overload of one should not grant authority over the other.
Migrations are systems, too
For a vector-index migration, snapshot a source boundary, bulk backfill into a versioned target, capture concurrent changes via an ordered log/outbox, validate counts and sampled recall/ACL results, shadow or dual-read, canary tenants, then switch a routing pointer. Retain rollback until the old index’s update stream and retention window can safely close. Dual-write alone is not proof: one side can fail silently, so reconciliation is mandatory.
For API, schema, prompt, model, or tool migrations, state compatibility direction and rollback unit. Backward-compatible expand/migrate/contract often beats a flag day. Record the migration version with outputs so evaluation and audit can reproduce behavior.
4. Worked design: a multi-tenant policy assistant with approved actions
This hypothetical system demonstrates a complete interview answer. It is not a claim about Purnendu’s delivered projects or metrics.
Requirements and illustrative scale
Employees ask policy questions with citations and may propose a small set of HR service actions. Answers must honor source permissions and regional retention. High-impact writes require human confirmation. Assume for sizing 200 enterprise tenants, 100,000 total users, 40 peak assistant requests/second, 2 million source documents, a 15-minute freshness target for changed policies, 99.9% illustrative serving availability, and region-specific data planes. Confirm all values in discovery.
Non-goals for the first release: open-ended web browsing, arbitrary SQL or HTTP tools, autonomous approval, payroll decisions, and training on customer content. Success combines evaluated grounded-answer quality by policy domain, zero unauthorized retrieval in adversarial testing, latency/SLO, pilot adoption, and an agreed workflow outcome measured against a baseline.
Architecture
flowchart TD
CP["Global control plane"] --> RA["Region A data plane"]
CP --> RB["Region B data plane"]
U["User and IdP"] --> GW["Gateway"]
GW --> OR["Assistant orchestrator"]
OR --> PE["Policy engine"]
OR --> HR["Hybrid retrieval and ACL"]
OR --> LG["LLM gateway"]
OR --> TB["Tool broker and approval"]
TB --> HS["Business systems"]
OR --> AU["Audit and telemetry"]
SRC["Enterprise sources"] --> CN["Connectors"]
CN --> VP["Versioned pipeline"]
VP --> IG["Index generation"]
VP --> EV["Evaluation gate"]
The global control plane maps a tenant to a deployment stamp and manages versioned configuration, but carries no customer prompt content. The regional gateway validates identity, derives tenant/user/roles, applies quotas, and routes only to the mapped data plane. The assistant orchestrator stores a checkpointed turn state with prompt/policy/model/retrieval versions and deadline.
Read and action flows
- Normalize the user request and run policy/risk classification. Derive authorization filters from identity.
- Run hybrid lexical/vector retrieval, pre-filter by tenant and ACL where supported, rerank, then reauthorize source fetch. Include provenance and current source version.
- The LLM gateway selects an allowed model route by capability, region, policy, health, and budget. It enforces deadline/token limits and records usage without raw content by default.
- Validate the answer structure, citations, policy, and uncertainty. If evidence is insufficient, return a scoped refusal or escalation rather than inventing.
- For an action, the model emits a typed proposal only. The tool broker reauthorizes the user, validates allowlisted resource IDs and limits, creates an idempotent pending action, and renders a deterministic approval preview.
- Approval is bound to action digest, approver, tenant, and expiry. Execution uses user-context or narrowly scoped service credentials, then records a receipt. Uncertain remote results enter reconciliation before retry.
Failure, security, and operations
| Scenario | Designed behavior | Signal |
|---|---|---|
| Vector store unavailable | Serve only verified fresh exact-cache entries or fail with an honest retryable status; no uncited generation | Valid-answer SLI, retrieval errors, cache age, circuit state |
| Primary model slow | Deadline-aware evaluated fallback for read-only eligible tasks; actions fail closed if capability/policy differs | Attempt amplification, route latency, fallback quality/cost |
| Malicious policy document | Treat content as untrusted; tools remain mediated; provenance supports quarantine and index-generation rollback | Injection eval, tool denials, document canary, lineage |
| Noisy tenant | Per-tenant admission and concurrency; weighted fair queues; dedicated stamp option | Tenant cohort latency, rejected/queued work, saturation |
| ACL changes during conversation | Reauthorize each retrieval/tool action; invalidate affected cache; do not trust old memory as authority | Policy/ACL version, authorization denials, stale-cache audit |
| Regional outage | Route only if approved replicated state and residency allow; otherwise communicate outage and recover to RTO | Regional SLO, replication lag, failover/failback drill |
Cost controls include per-tenant/model token budgets, prompt and retrieval limits, cache only where identity/policy/version keys make reuse safe, small-model routing for evaluated classes, incremental ingestion by content hash, and storage lifecycle. Report cost per successful eligible task alongside quality; optimizing cost per raw request rewards cheap failures.
Rejected alternatives
- One model call with all tenant documents: rejected for context limits, cost, stale data, poor provenance, and access-control risk.
- One dedicated stack per tenant from day one: rejected as the universal default because fleet cost and upgrades grow quickly; retained for isolation/residency tiers.
- Let the model call the HRIS directly: rejected because free-form authority, secrets, approval, idempotency, and audit cannot be enforced reliably.
- Active-active global writes immediately: rejected unless recovery requirements justify conflict, replication, residency, and operational complexity.
5. Forward Deployed execution: from pilot to durable ownership
FDE work succeeds when the customer can operate, trust, and extend the system after the initial team leaves. Technical depth and change management are one delivery problem.
Engagement sequence
- Align: stakeholder map, sponsor, user owner, security/data/platform owners, decision process, RACI, outcomes, constraints, and working cadence.
- Discover: workflow observation, representative data/ACL sample, architecture/security review, baseline measurement, and assumption/risk register.
- De-risk: thin vertical slice against the hardest unknown—often access-correct retrieval, malformed documents, tool authorization, or deployment connectivity.
- Pilot: limited users/data, shadow or read-only mode, explicit acceptance criteria, support channel, daily evidence review, and kill switch.
- Productionize: SLOs, capacity, threat model, incident/restore drills, runbooks, observability, cost guardrails, and ownership training.
- Roll out: cohorts with canary metrics, change communication, user education, feedback triage, rollback gates, and decision log.
- Handoff and expand: architecture record, operations pack, backlog, known limitations, support escalation, outcome review, and next hypothesis.
Migration and workshop artifacts
A useful architecture workshop produces a context/data-flow diagram, requirements and NFR table, identity/ACL map, scale worksheet, risk/assumption register, option matrix, decisions and rejected alternatives, rollout plan, and owners. A migration assessment adds inventory, data quality, dependencies, compatibility, cutover/rollback, parallel-run duration, validation/reconciliation, downtime, training, and decommission criteria.
Do not make the proof of concept a hidden production system. Mark synthetic versus customer data, temporary credentials, retention, unsupported scale, missing controls, and expiry. Promotion requires an explicit review against production criteria rather than enthusiasm.
Handle objections and unsafe requirements
Listen for the underlying need, restate it, and separate non-negotiable safety from negotiable implementation. “No human approval because it slows the workflow” may conceal a latency goal. Offer risk-tiered automation: auto-execute bounded reversible low-risk actions; batch approvals; improve preview UX; keep high-impact irreversible actions gated. Quantify residual risk and seek the authorized risk owner’s decision. Never quietly implement an unsafe exception.
When a request conflicts with product capability or evidence, say: what is known, what is unknown, the failure/blast radius, the recommended safe path, alternatives, and the decision required. Escalation is a delivery skill when it preserves trust and schedule.
Measure value without overselling causality
Choose one primary outcome close to the workflow—such as eligible cases resolved with verified citations or time from accepted request to confirmed completion—and guardrails for quality, safety, cost, and equity across cohorts. Establish baseline and instrument eligibility before rollout. Use phased cohorts or a credible comparison when possible; report adoption and outcome separately. Example projections must remain labeled assumptions until measured.
Executive updates fit one page: outcome and current evidence; user/rollout scope; material risk/decision; spend/capacity; next milestone and owner. Engineering appendices hold trace, schema, and benchmark detail. Good communication changes resolution, not truth.
Syllabus checkpoint: the non-negotiable design frame
Before drawing components, write the design contract. Functional requirements describe user-visible behavior and workflows. Non-functional requirements attach measurable constraints to quality, latency, availability, freshness, security, privacy, compliance, recovery, and cost. Make scale estimates for users, tenants, requests, documents/events, vector count, write rate, storage growth, model tokens, and concurrency; show units and peak-to-average assumptions.
Then define API contracts, the data model and ownership of each record, and the end-to-end data flow for ingestion and serving. Walk normal operation plus explicit failure modes: malformed or stale input, duplicate event, dependency timeout, quota exhaustion, partial regional failure, cross-tenant access attempt, model regression, and operator mistake. For each, state detection, containment, recovery, customer behavior, and data reconciliation.
Interview playbook
Use DISCOVER on the whiteboard:
- D — Desired outcome: users, workflow, source of truth, definition of done, non-goals.
- I — Inputs and identity: data, ACLs, authority, classification, residency, lifecycle.
- S — Scale and SLOs: estimates, peaks, quality, latency, freshness, RPO/RTO, budget.
- C — Contracts and core state: APIs/events, records, state machines, versioning, idempotency.
- O — Operational architecture: data/control planes, flows, failure isolation, telemetry, capacity.
- V — Verification and value: evaluations, security tests, reconciliation, outcome baseline.
- E — Evolution: migration, canary, rollback, cost, rejected alternatives, decision triggers.
- R — Rollout and responsibility: cohorts, RACI, training, incident/restore, handoff.
Common traps are drawing before asking questions, stating scale without arithmetic, omitting source permissions, treating a vector database as the whole RAG system, trusting the model with tenant identity or authorization, saying “multi-region” without state semantics, ignoring migration/rollback, promising ROI without a baseline, accepting unsafe customer requirements, or presenting only the chosen design without rejected alternatives and reconsideration triggers.
Question bank
These probes test architecture judgment and customer-facing execution together.
Q1What are your first ten minutes after a customer asks for “an enterprise RAG platform”?
Strong answer outline
- Clarify users, decisions/workflow, source of truth, citations/refusal, data/ACLs, freshness, action scope, deployment, and success baseline.
- State non-goals and highest-risk assumptions; request representative corpus and permission evidence.
- Define a thin test that retires the hardest risk before choosing the full stack.
Follow-up probes
- What if the customer has no gold dataset?
- Who must attend discovery?
You leave with evidence, owners, and definition of done—not a vendor shopping list.
Q2Design tenant isolation for an AI assistant serving regulated and standard customers.
Strong answer outline
- Classify isolation/residency/threat requirements per component and derive tenant mapping from authenticated control-plane state.
- Use shared regional stamps for standard tenants with logical isolation, quotas, RLS/index/cache controls; dedicated stamps/keys/routes where required.
- Automate provisioning, policy, observability, upgrades, deletion, and cross-tenant tests across the fleet.
Follow-up probes
- What remains shared?
- How do you move a tenant between stamps?
You treat isolation as component-specific and include operational fleet cost and migration.
Q3How would you estimate capacity for a streaming assistant?
Strong answer outline
- Estimate active users × turns, burst factor, simultaneous stream duration, token input/output, retrieval fan-out, tool rate, and regional split.
- Map to API concurrency, provider token/requests quotas, database/pool, cache/index, network, and observability capacity.
- Give ranges/sensitivity, reserve failure headroom, and propose representative load/soak tests.
Follow-up probes
- Why is requests/second insufficient?
- Which signal drives autoscaling?
Your arithmetic exposes the actual bottleneck and labels assumptions.
Q4Design a large vector-index migration with no authorization regression.
Strong answer outline
- Version target schema/index and snapshot a boundary; bulk backfill while capturing changes via outbox/log.
- Validate counts, lineage, sampled recall/latency, and adversarial ACL behavior; reconcile dual paths.
- Shadow/dual-read, canary tenants, switch routing pointer, monitor, retain rollback, then decommission safely.
Follow-up probes
- How do deletes propagate?
- What if embeddings change dimension?
You cover concurrent writes, correctness, ACLs, canary, rollback, and retirement.
Q5What belongs in an agent platform checkpoint?
Strong answer outline
- Logical run/tenant/user, state version, completed/pending steps, typed inputs/outputs references, budgets/deadline, policy/model/tool versions.
- Idempotency keys, external receipts, approval digest/status/expiry, retry count, and compensation/reconciliation state.
- Encrypt/minimize sensitive content, authorize resume, and migrate checkpoint schemas deliberately.
Follow-up probes
- How does a code deployment resume old runs?
- What if tool outcome is unknown?
Your checkpoint supports safe, auditable resume rather than merely saving chat messages.
Q6Design an LLM gateway without creating a single point of catastrophic policy failure.
Strong answer outline
- Centralize authenticated routing, capability/region registry, quotas, policy versions, budgets, deadlines, telemetry, and provider adapters.
- Keep deterministic local authorization in applications/tool brokers; make gateway horizontally available with cached last-known-safe config and fail-closed rules.
- Canary config/model routes, audit changes, isolate tenants/providers, and exercise fallback.
Follow-up probes
- Where can semantic caching leak data?
- How do you handle gateway control-plane outage?
You gain consistent policy without granting the gateway unbounded content or authority.
Q7How do you design a document AI human-review queue?
Strong answer outline
- Route by explicit quality/risk signals with original rendering, extracted region, confidence, warnings, lineage, and task instructions.
- Prioritize by business impact/deadline, enforce tenant/PII access, lease work, capture structured correction and reviewer identity.
- Measure agreement, turnaround, backlog age, correction outcome, and feed validated labels into evaluation—not automatically into production training.
Follow-up probes
- How do reviewers avoid seeing unnecessary PII?
- What if reviewers disagree?
You designed authority, evidence, queue operations, quality control, and feedback governance.
Q8How would you architect an evaluation platform as a release gate?
Strong answer outline
- Version datasets/items, provenance, segment/risk labels, prompts, retriever/index, model, tools, code, and environment.
- Run deterministic and calibrated judge/human metrics, compare paired results by segment with uncertainty and failure examples.
- Encode critical-regression and aggregate thresholds, require waiver owner/evidence, store reports, and correlate with online outcomes.
Follow-up probes
- How do you prevent benchmark leakage?
- When can an aggregate improve but release fail?
Your gate is reproducible, segment-aware, auditable, and not controlled by one opaque score.
Q9How do you make a connector part of a larger system design rather than a side box?
Strong answer outline
- Specify OAuth/service identity, source semantics, webhooks plus inventory, canonical records, checkpoints, inbox/outbox, and reconciliation.
- Connect schema/ACL/deletion changes to lineage, indexing generations, caches, evaluation, and audit.
- Include provider quotas/outage, noisy tenant fairness, reauthorization, rollout, and operational ownership.
Follow-up probes
- What is the source of truth?
- How is a missed delete discovered?
You trace connector uncertainty into downstream correctness and recovery.
Q10A customer demands fully autonomous payroll changes. How do you respond?
Strong answer outline
- Clarify desired speed/volume and classify impact, reversibility, authorization, regulatory/customer policy, and failure blast radius.
- State why model-only autonomous authority is unsafe; propose typed bounded actions, deterministic validation/auth, preview/approval, limits, idempotency, audit, and staged evidence.
- Offer automation for low-risk reversible classes, quantify residual risk, and escalate the explicit decision to the authorized owner.
Follow-up probes
- What evidence could relax the gate?
- What if the sponsor refuses?
You preserve the underlying outcome while holding a clear safety boundary and escalation path.
Q11Plan rollout for replacing an existing enterprise search tool.
Strong answer outline
- Inventory integrations, content/ACLs, user workflows, baseline quality/latency/cost, and decommission constraints.
- Backfill and reconcile, shadow queries, evaluate by segment, pilot representative cohorts, train/support, and canary default routing.
- Keep old path/read-only fallback through acceptance; define cutover, rollback, data retention/export, and owner sign-off.
Follow-up probes
- How do you prevent selection bias in pilot users?
- When can the old index be deleted?
You cover technical migration, users, value evidence, rollback, and decommission.
Q12How do you show customer value without fabricating ROI?
Strong answer outline
- Define eligible workflow and baseline before launch; choose one primary outcome plus quality/safety/cost guardrails.
- Instrument adoption separately from outcome; use phased comparison or credible counterfactual and report uncertainty/confounders.
- Label projections as assumptions, report observed sample/window/cohorts, and agree who owns the business calculation.
Follow-up probes
- What if usage is high but outcome is flat?
- How do you value avoided risk?
You distinguish measured evidence, inference, and projection and avoid claiming personal metrics.
Q13What must be in an FDE production handoff?
Strong answer outline
- Architecture/decision records, source/config/code ownership, inventory, data/identity map, known limits and risk register.
- SLO/dashboard/alerts, runbooks, incident/escalation, backup/restore, key rotation, access review, deployment/rollback, cost/capacity.
- Named RACI, trained operators, paired drills, acceptance evidence, support terms, backlog, and decommission of temporary POC access.
Follow-up probes
- How do you test handoff quality?
- What temporary artifacts are dangerous?
The customer team can operate and recover independently, and temporary risk is removed.
Q14Present a complex architecture decision to an executive in two minutes.
Strong answer outline
- Lead with customer outcome and the decision required, not component names.
- Give two or three options with material trade-off in risk, time, cost, and reversibility; state recommendation and evidence.
- Name residual risk, next validation/milestone, owner, and what would change the recommendation.
Follow-up probes
- What technical detail stays in appendix?
- How do you communicate uncertainty?
A non-specialist can make the right decision without the architecture being misrepresented.
Proof artifact: four-design FDE portfolio
Create a reusable portfolio using synthetic scenarios and clearly labeled illustrative estimates. Do not imply customer delivery or personal metrics that are not documented.
- Choose four timed designs covering different correctness centers: enterprise RAG, multi-tenant assistant/agent platform, document AI or connector, and LLM gateway/evaluation platform.
- For each, produce a one-page brief: discovery questions, outcomes/non-goals, assumptions with arithmetic, NFRs, APIs/events, core records/state machines, architecture/data flow, trust boundaries, failure table, SLOs, capacity/cost, rollout/rollback, and two rejected alternatives.
- Build one thin vertical slice for the highest-risk assumption, such as ACL-correct retrieval, resumable approved tool action, generation-based reindex, or policy-consistent model fallback.
- Create an FDE engagement pack: workshop agenda, assumption/risk register, decision log, pilot definition of done, RACI, migration/cutover checklist, incident/restore drill, executive update, and handoff acceptance.
- Record a 35-minute whiteboard answer and a two-minute executive version. Review question-first timing, arithmetic, trade-offs, failure recovery, and whether the design maps back to value.
Measure: time to requirements and first architecture, number of explicit assumptions, estimate consistency, critical flows covered, failure/recovery completeness, threat boundaries, rollback units, decision rationale, and communication fit. For the prototype, add task/quality, authorization-negative tests, latency, cost per successful eligible task, and recovery drill time.
Inject failures: tenant mapping error, missed connector event, malformed document, index migration drift, model/provider outage, duplicate tool execution, expired approval, regional loss, and an unsafe stakeholder request. Show designed containment, evidence, decision owner, rollback, and customer communication.
Present: four architecture sheets, live scale calculation, one trace/state-machine demo, option matrix, migration and rollback sequence, risk register before/after the thin slice, pilot scorecard with hypothetical labels, two-minute executive recording, and operator handoff drill.
Chapter review
Architecture depth is the ability to turn uncertain customer needs into testable contracts, quantified trade-offs, controlled authority, recoverable operations, phased change, and durable ownership. FDE depth adds the human system that makes the technical system valuable.
Glossary
- Assumption register
- A living list of uncertain beliefs, owners, evidence, risk if wrong, and validation status.
- Control plane / data plane
- The management/provisioning path and the path that processes tenant/user workload data.
- Definition of done
- Agreed functional, quality/safety, operational, and value/adoption acceptance criteria.
- Deployment stamp
- A repeatable unit of infrastructure serving one or more tenants to balance isolation and fleet scale.
- Forward Deployed Engineering
- Customer-facing engineering that discovers, builds, integrates, deploys, and hands off production solutions in context.
- Non-goal
- An explicit boundary for what a design or release does not attempt to support.
- Thin vertical slice
- The smallest end-to-end implementation that tests a high-risk assumption across real boundaries.
- Rollback unit
- The independently reversible version boundary for code, schema, configuration, data, index, or infrastructure.
Mastery checklist
- I begin with outcome, users, authority, data, failure tolerance, and success evidence.
- I calculate illustrative scale transparently and state the benchmark that will replace assumptions.
- I define APIs, events, records, and state machines before naming every component.
- I can compare isolation, index, agent, gateway, connector, pipeline, and evaluation trade-offs.
- I trace security, privacy, observability, cost, migration, failure recovery, and rollback through the design.
- I can state rejected alternatives and objective triggers for reconsideration.
- I can convert an unsafe request into a bounded option and escalate the residual-risk decision.
- I can plan pilot cohorts, value measurement, customer communication, and an operator-tested handoff.
Primary sources
Links checked 2026-08-04.
- Microsoft Azure Architecture Center — Architect multitenant solutions, architectural approaches, and control planes
- AWS Well-Architected
- Google Cloud Well-Architected Framework
- Google SRE — The Art of SLOs
- OWASP Top 10 for Large Language Model Applications
- IETF RFC 9700 — OAuth 2.0 Security Best Current Practice
- AWS Prescriptive Guidance — Transactional outbox pattern
CHAPTER 13 · SENIOR SIGNAL
Leadership, Coding & Communication
18 min read · 14 interview drillsLearning objectives
This chapter turns seniority into observable behavior: sound decisions, useful written artifacts, calm coding, and evidence that other people became more effective because of your work.
- Build six truthful leadership stories that separate personal ownership from team outcomes.
- Lead an ambiguous architecture decision without relying on title or authority.
- Write an async design memo, decision record, status update, and review comment that lets others act.
- Solve coding problems by clarifying contracts, selecting a pattern, proving correctness, and testing edges.
- Handle practical Python, SQL, API, and debugging screens with production judgment.
- Review AI-assisted code as accountable engineering work, not as trusted output.
Make staff-level scope visible
A senior answer is not made senior by saying “I led.” It becomes senior when the interviewer can trace how you framed an unclear problem, changed a consequential decision, created leverage beyond your own code, and verified the result.
flowchart LR
A["Ambiguity"] --> F["Framing"]
F --> D["Decision"]
D --> L["Alignment"]
L --> Y["Delivery"]
Y --> V["Leverage"]
V --> E["Evidence"]
The six-story portfolio
Prepare one story for each row. A story may cover more than one row, but do not force every question into the same heroic project.
| Story | Decision worth explaining | Evidence to bring |
|---|---|---|
| Architecture without authority | How you converted competing constraints into an accepted direction | Decision memo, rejected alternatives, rollout gate |
| Standards and leverage | Why a reusable pattern was better than another one-off | Reference implementation, adoption trail, maintenance effect |
| Mentoring | How you diagnosed a capability gap and transferred ownership | Review progression, learning plan, later independent decision |
| Disagreement | What evidence changed the discussion and what you conceded | Experiment, ADR, decision log, follow-up result |
| Technical debt or incident | Why remediation displaced other work | Risk model, incident timeline, prevention and detection controls |
| Customer or product outcome | How field evidence altered requirements or sequencing | Discovery notes, success measure, rollout feedback |
Use STAR-L, but keep the “A” inspectable
Situation gives only the context needed to understand the stakes. Task names your mandate and constraints. Action should consume roughly half the answer: questions asked, analysis performed, options rejected, people aligned, safeguards added, and course corrections. Result separates measured outcomes from impressions. Learning says what you would repeat or change.
Lead decisions, not meetings
Influence without authority comes from improving the decision environment. Make the problem legible, expose trade-offs, invite the right objections, and leave a durable record.
Frame
Name the decision, owner, deadline, non-goals, constraints, and reversible versus irreversible parts.
Compare
Evaluate two or three viable options against explicit criteria such as quality, latency, operability, privacy, and migration cost.
De-risk
Run the smallest experiment that resolves the largest uncertainty. Do not prototype what documentation already answers.
Commit
Record the chosen option, dissent, triggers for revisiting it, rollout, rollback, and who owns follow-through.
A compact architecture decision record
Title: Choose the retrieval serving path
Status / owner / decision date:
Context: users, scale, sensitivity, current failure
Decision drivers: quality, p95 latency, cost, operations
Options: A / B / C, with evidence and migration effort
Decision: chosen option and why now
Consequences: gains, accepted debt, new risks
Rollout: shadow → limited cohort → broader release
Rollback and revisit triggers:
The record is not a ceremony. It is a compression mechanism: a teammate in another time zone should be able to challenge or execute the decision without reconstructing a meeting.
Resolve disagreement with a ladder
- Restate the shared objective and the other position until its owner agrees.
- Classify the disagreement: facts, forecasts, values, constraints, or ownership.
- Seek disconfirming evidence and define a time-boxed test when the decision is reversible.
- Ask the accountable owner to decide when evidence cannot remove uncertainty.
- Disagree, commit, and log a revisit trigger rather than relitigating continuously.
Write so work can continue without you
Current Remote and Sourcegraph role pages explicitly emphasize structured writing, async work, customer communication, autonomy, and substantive review. Treat writing as part of system reliability: missing context creates coordination failures just as missing timeouts create runtime failures. See the live Remote Senior Forward Deployed Engineer role and Sourcegraph Agent Engineer role.
The five-block async update
Outcome: what changed for the user or project
Evidence: test, metric, trace, screenshot, or decision
Risk: what could still invalidate the result
Next: owner and date for the next concrete action
Ask: one explicit decision or help request, if needed
Lead with the outcome, not an activity diary. Replace “worked on evaluation” with “the release gate now catches the three known citation regressions; the multilingual slice is still below its proposed threshold.”
Design memo anatomy
A two-page memo should cover problem and non-goals, users and success measures, constraints and assumptions, proposed flow, alternatives, failure modes, security, operations, rollout, and unresolved questions. Put the recommendation near the top. Attach detailed benchmarks rather than burying the decision beneath them.
Code-review comments that create leverage
Tag the nature of the comment: blocker for correctness or safety, important for maintainability, suggestion for a worthwhile alternative, and nit for optional polish. State the consequence and, where useful, a concrete path forward. Google’s engineering review guide and GitHub’s pull-request review documentation are useful primary references for review practice and review states.
Translate trade-offs for customers
Use consequence language. “A reranker adds another model call” is technical description. “A reranker may improve the difficult queries we sampled, but adds latency and cost to every request; we propose enabling it only for low-confidence queries and measuring both task success and p95 latency” is a decision a customer can evaluate.
Make coding reasoning observable
Senior coding screens still require fundamentals, but the strongest signal is controlled problem solving: establish the contract, choose the simplest correct structure, prove the invariant, and test the boundaries.
The 40-minute loop
- Clarify: input size, ordering, duplicates, empty values, mutation, return contract, and error behavior.
- Example: walk one normal case and one adversarial case by hand.
- Baseline: give a correct simple approach and its time/space cost.
- Pattern: choose map/set, two pointers, sliding window, stack, heap, binary search, traversal, backtracking, greedy, or dynamic programming because a specific invariant fits.
- Implement: name state by meaning and narrate only consequential choices.
- Verify: trace edges, state complexity, and identify what would change at production scale.
| Signal | Likely pattern | Invariant to explain |
|---|---|---|
| Longest/shortest contiguous range | Sliding window | When moving the left edge restores validity |
| Top k or repeatedly smallest | Heap | Heap contains only the best candidates seen |
| Reachability or dependencies | BFS/DFS/graph | Visited state prevents repeated work or cycles |
| Monotone answer predicate | Binary search on answer | All values on one side share feasibility |
| Overlapping choices | Dynamic programming | State captures all information future choices need |
Practise the work-shaped screens
The syllabus calls out async Python, SQL, APIs, transformations, tests, unfamiliar code, and reading Go or TypeScript. These exercises reward production instincts more than puzzle tricks.
Async Python: concurrency needs ownership
import asyncio
async def fetch_all(ids, fetch_one):
async with asyncio.TaskGroup() as group:
tasks = {item_id: group.create_task(fetch_one(item_id))
for item_id in ids}
return {item_id: task.result() for item_id, task in tasks.items()}
TaskGroup gives the subtasks a lifetime owned by the context. Discuss timeouts, bounded concurrency, partial-result policy, and cancellation cleanup; the Python documentation explains that task-group failure cancels remaining tasks and that cleanup should propagate cancellation after finally work. Read the official coroutines and tasks documentation.
SQL: express the business question first
WITH ranked AS (
SELECT tenant_id, run_id, cost_usd,
row_number() OVER (
PARTITION BY tenant_id ORDER BY cost_usd DESC
) AS cost_rank
FROM agent_runs
WHERE started_at >= :window_start
)
SELECT tenant_id, run_id, cost_usd
FROM ranked
WHERE cost_rank <= 3;
This asks for the three most expensive runs per tenant, not the three most expensive globally. Explain tie semantics (row_number versus rank), indexes supporting the filter, and why an execution plan matters. Use the official PostgreSQL guides to window functions and EXPLAIN.
API implementation checklist
- Validate request shape and semantic constraints; return stable error contracts.
- Derive identity and tenant scope from authentication, not caller-supplied fields.
- Define idempotency, timeout, retry, concurrency, and transaction boundaries.
- Test happy path, malformed input, duplicate request, dependency timeout, and forbidden access.
- Expose correlation identifiers and useful telemetry without logging secrets.
Debugging unfamiliar code
Restate the symptom, bound the blast radius, reproduce with the smallest input, trace data across boundaries, and form ranked hypotheses. Change one variable at a time. A senior candidate distinguishes mitigation from root-cause correction and adds a regression test plus a detection improvement.
AI-assisted code is still your code
Record the prompt or intent when relevant, inspect every changed line, verify dependencies and licenses, threat-model new data paths, run targeted and broader tests, and be ready to explain the result without the assistant. Current Sourcegraph and Automattic application material asks candidates for concrete opinions about coding agents; speak from a real workflow, while never presenting generated code as independently trustworthy.
Syllabus checkpoint: complete coding-pattern coverage
Keep a compact mental map rather than memorizing isolated solutions. Arrays and strings reward index discipline, maps/sets capture membership and counts, and two pointers or sliding windows exploit ordered or contiguous structure. Stacks model nested or monotonic state; queues model FIFO work; linked lists test pointer ownership and edge cases. Binary search needs a monotone predicate, while interval problems require explicit boundary semantics.
Trees and graphs share traversal machinery but different invariants: tree DFS supports subtree/postorder reasoning, BFS supports levels and minimum unweighted hops, and general graphs require cycle/visited handling. Heaps maintain an extremum or top-k frontier. Backtracking explores a choice tree with pruning; greedy algorithms require an exchange or staying-ahead argument; dynamic programming requires a state that contains all information future choices need. For every pattern, explain correctness, time/space complexity, empty/singleton/duplicate cases, and how you would test it.
Interview playbook
Choose the framework that matches the signal being tested, then keep the answer evidence-led.
Leadership answer: Scope → Judgment → Leverage → Evidence → Learning
- Scope: users, stakes, constraints, your mandate, and what was ambiguous.
- Judgment: alternatives and the pivotal decision, including what you rejected.
- Leverage: how you aligned people or created a pattern others could use.
- Evidence: measured result, rollout observation, or honest qualitative outcome.
- Learning: a specific correction you would make next time.
Coding answer: Contract → Baseline → Invariant → Code → Tests → Cost
Clarify the contract, state a correct baseline, name the invariant behind the better approach, implement, trace tests, then give time and space complexity. Mention production concerns only after solving the asked problem.
Common traps
- Using “we” throughout so the interviewer cannot locate personal ownership.
- Claiming consensus when a real decision owner or disagreement existed.
- Giving activity metrics instead of an outcome or risk reduction.
- Narrating every keystroke during coding instead of decisions and invariants.
- Ignoring cancellation, authorization, duplicates, or failure behavior in practical tasks.
- Claiming AI assistance made work faster without showing verification or a quality boundary.
Question bank
Answer aloud. Keep leadership answers near two minutes initially; allow follow-up probes to reveal the deeper evidence.
Q1Tell me about an architecture you led without formal authority.
Strong answer outline
- Define the ambiguous decision, affected teams, constraints, and your actual mandate.
- Show the options, evidence, dissent, and the mechanism used to reach a decision.
- Close with adoption, measured or observed outcome, and one learning.
Follow-up probes
- Who initially disagreed, and why?
- Which outcome belongs to you versus the team?
A strong answer makes influence mechanisms and personal actions inspectable; “I convinced everyone” without evidence does not pass.
Q2Two senior engineers disagree about an agent framework. How do you move the decision forward?
Strong answer outline
- Align on workload, reliability boundary, team capability, and decision deadline.
- Turn preferences into criteria; test the highest-risk unknown with a thin vertical slice.
- Let the accountable owner decide, record dissent and revisit triggers, then commit.
Follow-up probes
- What if the benchmark is inconclusive?
- When would you override consensus?
Include ownership and reversibility. Endless consensus-seeking or framework feature comparison alone is insufficient.
Q3How would you make a build-versus-buy decision for an LLM gateway?
Strong answer outline
- Define required routing, policy, observability, data handling, provider support, and exit constraints.
- Compare total ownership cost, integration risk, control, roadmap fit, and vendor lock-in.
- Propose a reversible pilot with acceptance thresholds and an exit plan.
Follow-up probes
- What evidence would reverse your choice?
- How do security review and incident ownership change the answer?
The answer must address lifecycle and migration, not just license price or feature count.
Q4A capable teammate repeatedly ships prompt changes without evaluation. How do you mentor them?
Strong answer outline
- Use a concrete escaped regression to establish the capability gap without personal blame.
- Pair on a minimal golden set, threshold, and review checklist; explain why each layer exists.
- Transfer ownership and later inspect whether they can design the next gate independently.
Follow-up probes
- What if delivery pressure rewards the shortcut?
- How do you know mentoring worked?
Pass only if the story changes the system and builds independent judgment, rather than merely correcting one pull request.
Q5Write the verbal version of an async update after a failed canary deployment.
Strong answer outline
- Lead with outcome: canary rolled back and stable version remains serving.
- Give evidence and scope: triggering SLI, affected cohort, and current customer impact.
- Name leading hypothesis, next owner/date, and one explicit ask; avoid declaring root cause prematurely.
Follow-up probes
- What details belong in the incident channel rather than the executive update?
- When do you update again?
The update must let readers understand safety, uncertainty, ownership, and next action in under a minute.
Q6What makes a code-review comment blocking rather than optional?
Strong answer outline
- Block correctness, security, privacy, data loss, broken contracts, or unacceptable operability risk.
- Explain consequence and evidence, not authority or taste.
- Offer a correction or clarify the acceptance condition; label non-blocking design ideas honestly.
Follow-up probes
- What if the style issue violates a team standard?
- How do you handle a disputed blocker?
A good answer preserves a high bar without using review as a vehicle for preference or scope expansion.
Q7Design an O(n) solution for the longest substring with at most k distinct characters.
Strong answer outline
- Clarify empty input, k ≤ 0, and character model.
- Maintain a frequency map for a window; advance right, then move left until distinct count is valid.
- Each pointer moves at most n times: O(n) time and O(min(n, alphabet)) space.
Follow-up probes
- How would you return the substring, not its length?
- Which invariant must hold before updating the maximum?
Trace a repeated-character case and a window requiring multiple left moves; state the invariant precisely.
Q8When would you choose BFS rather than DFS for a dependency graph?
Strong answer outline
- Use BFS for minimum unweighted hop count or level-order processing.
- Use DFS for exhaustive exploration, cycle detection patterns, or postorder dependencies when recursion depth is controlled.
- State directedness, cycle behavior, visited state, and memory trade-off.
Follow-up probes
- How do you produce a topological order?
- What changes for weighted edges?
Do not claim one traversal is universally faster; tie the choice to the required output and graph shape.
Q9An async endpoint fans out to three model providers. Define timeout and cancellation behavior.
Strong answer outline
- Set an end-to-end deadline and derive smaller per-attempt budgets; bound concurrency.
- Choose first-valid, quorum, or all-results semantics before selecting a primitive.
- Cancel work that can no longer affect the response, propagate cancellation after cleanup, and record provider outcomes.
Follow-up probes
- When would shielding be justified?
- How do retries interact with the deadline?
Pass if the answer covers ownership, cleanup, partial results, and retry amplification—not merely asyncio.gather.
Q10Write SQL for the three highest-cost agent runs per tenant and explain tie behavior.
Strong answer outline
- Use a window function partitioned by tenant and ordered by cost descending.
- Select
row_numberfor exactly three rows orrank/dense_rankwhen ties should be preserved. - Filter in an outer query; discuss time predicate and supporting index using
EXPLAIN.
Follow-up probes
- How do null costs sort?
- What if tenant cardinality is extremely skewed?
The query and verbal contract must agree on ties, time window, and deterministic ordering.
Q11An unfamiliar webhook service creates duplicate payroll actions. How do you debug it?
Strong answer outline
- Mitigate harmful processing, preserve evidence, and bound affected events and tenants.
- Trace provider delivery IDs through ingress, queue, worker, and database transaction; test retry and crash boundaries.
- Fix with durable idempotency at the side-effect boundary, add replay tests and duplicate-rate telemetry.
Follow-up probes
- What if the provider supplies no stable event ID?
- How do you reconcile already duplicated actions?
Distinguish duplicate delivery from duplicate effect and mitigation from root cause.
Q12How do you review code produced by a coding agent?
Strong answer outline
- Re-establish requirements and inspect the diff, dependencies, and changed trust boundaries.
- Run targeted tests, adversarial cases, static checks, and broader regression tests appropriate to risk.
- Refactor or reject code you cannot explain; record material AI assistance when policy requires it.
Follow-up probes
- Which failures are tests unlikely to reveal?
- When is generated code inappropriate?
A credible answer includes a real verification workflow and makes the engineer—not the tool—accountable.
Q13Explain a relevance-versus-latency trade-off to a non-technical customer.
Strong answer outline
- Start with the user consequence: harder questions may improve while every response could slow.
- Show representative evidence and important slices, not a single aggregate score.
- Offer a bounded decision: selective reranking, a latency guardrail, cohort rollout, and stop condition.
Follow-up probes
- What if the customer asks for “best quality” at any cost?
- How will users notice the difference?
Avoid jargon-only explanations; include choice, consequence, evidence, and control.
Q14How do you decide whether technical debt should displace roadmap work?
Strong answer outline
- Quantify recurring delivery drag, incident exposure, security risk, and option value.
- Compare remediation size and timing against roadmap impact; separate urgent containment from durable repair.
- Propose a measurable slice with owner, success signal, and stop rule.
Follow-up probes
- What if the risk has never caused an incident?
- How do you avoid a vague “20% debt” program?
Prioritize by expected consequence and leverage, not by developer annoyance or architectural purity.
Proof artifact: the senior-signal packet
Create one reviewable packet that combines leadership, writing, implementation, and explanation. Use a real project only where you can disclose it; otherwise build a clearly labeled sandbox case.
Build it
- Choose a decision such as adding a reranking stage to a multi-tenant RAG service.
- Write a two-page design memo and one ADR with criteria, alternatives, security, operations, rollout, rollback, and revisit triggers.
- Implement a small async API plus a SQL report. Add tests for invalid input, timeout, cancellation, duplicate requests, and cross-tenant access.
- Request or simulate review; classify comments and record which feedback changed the design.
- Prepare a two-minute leadership account that states personal ownership without inventing project results.
Measure it
- Memo: decision visible in the first 200 words; every risk has an owner or acceptance statement.
- Code: test pass rate, branch coverage for critical failure paths, static checks, and p50/p95 latency under a declared load.
- Communication: a reviewer can state the decision, main trade-off, and next action after one read.
- Practice: solve and explain the coding task within a fixed time; log clarification, implementation, and verification minutes separately.
If you need numbers before measurement, label them as targets. For example: “hypothetical target: p95 below 800 ms at 20 requests/second,” never as a achieved result.
Inject failure deliberately
Make one downstream call exceed its timeout, send the same idempotency key twice, cancel the client request mid-flight, and attempt a cross-tenant identifier. Verify cleanup, stable error behavior, one durable side effect, and useful trace context. Then seed an AI-generated implementation with a subtle missing tenant predicate and demonstrate that review or tests catch it.
Present it
Bring the memo, ADR, small repository, test report, one trace, review log, and a ten-minute recording. Present the decision in two minutes, code invariant in two, failure evidence in three, and learning in one; reserve two minutes for questions.
Chapter review
Leadership is decision quality plus leverage. Coding is a visible chain from contract to proof. Remote communication is successful when another person can decide or act without recovering missing context.
Glossary
- ADR
- A short, durable record of an architecture decision, its context, consequences, and revisit conditions.
- Invariant
- A property that remains true while an algorithm progresses and supports its correctness argument.
- Leverage
- An intervention that improves the output or judgment of people beyond the author’s own task.
- Structured concurrency
- A model in which concurrent tasks have explicit ownership and bounded lifetimes.
- Revisit trigger
- Observable evidence that should cause a prior decision to be re-examined.
- Window function
- A SQL calculation across related rows while retaining each input row.
Mastery checklist
- I have six distinct, truthful stories with personal actions and defensible evidence.
- I can name the decision owner, rejected option, and revisit trigger in each architecture story.
- I can write an outcome-led update and a two-page memo without a meeting transcript.
- I can distinguish blocker, important, suggestion, and nit review feedback.
- I can solve representative map/window, graph, heap, and basic DP problems while explaining invariants.
- I can reason about async cancellation, SQL ties, API idempotency, authorization, and failure tests.
- I can explain exactly how I verify AI-assisted code.
Primary sources
- Remote — Senior Forward Deployed Engineer (role active when checked).
- Sourcegraph — Agent Engineer IC4 (role active when checked).
- Python documentation — coroutines and tasks.
- PostgreSQL documentation — window functions and using EXPLAIN.
- Google Engineering Practices — code review.
- GitHub Docs — pull-request reviews.
Checked: 2026-08-04. Role requirements and URLs are volatile; re-open the official posting before applying.
CHAPTER 14 · TARGETING
Role-to-Topic Preparation Map
18 min read · 14 interview drillsLearning objectives
Role targeting is a translation problem: convert a volatile job description into a short preparation queue, a truthful evidence set, and interview hypotheses that can be tested.
- Separate live requirements from historical syllabus signals and stale search results.
- Compile a job description into capabilities, proof, stories, likely exercises, and gaps.
- Prioritize the Qdrant, Remote, Sourcegraph, Canonical, Supabase, Automattic, and Deel lanes appropriately.
- Reuse proof artifacts while changing the emphasis for each role.
- Handle location, time-zone, travel, stack, and tenure constraints honestly.
- Run a focused 48-hour preparation sprint after selecting an active role.
Compile the role before studying
A job description is neither a complete curriculum nor a keyword list. It is a noisy statement of outcomes, constraints, and organizational anxieties. Your first task is to turn each sentence into an interviewable capability.
flowchart LR
JD["Live job description"] --> AR["Atomic requirements"]
AR --> EV["Proof and stories"]
EV --> IH["Interview hypotheses"]
IH --> GP["Gap prioritization"]
GP --> Q["48-hour prep queue"]
Parse five kinds of signal
| Signal | Question to ask | Preparation output |
|---|---|---|
| Outcome | What must become measurably better? | A proof artifact and outcome story |
| Domain depth | Which concepts must survive technical probing? | A topic drill and failure diagnosis |
| Operating context | Who, what scale, what sensitivity, what ownership? | A system design with constraints |
| Behavior | How is work framed, communicated, or influenced? | A leadership story and writing sample |
| Eligibility | Can the company hire this location and can the schedule/travel work? | A go/no-go check before deep preparation |
Make requirements atomic
“Build reliable AI-enabled enterprise integrations” hides at least eight probes: authentication, webhook verification, idempotency, event ordering, tenant isolation, AI evaluation, observability, and rollout. Split it. For each atomic requirement, record one of four evidence states:
- Demonstrated: a real project and artifact can be explained.
- Practised: a sandbox artifact can be shown, clearly labeled as practice.
- Conceptual: trade-offs are understood but implementation evidence is absent.
- Unknown: neither understanding nor evidence is ready.
Live role status snapshot
The supplied syllabus was prepared from selected official pages on 2026-08-03. The following direct checks were repeated on 2026-08-04. “Active” means the official page exposes a named role and application path. “Changed” means the source remains useful but its URL or target role family has moved. “Unavailable” means the supplied role cannot be verified as an open posting; do not prepare or apply as though it were live.
| Syllabus target | Status | What the live source says | Action |
|---|---|---|---|
| Qdrant — Forward Deployed Engineer, India | Active | Remote–India listing; customer delivery, vector search, deployment, relevance, workshops, Python plus another language; Kubernetes/search systems are useful. | Run the Qdrant lane now. |
| Remote — Senior Forward Deployed Engineer | Active | Customer discovery through rollout; integrations, applied AI, evaluation, reliability, security, and structured writing. The page says ongoing applications and approximately 10% travel. | Validate practical location/time overlap in the form, then run the FDE lane. |
| Sourcegraph — Agent Engineer IC4 | Active | Agent systems, retrieval, evaluation judgment, model choice, cost/latency, staff-scope influence, and Go/TypeScript/GraphQL/Postgres/Docker. Europe/North America are preferred and at least 20 hours/week EST overlap is stated. | Proceed only if schedule overlap is genuinely workable. |
| Canonical — Cloud Solutions Architect, Alliances | Active; URL shortened | Worldwide home-based field architecture across Linux, networking, Kubernetes, public/private cloud, open-source data systems, workshops, and partner presentations; global travel is part of the role. | Use the canonical ID URL and test breadth/presentation readiness. |
| Supabase careers | Changed role family | The syllabus cited a careers hub, not one role. The current hub includes an active AI Platform Engineer and active PostgreSQL-focused roles. | Re-map to the exact opening; do not prepare for a generic “Supabase platform role.” |
| Automattic — Applied AI Engineer | Unavailable | The supplied detail URL redirects to the current jobs directory and the role was absent from the live embedded listing at check time. | Retain the product/async signals as historical practice only; wait for a new official opening. |
| Deel — Senior Backend Engineer, AI focus | Unavailable | The supplied URL returned a generic jobs shell without named posting metadata, so active availability could not be verified. | Search Deel’s current official careers site for a new requisition before tailoring. |
Choose the correct preparation lane
Use the common foundation—production AI, evaluation, backend reliability, security, and communication—but change the center of gravity. The point is not to imitate the job description. It is to surface the most relevant evidence you actually have.
Qdrant FDE: search depth in a customer room
Emphasize: dense/sparse/hybrid retrieval, HNSW, filters and payload indexes, quantization, relevance metrics, capacity, Kubernetes, migration from another search system, and workshop facilitation. Proof: one Qdrant benchmark with a golden set, recall/quality and p95 latency, plus a deployment and rollback diagram. Likely probe: a customer’s filtered search is slow and relevance degraded after migration—how do you isolate data, query, index, resource, and evaluation causes?
Remote FDE: complete enterprise delivery
Emphasize: discovery, APIs, OAuth/service accounts, webhooks, event delivery, idempotency, messy data, RAG/agents, evaluation, multi-tenancy, observability, change management, and reusable patterns. Proof: an enterprise connector and customer reference architecture with definition of done, threat model, golden set, rollout, and reconciliation. Likely probe: turn an imprecise HR or payroll workflow into a secure production plan and show how success will be measured.
Sourcegraph IC4: opinionated agent engineering
Emphasize: multi-step agent reliability, code retrieval and context packing, eval pragmatism, model selection, caching/distillation decisions, cost/latency budgets, technical direction, mentoring, and stack adaptability. The live application asks for hands-on coding-agent experience and where deterministic code or human judgment belongs. Proof: a traceable multi-step code agent, targeted evals, a cost/latency profile, and a two-minute point of view grounded in a real failure.
Canonical alliances architect: breadth with a workshop spine
Emphasize: Linux troubleshooting, DNS/TCP/TLS, Kubernetes, OpenStack, storage, cloud primitives, automation, PostgreSQL/Kafka/NGINX integration, reference architectures, and partner enablement. Proof: a hybrid-cloud reference architecture and a ten-minute workshop segment that explains failure, operations, and cost to mixed audiences. Likely probe: discover a partner environment, make an architecture defensible, and respond thoughtfully when asked outside your deepest specialty.
Supabase AI Platform Engineer: governed internal agents
The current role is a much sharper target than the syllabus’s generic Supabase row. Emphasize: event-triggered execution, durable state, human review, atomic rollback, full run logs, golden suites, CI gates, permission-enforced risk tiers, Python/GCP/infrastructure as code, MCP/API integrations, and value instrumentation. Proof: a registered-agent platform slice in which a forbidden write is impossible at the credential or tool-schema layer, not merely prohibited in a prompt.
Supabase PostgreSQL alternatives: do not confuse adjacency with fit
The active Postgres Engineer role asks for deep internals, extensions in C and Rust, planner/executor/storage mechanics, WAL/MVCC, managed deployment troubleshooting, and large-scale idempotent rollouts. General PostgreSQL, RLS, and pgvector preparation is not equivalent. Mark such requirements honestly as demonstrated, practised, conceptual, or unknown.
Automattic and Deel: preserve signals, discard stale assumptions
The unavailable Automattic posting remains a useful historical prompt for product-first AI, user-facing scale, full-stack breadth, written application quality, and accountable AI-assisted coding. The unavailable Deel posting historically emphasized Node.js, PostgreSQL, AI API integration, ETL, messy data, and document parsing. Do not quote old eligibility, tenure, salary, or location terms as current facts. If a new role appears, recompile it from zero.
Build one evidence pack, then change the lens
A credible evidence pack is a connected body of work, not eight unrelated toy repositories. One enterprise AI system can yield search, evaluation, agent, integration, reliability, security, architecture, and leadership views.
| Artifact | Qdrant | Remote | Sourcegraph | Canonical | Supabase AI |
|---|---|---|---|---|---|
| Retrieval benchmark | Primary | Supporting | Primary | Context | Supporting |
| Evaluation harness | Relevance gate | Behavior/safety | Pragmatic agent eval | Validation | Primary CI gate |
| Resilient workflow | RAG pipeline | Enterprise task | Code agent | Automation | Primary platform slice |
| Connector | Migration/import | Primary | Code host | Partner system | MCP/API tool |
| Reference architecture | Search deployment | Customer rollout | Agent service | Primary | Governed runtime |
| Incident/postmortem | Recall/latency | Duplicate/privacy | Runaway cost | Cluster/network | Forbidden action |
| Design memo | Migration choice | Definition of done | Agent boundary | Partner proposal | Autonomy classes |
Score gaps by expected interview loss
Use priority = probability of probe × consequence of weakness × improvement per hour. The arithmetic is a forcing function, not scientific precision. A missing must-have with a likely live exercise ranks above an attractive adjacent technology. Eligibility failures rank before preparation: no amount of study fixes an impossible location or schedule constraint.
Numbers without invention
For every real story, prepare the measurement definition, source, time window, baseline, and your contribution. Useful dimensions include corpus or event volume, active users/tenants, quality by slice, p50/p95/p99 latency, cost per successful task, error or duplicate rate, recovery time, delivery lead time, adoption, and customer outcome. If the source metric is inaccessible, use a bounded qualitative statement. Do not convert a hypothetical benchmark into career history.
Run a 48-hour application sprint
Once an active role passes eligibility, stop browsing broadly. Produce a role-specific packet with explicit time boxes.
- Hour 0–1 — capture: save the official URL, check date, title, location, application fields, hiring stages, and the exact text of consequential requirements.
- Hour 1–2 — compile: split requirements, label must/preferred/context, assign evidence state, and predict interview formats.
- Hour 2–4 — select: choose three artifacts and four stories; link every selection to a requirement. Drop weak or duplicative material.
- Hour 4–6 — tailor: reorder resume bullets and portfolio links without changing facts. Mirror the employer’s problem language only where it accurately describes the work.
- Hour 6–8 — close one gap: practise the highest-value missing drill: a Qdrant diagnosis, FDE discovery, agent eval, Linux troubleshooting, or governance design.
- Hour 8–10 — simulate: run a resume deep dive, role-specific technical question, architecture case, and concise written response.
- Before submission — verify: re-open the page, eligibility, requested format, and AI-use policy. Answer application questions in your own voice and obey any explicit restrictions.
The one-page role brief
Role / official URL / checked date / status:
Eligibility: location, overlap, travel, employment constraints
Top outcomes: 1 / 2 / 3
Must-have capabilities:
Three proof artifacts:
Four stories:
Likely technical and behavioral exercises:
Red gaps and honest framing:
Questions for the interviewer:
Syllabus checkpoint: prepare the numbers behind every role story
For each selected artifact or employment story, prepare a disclosure-safe evidence sheet covering data volume, user and tenant scale, quality change, latency before and after, cost change, reliability or failure-rate change, delivery time, team/stakeholder scope, and customer impact. Record the metric definition, baseline, time window, source, attribution, and confidence. If a dimension was not measured, say so and describe what evidence exists; never fill a blank with a plausible number.
These numbers should remain consistent across resume, application form, recruiter screen, portfolio, and technical loop. Rehearse both an executive statement (“what changed and why it mattered”) and an engineering drill-down (“how measured, what else changed, what failed, and what your contribution was”).
Interview playbook
Answer role-fit questions by joining role evidence to personal evidence without pretending they are identical.
Requirement → Evidence → Judgment → Relevance → Gap
- Requirement: paraphrase the employer’s real outcome.
- Evidence: give one truthful project, artifact, or practice example.
- Judgment: explain the consequential trade-off or failure handled.
- Relevance: connect it specifically to this environment.
- Gap: name any meaningful difference and how you would de-risk it.
Role-specific openings
- Qdrant: lead with relevance and performance evidence, then customer delivery.
- Remote: lead with an ambiguous enterprise outcome owned through rollout.
- Sourcegraph: lead with an opinionated agent decision backed by eval, cost, and failure evidence.
- Canonical: lead with breadth, troubleshooting method, and workshop clarity.
- Supabase AI: lead with runtime/evaluation/governance mechanisms, especially structurally enforced permissions.
Common traps
- Keyword recitation with no project decision or artifact.
- Describing a closed role as active because a cached search result still exists.
- Hiding a stack, tenure, location, or schedule gap until late in the process.
- Changing claims between resume, application form, and interview.
- Preparing every target equally and becoming shallow in all of them.
Question bank
These questions test whether targeting is evidence-based rather than cosmetic.
Q1How do you turn “own production agent quality” into a preparation plan?
Strong answer outline
- Split ownership into dataset design, behavioral/safety checks, judge calibration, release thresholds, telemetry, and incident response.
- Map each item to demonstrated, practised, conceptual, or unknown evidence.
- Build one gated change with a failure case and prepare the release decision.
Follow-up probes
- Which part is most likely to be interviewed live?
- What would count as production ownership?
The plan must produce evidence and judgment, not a reading list of evaluation tools.
Q2How do you distinguish a must-have from aspirational job-description language?
Strong answer outline
- Weight explicit “must,” repeated responsibilities, first-90-day outcomes, application questions, and interview stages.
- Treat “nice to have” and broad company context differently, while noting hidden dependencies.
- Validate ambiguous requirements with the recruiter rather than silently assuming.
Follow-up probes
- What if the title and responsibilities imply different seniority?
- How do application questions change weighting?
Show a repeatable method and preserve uncertainty; confident guessing does not pass.
Q3What three proofs would you lead with for the active Qdrant FDE role?
Strong answer outline
- A measured hybrid-search benchmark with relevance slices and latency/recall trade-offs.
- A deployed Qdrant lab covering filters, sizing, observability, backup, and failure diagnosis.
- A customer-style migration/workshop artifact with requirements, rollout, rollback, and acceptance tests.
Follow-up probes
- What if your production system used another vector database?
- Which Qdrant-specific gap must be closed?
At least one proof must show customer communication and one must show Qdrant-specific technical depth.
Q4How would you answer Sourcegraph’s question about where coding agents shine?
Strong answer outline
- Name a bounded task where search, iteration, and verification make an agent useful.
- Name a failure you observed and the control added: tests, permissions, budget, deterministic step, or human gate.
- State a principled boundary for irreversible, ambiguous, or high-consequence actions.
Follow-up probes
- What have you changed your mind about?
- How did you measure usefulness?
The answer needs hands-on mechanics and evidence, not a generic pro/anti-agent opinion.
Q5What would you practise for a Remote FDE technical case?
Strong answer outline
- Run discovery around users, systems of record, data sensitivity, workflow, failure tolerance, and definition of done.
- Design auth, events, idempotency, AI behavior/eval, observability, and reconciliation.
- Close with phased rollout, customer ownership, reusable components, and measurable outcome.
Follow-up probes
- Where is a human approval mandatory?
- What becomes product versus customer-specific code?
A complete answer spans discovery through operations; an architecture diagram alone is insufficient.
Q6How do you prepare for Canonical’s breadth without memorizing every product?
Strong answer outline
- Build durable layers: Linux, networking, compute, storage, Kubernetes, data services, IAM, and operations.
- Practise a troubleshooting tree and reference architecture that make assumptions explicit.
- Learn Canonical-specific components enough to position them honestly and ask precise follow-ups.
Follow-up probes
- What do you say when you do not know an answer?
- How do you prepare a partner workshop?
Demonstrate method, breadth, and communication—not bluffing or a list of product definitions.
Q7What is the strongest proof for Supabase’s current AI Platform Engineer role?
Strong answer outline
- A durable event-triggered agent run with restart, human approval, rollback, and reconstructable logs.
- A golden/safety suite that blocks a known regression in CI.
- A permission model where a forbidden commitment write has no executable path, plus cost/value instrumentation.
Follow-up probes
- How do you grade the evaluator?
- How do you limit interruption cost to people?
The artifact must enforce governance in code; a prompt saying “do not write” fails the bar.
Q8Should you keep preparing for Automattic’s unavailable Applied AI role?
Strong answer outline
- Stop role-specific application work because the official detail page no longer verifies an opening.
- Retain durable product-AI, full-stack, async-writing, and accountable coding-agent drills if useful for other targets.
- Set a lightweight careers-page check rather than repeatedly tailoring to a closed requisition.
Follow-up probes
- Which materials can be reused elsewhere?
- What evidence would restart the application sprint?
Separate transferable learning from authorization to claim a current opening.
Q9A closed Deel listing historically requested more tenure than your verified profile shows. How should you handle a future similar role?
Strong answer outline
- First verify the new official requisition; do not transfer the old threshold automatically.
- State actual dates and scope consistently, with no rounding designed to cross a threshold.
- If eligible to apply, lead with relevant production ownership while accepting that tenure may remain a hard filter.
Follow-up probes
- Would you address the gap in a cover letter?
- When should you self-select out?
Integrity and consistency are mandatory; “compensate” never means rewriting chronology.
Q10How do you choose between two active roles this week?
Strong answer outline
- Apply eligibility gates: location, schedule, travel, work authorization, and hard experience requirements.
- Score must-have evidence coverage, gap severity, role interest, and artifact reuse.
- Choose one primary lane for deep preparation and time-box the second.
Follow-up probes
- How do you avoid optimizing only for apparent fit?
- What makes you revisit the choice?
The choice should be traceable to evidence and constraints, not brand preference or fear.
Q11How can the same RAG project support Qdrant, Remote, and Sourcegraph interviews?
Strong answer outline
- Qdrant lens: retrieval variants, index/filter tuning, deployment, and migration.
- Remote lens: customer requirement, integration, tenancy, evaluation, rollout, and operations.
- Sourcegraph lens: multi-step agent/context decisions, eval pragmatism, cost/latency, and technical leadership.
Follow-up probes
- Which facts must stay identical across versions?
- When does reframing become misrepresentation?
Change emphasis, not history, metrics, technology, or ownership.
Q12How do you discuss a required technology you have only read, not operated?
Strong answer outline
- Label the evidence state directly and avoid substituting adjacent experience as identical.
- Explain the transferable mental model and a concrete sandbox exercise completed.
- State the production unknowns and a focused ramp plan tied to the role.
Follow-up probes
- What adjacent experience is genuinely relevant?
- Which claim would you refuse to make?
The interviewer should be able to distinguish knowledge, practice, and production ownership.
Q13How do location and time-zone requirements affect fit scoring?
Strong answer outline
- Treat explicit applicant countries, overlap hours, and travel as eligibility or sustainability gates.
- Verify wording and form options on the live official page; ask recruiting when ambiguous.
- Do not promise an unhealthy schedule merely to pass screening.
Follow-up probes
- What does Sourcegraph’s EST overlap imply operationally?
- How do you record an unresolved eligibility question?
A good answer prioritizes legal and sustainable reality before topic overlap.
Q14How do you present impact when you do not have a trustworthy baseline metric?
Strong answer outline
- State that the baseline was not instrumented and do not manufacture a delta.
- Use available evidence: before/after incidents, adoption, qualitative feedback, or a later measurement with its limits.
- Explain the instrumentation you would add and keep sandbox targets explicitly hypothetical.
Follow-up probes
- Can a testimonial be evidence?
- How do you attribute a team outcome?
Uncertainty must remain visible; precision without provenance is a negative signal.
Proof artifact: the live role compiler
Create a versioned dossier for one active role. It should let another reviewer reproduce why you prioritized certain preparation and whether every application claim is supported.
Steps
- Capture the official posting as a dated PDF or text snapshot for personal analysis, respecting site terms; record the live URL and check time.
- Extract atomic requirements into a spreadsheet or Markdown file. Tag outcome/domain/context/behavior/eligibility, must/preferred, and evidence state.
- Link three artifacts and four stories. For each, record exact personal action, source of any metric, disclosure boundary, and one gap.
- Generate a one-page role brief, six likely technical probes, three questions for the employer, and a 48-hour preparation queue.
- Have a reviewer compare resume, application answers, dossier, and spoken story for consistency.
Metrics
- 100% of hard eligibility items explicitly resolved as pass, fail, or recruiter question.
- Every must-have mapped to evidence state; no unmarked blanks.
- At least three high-probability requirements supported by inspectable artifacts.
- Every quantitative claim has source, definition, time window, and attribution note.
- A reviewer can explain the top preparation priority and why in under two minutes.
Deliberate failure injection
Replace the saved posting with a closed-role shell, change a location requirement, and insert one unsupported resume claim. Your workflow should flag missing posting metadata, invalidate the prior eligibility decision, and fail the evidence audit. Then simulate a renamed role with the same URL and require a fresh diff rather than silently trusting the old brief.
What to present
Show the dated source register, requirement matrix, evidence links, one-page brief, change diff, and audit result. In an interview, present only your own evidence—not the internal scoring machinery—unless asked how you prepared.
Chapter review
Targeting is disciplined selectivity. Verify the opening, gate eligibility, translate verbs into capabilities, connect truthful evidence, and spend preparation time where it changes likely interview performance.
Glossary
- Atomic requirement
- One independently assessable capability or constraint extracted from a broader sentence.
- Evidence state
- Demonstrated, practised, conceptual, or unknown readiness for a requirement.
- Eligibility gate
- A location, schedule, legal, travel, or hard-experience condition evaluated before deep preparation.
- Role compiler
- The process that converts a live job description into evidence, gaps, interview hypotheses, and actions.
- Historical signal
- A useful topic from a closed or changed role that must not be represented as a current requirement.
- Evidence provenance
- The source, definition, time window, and attribution behind a claim.
Mastery checklist
- I verify title, description, location, and application path—not merely HTTP status.
- I can compile a role into atomic outcomes, depth, context, behavior, and eligibility.
- I know which of the seven syllabus sources are active, changed, or unavailable as of the check date.
- I can name the three strongest artifacts and four stories for my primary active target.
- I distinguish production evidence, sandbox practice, conceptual knowledge, and unknowns.
- I can state stack, tenure, location, travel, and schedule gaps without distortion.
- I re-check the official page and application instructions immediately before submission.
Official role sources and status
- Qdrant — Forward Deployed Engineer, India: active.
- Remote — Senior Forward Deployed Engineer: active.
- Sourcegraph — Agent Engineer IC4: active.
- Canonical — Cloud Solutions Architect, Alliances: active; canonical URL changed from the supplied long slug.
- Supabase careers, AI Platform Engineer, and Postgres Engineer: careers hub changed into specific active targets.
- Automattic — Applied AI Engineer supplied URL: unavailable; redirects to the jobs directory and was absent from its current listing.
- Deel — Senior Backend Engineer, AI focus supplied URL: unavailable; named posting could not be verified.
Checked: 2026-08-04. Re-check every role immediately before investing preparation time or submitting an application.
CHAPTER 15 · EXECUTION
A Practical 12-Week Execution Sequence
19 min read · 14 interview drillsLearning objectives
This sequence converts twelve weeks of two-hour work blocks into an interview-ready evidence trail. The unit of progress is a tested proof, not a completed playlist.
- Run a repeatable two-hour daily cadence with a weekly acceptance test.
- Build retrieval, Qdrant, agent, evaluation, integration, data, platform, reliability, and security proofs in dependency order.
- Measure quality, latency, cost, and failure behavior without manufacturing outcomes.
- Use deliberate failures to turn implementations into debugging and incident stories.
- Adapt the schedule when a live interview arrives or a week slips.
- Finish with a coherent portfolio and full-loop interview simulation.
Use an execution operating system
The syllabus proposes two focused hours per day. Protect that constraint: it forces selection, exposes over-engineering, and makes twelve weeks sustainable. A typical five-day week yields ten core hours; keep any sixth session as recovery or mock-interview time rather than silently expanding scope.
The daily 110 + 10 block
| Minutes | Work | Output |
|---|---|---|
| 0–10 | Read yesterday’s evidence and choose one falsifiable objective | One sentence: “By the end, I will know whether…” |
| 10–80 | Implement, benchmark, diagnose, or rehearse | Code, test, trace, query plan, diagram, or recording |
| 80–105 | Test the edge or inject the planned failure | Observed behavior and correction |
| 105–110 | Commit or checkpoint the artifact | Reproducible state |
| 110–120 | Write the evidence log and next smallest action | Metric, uncertainty, decision, next step |
The weekly rhythm
- Monday — baseline: define contract, dataset/load, success measure, and simplest working version.
- Tuesday — controlled change: change one consequential variable.
- Wednesday — failure: inject a realistic fault and trace it across the system.
- Thursday — evidence: rerun, compare, analyze slices, and capture limitations.
- Friday — explain: produce a one-page report and a five-to-ten-minute verbal walkthrough.
Before week one, record a baseline mock: one 35-minute architecture, one 30-minute coding problem, one SQL task, and two leadership answers. Score clarification, correctness, depth, failure coverage, and communication from 0–3. This is diagnostic, not a career metric.
Weeks 1–3: retrieval foundations and Qdrant depth
Search comes first because later agent and evaluation work needs a measurable grounding layer. Use one versioned corpus and query set across all three weeks so improvements remain comparable.
Week 1 — establish the retrieval laboratory
Core work: choose a disclosure-safe document corpus; define query intents and judgments; implement lexical, dense, and hybrid retrieval; calculate recall@k, MRR or nDCG as appropriate; capture p50/p95 latency and index size. Include identifier-heavy, semantic, filtered, long-document, and no-answer slices.
Acceptance test: a single command rebuilds the index and produces a report with dataset version, configuration, per-slice quality, latency, and at least five inspected errors. Explain why each metric matches the user task.
Failure: corrupt or omit a subset of relevance judgments. The report must expose dataset coverage rather than quietly comparing incomparable runs.
Week 2 — tune the complete RAG retrieval path
Core work: compare two chunking strategies; query rewrite only where justified; dense/sparse fusion; a reranker; context packing with source boundaries; and incremental indexing. Change one variable per run. Use an error taxonomy such as retrieval miss, wrong rank, filter exclusion, bad chunk boundary, stale data, and correct evidence lost during packing.
Acceptance test: recommend a pipeline for at least two query slices and reject one apparently better aggregate configuration because of latency, cost, or a critical regression. Keep all numbers as measured sandbox results, never employment claims.
Failure: add near-duplicate documents and one stale version. Demonstrate how deduplication, metadata, or recency policy affects citations.
Week 3 — operate Qdrant, not just call it
Core work: model collections, named dense/sparse vectors and payloads; create payload indexes for filters; tune HNSW/search parameters; test quantization; exercise snapshot/restore; and reason about shards, replicas, and multi-tenant layout. Qdrant’s current documentation describes dense+sparse fusion and multi-stage queries in the Hybrid and Multi-Stage Queries guide and memory/performance choices in Quantization.
Acceptance test: publish a before/after matrix for quality, filtered-query p95, throughput, memory or storage, configuration, and workload. Restore from a snapshot into a clean environment and verify document count plus sampled results.
Failure: run a selective filter without the appropriate payload index, then interrupt a migration or restore. Capture symptoms, diagnosis, mitigation, and the safer runbook.
Weeks 4–6: reliable agents, evaluation, and enterprise integration
These weeks join probabilistic behavior to deterministic controls. Use one workflow—for example, an evidence-backed enterprise request triage—so the evaluation and connector are part of the same system.
Week 4 — make the workflow resumable and bounded
Core work: draw explicit states and terminal outcomes; separate deterministic routing, model decisions, tools, and human approval; add schema validation, per-step timeouts, retry budgets, maximum steps, cost budget, checkpointing, idempotent side effects, and trace correlation. LangGraph’s current persistence documentation explains threads, checkpoints, state history, and replay concepts; verify the installed version before coding.
Acceptance test: terminate the process after a checkpoint, restart with the same run identity, and prove that completed side effects are not repeated. Every terminal state must be named: success, rejected, budget exhausted, invalid input, dependency failure, or human cancellation.
Failure: return malformed tool output, make a tool time out, and create a cycle. Show validation, bounded retry, and termination rather than an infinite repair loop.
Week 5 — build an evaluation gate that can say no
Core work: version a golden set; split retrieval, generation, tool-use, task completion, safety, latency, and cost measures; create deterministic assertions; write a human rubric; calibrate any model judge against labeled cases; add slice thresholds and a CI report. Include prompt injection, cross-tenant, no-answer, and malformed-input cases.
Acceptance test: seed a known regression and prove the gate blocks it for the correct reason. Then seed a harmless change and ensure the suite does not fail noisily. Document override authority and the evidence required.
Failure: deliberately bias a judge with order or verbosity and show disagreement against human labels. Tighten rubric or use a deterministic check where possible.
Week 6 — integrate across unreliable boundaries
Core work: design an OAuth/service-account connection, signed webhook ingestion, event queue, idempotent worker, outbox or equivalent handoff, dead-letter path, backfill, replay, and reconciliation. Define API versions, rate limits, tenant scoping, audit events, and long-running job status.
Acceptance test: the same event delivered repeatedly produces one intended side effect; a missed webhook is found by reconciliation; a dead-lettered item can be repaired and replayed with an audit trail.
Failure: crash after the external side effect but before local acknowledgement. Explain why “exactly once” is not a magic transport property and how idempotency plus reconciliation controls the outcome.
Weeks 7–9: data, platform, reliability, and security
Now deploy the same system under realistic operational constraints. The interview goal is to explain what happens after the happy-path demo.
Week 7 — PostgreSQL and replayable ingestion
Core work: normalize the transactional core; use JSONB deliberately; add indexes from query shapes; inspect EXPLAIN (ANALYZE, BUFFERS) safely on test data; exercise transactions, locks, and a deadlock; apply tenant isolation and row-level security; build a checkpointed ingestion path with validation, deduplication, lineage, and quarantine.
Acceptance test: show the slow query, plan, hypothesis, change, new plan, and trade-off. Replay an ingestion partition without duplicating rows. Prove a cross-tenant negative test. PostgreSQL’s official Using EXPLAIN chapter is the primary reference for reading plans.
Failure: feed a malformed PDF-derived record, duplicate an input file, and create lock contention. The pipeline should quarantine or retry without losing lineage.
Week 8 — deploy with explicit lifecycle behavior
Core work: build a small multi-stage image; define configuration and workload identity; deploy the API, worker, and dependencies; set requests/limits; create startup, readiness, and liveness probes with different semantics; add autoscaling assumptions, rolling update, rollback, and a rough cost model. The Kubernetes guide distinguishes how startup, readiness, and liveness probes affect container lifecycle and traffic.
Acceptance test: a slow-starting process is not killed prematurely, an unready instance receives no traffic, a deadlocked process recovers, and a bad release rolls back. State which behavior each probe is designed to observe.
Failure: point readiness at a fragile downstream dependency and observe the amplification. Redesign it to represent whether this instance can serve its contract without creating a fleet-wide outage.
Week 9 — observe, budget, threaten, and recover
Core work: instrument request and agent-step traces, structured logs, metrics, quality samples, token/cost use, and correlation IDs; define user-centered SLIs/SLOs; create actionable alerts; threat-model prompt injection, retrieval poisoning, data leakage, tool abuse, secrets, and cross-tenant access. The OpenTelemetry Python guide provides current instrumentation examples; the OWASP GenAI prompt-injection entry is a useful adversarial checklist.
Acceptance test: use one trace to locate a latency or failure cause, run an incident drill, and produce a short postmortem with detection, mitigation, root cause, correction, and learning. Show preventive and detective controls for a high-risk AI path.
Failure: disable a dependency, exhaust a quota, and place hostile instructions in retrieved content. Verify degradation, bounded retries, human-safe messaging, and alerts tied to action.
Weeks 10–12: convert engineering proof into interview performance
The final phase does not add a new platform. It compresses what you built into timed designs, coding fluency, truthful stories, and role-specific simulations.
Week 10 — four timed system designs
Design an enterprise RAG platform, a governed agent platform, an enterprise connector, and an evaluation platform. Use 45 minutes each: discovery, estimates, APIs/data flow, security/tenancy, failure handling, SLO/observability, cost, rollout/migration/rollback, and rejected alternatives. Record the session and score whether assumptions preceded components.
Acceptance test: each design contains a definition of done, one scale estimate, one critical trust boundary, three failure modes, and a phased rollout. Re-run the weakest design later without reading the first solution.
Failure: have a mock interviewer change a core constraint at minute 20—for example, data cannot leave a private network. Adapt the design without discarding the entire reasoning chain.
Week 11 — coding, SQL, API, and debugging under time
Run mixed sets: maps/windows, stacks/intervals, trees/graphs, heap/greedy, and one basic DP; SQL joins/aggregation/windows/plans; an async API; and an unfamiliar bug. Five high-quality algorithm problems in the week is the syllabus baseline, but quality means a second attempt, edge cases, invariant, and complexity—not merely an accepted submission.
Acceptance test: keep a scorecard for clarification, pattern selection, correctness, tests, complexity, and communication. Re-solve misses from a blank editor after 48 hours. For practical work, include timeout, authorization, idempotency, and failure tests.
Failure: introduce a misleading test, a cancellation bug, and a SQL tie ambiguity. Practise detecting a flawed premise rather than coding around it.
Week 12 — full-loop simulation and application pack
Run at least two role-specific loops: recruiter pitch, resume deep dive, technical discussion, coding/practical task, system design or FDE case, cross-functional scenario, leadership story, and written follow-up. Use the live role compiler from Chapter 10 and re-check availability before each simulation.
Acceptance test: every major topic—search, agents, evaluation, integration, data/platform, reliability, security, customer architecture, leadership—has one credible story or artifact. Gaps are stated, not hidden. The second simulation shows a specific improvement from the first.
Failure: remove a favorite story, challenge a metric’s provenance, and ask for a design in a different domain. The pack should survive without memorized wording or invented precision.
Control the plan instead of obeying it blindly
A sequence is useful only while it targets current constraints. Review the evidence every Friday and change the next week only for a recorded reason.
The weekly scorecard
| Dimension | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| Reproducibility | No artifact | Works only locally/manual | Documented repeat | Clean rebuild and versioned inputs |
| Measurement | Opinion | One raw number | Baseline plus slices | Trade-off and uncertainty |
| Failure depth | Happy path | Error observed | Diagnosed and tested | Recovery/runbook/telemetry |
| Explanation | Cannot explain | Feature tour | Decision and trade-off | Adapts under probing |
| Role relevance | Unlinked | Keyword overlap | Mapped requirement | Direct evidence for active target |
When time is lost
Do not compress every missed task into the next week. Preserve dependency order and acceptance tests. Drop polish, extra frameworks, frontend work, and duplicate artifacts first. If an interview arrives early, branch into a three-day role sprint: live requirements, top two proof gaps, one mock, then return to the plan.
What not to over-prepare
For these target roles, foundation-model pretraining pipelines, optimizer derivations, distributed training internals, CUDA kernels, state-space-model mathematics, and transformer history have lower expected return unless a live posting explicitly asks. Learn enough to make an informed build/use/fine-tune decision, then return to production retrieval, evaluation, integration, reliability, and architecture.
Interview playbook
Use the twelve-week work as evidence, not as a claim of prior production experience. Label sandbox work explicitly and connect it to real judgment you can defend.
Proof answer: Contract → Baseline → Change → Failure → Evidence → Decision
- Contract: user task, constraints, and success definition.
- Baseline: simplest measured system and dataset/load.
- Change: one variable and the hypothesis behind it.
- Failure: injected or observed fault and diagnosis.
- Evidence: quality/latency/cost/operability result with limits.
- Decision: ship, reject, narrow, or gather more evidence.
Common traps
- Presenting twelve mini-demos with no shared dataset, architecture, or narrative.
- Reporting only aggregate quality and average latency.
- Calling an injected sandbox failure a real production incident.
- Changing multiple variables and attributing the result to one of them.
- Spending the final week polishing slides instead of running mocks.
- Using a library’s feature names as a substitute for understanding lifecycle and failure behavior.
Question bank
Use these as weekly retrospectives and mock-interview prompts.
Q1How do you make progress in only two focused hours per day?
Strong answer outline
- Define one falsifiable session outcome and start from yesterday’s checkpoint.
- Reserve time for an edge/failure and an evidence log, not only implementation.
- Use weekly acceptance tests and drop optional polish when scope grows.
Follow-up probes
- What work is deliberately excluded?
- How do you recover after a missed day?
The answer needs a mechanism for focus, evidence, and scope control—not “be disciplined.”
Q2What makes a weekly proof complete?
Strong answer outline
- A reproducible artifact with versioned inputs and a clear contract.
- A baseline, relevant measures and slices, plus at least one diagnosed failure.
- A concise decision, limitations, and explanation another engineer can follow.
Follow-up probes
- Is a notebook sufficient?
- What if the experiment disproves the hypothesis?
A negative result can pass; an irreproducible impressive result cannot.
Q3Your hybrid retriever improves aggregate nDCG but hurts identifier queries. What do you do?
Strong answer outline
- Validate judgments and isolate the affected slice rather than accepting the aggregate.
- Inspect fusion, sparse candidate depth, tokenization, filters, and reranking behavior.
- Consider query classification or weighted routing; decide against user-critical thresholds.
Follow-up probes
- How do latency and cost enter the decision?
- What test prevents recurrence?
Protect important slices and avoid tuning blindly to a single summary metric.
Q4Which Qdrant failure would you deliberately practise in week 3?
Strong answer outline
- Use a selective metadata filter without its payload index and measure the symptom.
- Inspect query shape, index configuration, load, and quality before changing parameters.
- Add the index, rerun the same workload, and document operational/memory consequences.
Follow-up probes
- How would you test snapshot recovery?
- What if latency improves but recall drops?
Failure, diagnosis, controlled change, and comparable evidence must all appear.
Q5How do you prove an agent workflow is resumable rather than merely retryable?
Strong answer outline
- Persist state and step identity at defined boundaries with idempotent external effects.
- Kill the process after a side effect but before completion, then resume the same run.
- Verify the effect is not repeated and the trace shows checkpoint history and recovery.
Follow-up probes
- Which operations cannot be replayed safely?
- How do you version state?
Restart evidence and side-effect semantics are required; generic retry code is not enough.
Q6Your evaluation gate blocks harmless prompt edits. How do you reduce noise without weakening it?
Strong answer outline
- Classify failures: flaky infrastructure, judge variance, ambiguous rubric, or real slice sensitivity.
- Move stable requirements to deterministic checks, calibrate judges, and require repeated or confidence-aware evidence where appropriate.
- Keep critical safety failures hard-blocking and document override authority.
Follow-up probes
- When is an override acceptable?
- How do you detect a weak judge?
The answer must preserve risk-based rigor while improving signal-to-noise.
Q7A webhook worker crashes after writing externally but before acknowledging the event. What should the lab demonstrate?
Strong answer outline
- Redelivery is expected; use a durable idempotency identity at the side-effect boundary.
- Record attempts and outcome so the worker can distinguish retry, conflict, and unknown state.
- Use reconciliation for ambiguity and prove repeated delivery does not multiply the intended effect.
Follow-up probes
- What if the external API lacks idempotency support?
- Where does the transaction end?
Do not promise transport-level exactly-once delivery; show tolerated duplication and repair.
Q8How do you show that a PostgreSQL index actually helped?
Strong answer outline
- Fix representative data, parameters, cache caveats, and query contract.
- Capture plans and timings before/after; inspect estimates, scans, rows, buffers, sort, and selectivity.
- State write/storage cost and whether the improvement holds across important parameter values.
Follow-up probes
- Why can one
EXPLAIN ANALYZEmislead? - What if estimates are wrong?
Evidence must include plan interpretation and trade-off, not merely lower elapsed time once.
Q9What is wrong with using the same Kubernetes endpoint for liveness and readiness?
Strong answer outline
- Readiness answers whether this instance should receive traffic; liveness answers whether restart may repair it.
- A downstream outage in both probes can remove traffic and restart every pod, amplifying failure.
- Design each probe around distinct recovery semantics and use startup for slow initialization.
Follow-up probes
- Should readiness check the database?
- What does a startup probe protect?
The answer must reason from orchestrator action, not endpoint naming convention.
Q10How do you test indirect prompt injection in week 9?
Strong answer outline
- Place hostile instructions in retrieved content and define a prohibited tool/data outcome.
- Restrict tool permissions and data scope structurally; separate instructions from untrusted content.
- Trace the run, verify no forbidden effect, and add the case to the safety suite and alerting.
Follow-up probes
- Why is output filtering insufficient?
- Which human gate remains?
Test the effect boundary and controls; merely detecting suspicious text does not pass.
Q11What should a week-11 coding scorecard reveal?
Strong answer outline
- Separate clarification, pattern selection, implementation correctness, tests, complexity, and communication.
- Tag failures by cause rather than only problem topic.
- Schedule blank-editor reattempts and track whether the cause disappears after 48 hours.
Follow-up probes
- How many problems are enough?
- What if speed rises but explanation worsens?
The scorecard must drive a targeted next drill, not become a vanity count of solved problems.
Q12You lose an entire week. How do you re-plan?
Strong answer outline
- Keep dependency order and identify which acceptance tests serve the active target.
- Drop duplicate artifacts, extra tools, frontend polish, and low-return reading first.
- Merge only compatible work—for example, use the connector as the week-9 failure target—then record the trade-off.
Follow-up probes
- Which week must not be skipped?
- How do you prevent permanent catch-up mode?
Protect evaluation, failure, and final simulation; do not compress every task into longer days.
Q13Why are deep pretraining and CUDA lower priority in this plan?
Strong answer outline
- The selected live roles emphasize applied systems, retrieval, evaluation, integration, platform operation, and customer/technical leadership.
- Preparation time should follow likely probes and evidence gaps, not field prestige.
- Reprioritize immediately if a new official role makes training or kernel depth a core outcome.
Follow-up probes
- What model knowledge remains necessary?
- When would fine-tuning enter the plan?
Explain opportunity cost from current role evidence; do not dismiss the technical value of the topics.
Q14What is the final readiness signal after week 12?
Strong answer outline
- One credible story or artifact for every major syllabus domain, with truthful evidence state.
- Two complete role-specific mocks showing correction of identified weaknesses.
- A current role brief, consistent claims, a gap statement, and the ability to adapt under follow-up.
Follow-up probes
- Which weakness would delay an application?
- What does “credible” mean for sandbox work?
Readiness is demonstrated under simulation, not inferred from finishing the calendar.
Proof artifact: the twelve-week evidence repository
Build one repository or portfolio folder that tells a coherent engineering story. Keep sensitive employer material out; use disclosure-safe or synthetic data and label sandbox claims.
Structure and steps
README.md # user problem, architecture, how to reproduce
evidence/weekly-log.md # hypothesis, result, limitation, next decision
retrieval/ # corpus manifest, judgments, benchmark configs
agent/ # state model, permissions, checkpoints, evals
connector/ # contracts, idempotency, replay, reconciliation
platform/ # deployment, probes, telemetry, threat model
designs/ # four timed architecture records
interview/ # story index, scorecards, mock retrospectives
- Tag a baseline before each controlled change and record dependency versions.
- Automate one clean setup and one evaluation command; include expected runtime and resource needs.
- Add a decision ledger that links every claimed improvement to raw evidence and limitations.
- Record five short walkthroughs: retrieval, agent/eval, integration, operations/security, and architecture.
- Create a final role index showing which files support which active requirement.
Metrics
- Twelve weekly acceptance tests with pass/fail and evidence link.
- At least one quality, latency, cost/resource, and reliability measure where relevant.
- At least one important slice and one limitation in every benchmark report.
- Time-to-reproduce from a clean environment and time-to-explain in a ten-minute walkthrough.
- Mock score improvement by dimension; do not turn practice scores into employment claims.
Deliberate failure injection
Maintain a failure manifest: missing judgments, stale duplicate documents, slow filtered search, killed agent process, malformed tool output, biased judge, duplicate webhook, malformed ingestion record, lock contention, bad readiness dependency, provider outage, and indirect prompt injection. For each, capture expected behavior, actual signal, containment, repair, regression test, and remaining risk.
What to present
Open with a one-page architecture and evidence map. Demonstrate one clean run and one failure/recovery, then show the decision ledger and a role-specific index. Keep the full repository available for follow-up, but lead with the smallest evidence that answers the interviewer’s question.
Chapter review
The twelve weeks form a dependency chain: measurable retrieval, bounded agents, gates, reliable boundaries, operable deployment, and finally interview compression. The schedule is successful when it produces repeatable evidence and better decisions.
Glossary
- Acceptance test
- The observable condition that must pass before a week or artifact is considered complete.
- Evidence log
- A dated record of hypothesis, configuration, result, limitation, decision, and next action.
- Controlled change
- An experiment that changes one consequential variable while keeping comparison conditions stable.
- Failure manifest
- A catalog of injected faults, expected behavior, observed evidence, and recovery controls.
- Blank-editor reattempt
- Solving a missed problem again from scratch after delay to test retained reasoning.
- Scope guard
- An explicit rule for what is dropped when time or complexity exceeds the plan.
Mastery checklist
- I can state the weekly proof and acceptance test for all twelve weeks.
- My corpus, judgments, configurations, loads, and dependency versions are recorded.
- Every important artifact includes a deliberate failure and recovery evidence.
- I distinguish sandbox results from production experience every time.
- I use role relevance and evidence gaps to adjust the sequence.
- I have four timed designs, a coding/SQL scorecard, six truthful stories, and two full mocks.
- I can reproduce and explain the final repository without hidden manual steps.
Primary technical sources
- Qdrant — Hybrid and Multi-Stage Queries and Quantization.
- LangGraph — Persistence.
- PostgreSQL — Using EXPLAIN.
- Kubernetes — Configure liveness, readiness, and startup probes.
- OpenTelemetry — Python getting started.
- OWASP GenAI — Prompt Injection.
Checked: 2026-08-04. Product behavior and APIs change; pin versions in the repository and re-check official documentation before running each lab.
CHAPTER 16 · FINAL GATE
Official References & Final Readiness
21 min read · 14 interview drillsLearning objectives
The final gate is an audit, not a pep talk. Verify the opportunity, verify every claim, expose gaps, and rehearse the complete interview loop against the role that exists now.
- Determine whether an official role is active, changed, unavailable, or ambiguous.
- Maintain a dated source register with location, schedule, travel, stack, and application constraints.
- Reconcile the supplied syllabus with live job-description changes.
- Audit every resume, portfolio, metric, and leadership claim for provenance and ownership.
- Apply a role-specific readiness gate across technical, customer, coding, leadership, and writing signals.
- Run a full-loop simulation and decide honestly whether to apply now, prepare briefly, or stop.
Treat job descriptions as volatile production inputs
A role page can change between preparation and submission. Titles move, applicant-tracking systems retain empty shells, location lists narrow, and application questions introduce constraints that were not visible in the main description. The correct response is a source protocol.
flowchart LR
OC["Official careers source"] --> NP["Named posting"]
NP --> AP["Application path"]
AP --> CT["Constraints extraction"]
CT --> DC["Dated capture"]
DC --> DF["Semantic diff"]
DF --> GD["Go or no-go decision"]
Source hierarchy
- Live official company career page or company-linked ATS: authority for current title, responsibilities, eligibility, and application instructions.
- Official company handbook and technical documentation: useful for durable work practices and product behavior, but not proof that a requisition is open.
- Supplied syllabus snapshot: authoritative for this handbook’s intended curriculum, not current availability.
- Search result, aggregator, repost, or social post: discovery only. Use it to find the official page, never as final authority.
The four-state status model
- Active: named role, substantive description, current location/arrangement, and application path are present.
- Changed: the role or family remains relevant, but title, URL, scope, location, or target opening differs from the syllabus.
- Unavailable: the supplied role redirects to a directory, returns a generic shell, is absent from the current list, or has no usable application path.
- Ambiguous: evidence conflicts; pause role-specific submission and ask the employer or wait for the official system to resolve.
Minimum verification record
company / role / requisition ID:
official URL / final redirected URL:
checked timestamp and timezone:
status and evidence:
locations / overlap / travel / employment type:
must-have outcomes and stack:
application stages and special instructions:
changes from previous capture:
next re-check date / owner:
Official role register: checked 2026-08-04
This register distinguishes live evidence from syllabus history. Re-check it immediately before applying; the date is part of every status statement.
Active roles
| Official source | Verified signal | Constraint to confirm | Final readiness emphasis |
|---|---|---|---|
| Qdrant — Forward Deployed Engineer, India | Named Remote–India posting with application path; search/customer delivery, performance, relevance, workshops, Python and another language. | The listing’s applicant location is India while the prose also references APAC; use the form as final eligibility evidence. | Vector/search diagnosis plus customer workshop and deployed proof. |
| Remote — Senior Forward Deployed Engineer | Named active page with apply action; full lifecycle from discovery to secure, observable AI/integration rollout; page says applications are ongoing. | Customer-hour overlap and roughly 10% travel; validate country-specific hiring in the application flow. | Enterprise connector, definition of done, evaluation, rollout, operations, structured writing. |
| Sourcegraph — Agent Engineer IC4 | Named remote posting with application form; agent systems, retrieval, eval judgment, model decisions, cost/latency and staff-scope leadership. | Europe/North America preference and at least 20 hours each week overlapping EST; decide whether this is sustainable before applying. | Opinionated production agent story, pragmatic evals, cost profile, Go/TypeScript adaptability, mentorship. |
| Canonical — Cloud Solutions Architect, Alliances | Named worldwide home-based posting and application form; Linux, networking, Kubernetes, public/private cloud, data stack and partner workshops. | Global travel and breadth. The current application form includes an own-words agreement and warns that AI/generated application content is disqualifying. | Reference architecture, troubleshooting tree, live presentation, honest breadth. |
Changed target
Supabase careers remains a live official hub, but the syllabus cited a company role family rather than a requisition. On the check date it listed a directly relevant AI Platform Engineer role centered on durable agent execution, human review, rollback, complete logging, evaluation gates, enforced risk tiers, integrations, infrastructure ownership, and value measurement. It also listed a Postgres Engineer role demanding deep database internals and C/Rust extension work. These are different preparation lanes; “Supabase experience” is not a single target.
Unavailable supplied roles
- Automattic — Applied AI Engineer: the supplied role URL redirected to the current jobs directory, and the role was absent from the live embedded listing when checked. Keep product-first AI and async-writing topics as historical syllabus signals, but do not describe the opening as active.
- Deel — Senior Backend Engineer, AI focus: the supplied ATS URL returned a generic jobs shell without named posting metadata. The historic syllabus topics—Node.js, PostgreSQL, AI APIs, ETL, messy data, document parsing—remain useful, but current eligibility and requirements are unverified.
Detect changes that alter preparation
Not every edit matters. Rank changes by whether they affect eligibility, expected interview signal, evidence selection, or the decision to apply.
| Change | Impact | Required response |
|---|---|---|
| Location, work authorization, overlap, travel | Potential hard gate | Re-evaluate go/no-go before further preparation |
| Must-have outcome or seniority | Evidence and interview depth | Recompile requirements and re-score gaps |
| Application question or AI-use instruction | Integrity and submission process | Follow literally; write in your own words where required |
| Stack wording | Exercise likelihood and ramp story | Separate hard implementation need from adaptable breadth |
| Hiring stages | Simulation design | Add, remove, or reorder mock rounds |
| Compensation or benefits | Candidate decision, highly volatile | Use current official terms and recruiter confirmation; do not freeze them in prep notes |
| URL canonicalization only | Low, if content/application persists | Update source register; preserve requisition identity |
Diff semantically, not just textually
A reordered paragraph is noise; “remote” changing to a country list is not. Keep a normalized requirement matrix and compare atomic requirements, locations, application questions, and interview stages. Save a dated summary rather than relying on memory. Do not publish copied job-description text; record brief paraphrases and links.
Resolve conflicts conservatively
If the role page says worldwide but the application country selector excludes your location, treat eligibility as ambiguous and ask recruiting before investing heavily. If an old syllabus says eight years and a new requisition says something else, the new official requisition controls. If an aggregator claims active but the official page is a generic shell, mark unavailable.
Audit claims before rehearsing them
The final evidence audit protects both credibility and interview performance. Any claim likely to attract a follow-up must have a source and an explanation of personal scope.
The claim ledger
| Field | Question | Failure to catch |
|---|---|---|
| Claim | What exactly are you asserting? | Vague “improved performance” language |
| Context | Which system, users, time window, and constraints? | Borrowed or timeless outcome |
| Personal action | What did you decide, implement, review, or influence? | Team work presented as individual work |
| Evidence | Which report, dashboard, commit, design, test, or stakeholder record supports it? | Memory-only precision |
| Metric definition | Baseline, formula, slice, time window, and source? | Incomparable before/after values |
| Attribution | What else changed, and how certain is causality? | Claiming the whole delta |
| Disclosure | Can this be shared without customer, employer, or security harm? | Leaking confidential detail |
| Probe | What technical question should this claim trigger? | A bullet the candidate cannot explain |
Use three evidence labels
- Production: work used by real users or operations, described within disclosure limits.
- Practice: a reproducible sandbox or open project built for learning; all numbers are lab measurements.
- Hypothetical: an architecture, target, or worked example not implemented; numbers are assumptions.
Do not let polished presentation blur these categories. A strong practice artifact can demonstrate method, but it does not become production tenure. A hypothetical capacity estimate can demonstrate design reasoning, but it is not an achieved scale.
Red-team the story bank
For each of six leadership stories, ask: Who owned the decision? Who wrote the code? Who measured the result? What did you initially get wrong? What evidence would contradict the story? What can you safely disclose? Remove claims that survive only because no one probes them.
Use role-specific readiness gates
Readiness is not feeling comfortable with every topic. It is the ability to define, design, implement a small version, debug a failure, and explain a production trade-off for the capabilities the active role is likely to test.
The five-evidence ladder
- Define: explain the concept and why it matters in plain language.
- Design: place it in a complete system with constraints and alternatives.
- Build: implement or configure a small reproducible version.
- Debug: diagnose a realistic failure using evidence.
- Judge: decide when not to use it and defend the trade-off.
| Domain | Minimum evidence | Red flag |
|---|---|---|
| Retrieval/search | Versioned benchmark, slices, latency/quality trade-off, slow-query diagnosis | Only framework calls or aggregate score |
| Agents | Explicit state, permissions, checkpoint/restart, bounded failure, human boundary | Unbounded loop or prompt-only safety |
| Evaluation | Golden set, calibrated rubric, regression gate, adversarial cases | “Outputs looked good” |
| Integration/backend | Auth, idempotency, replay, reconciliation, contract tests | Happy-path webhook |
| Data/platform | Query plan, replayable ingestion, deployment probes, rollback | Tool names without operating behavior |
| Reliability/security | SLI/SLO, trace, incident drill, threat model, tenant negative test | No failure or trust boundary |
| FDE/system design | Discovery, definition of done, estimates, rollout, customer communication | Components before questions |
| Leadership/writing | Six truthful stories, design memo, ADR, outcome-led update, review example | Title-based leadership or unverifiable metrics |
| Coding/SQL | Timed mixed set, invariants, edge tests, complexity, practical API/SQL depth | Memorized solution without contract |
Green, amber, and red decisions
Green: all eligibility gates pass, core must-haves reach at least design/build/debug, and the loop has been simulated. Amber: one material gap can plausibly improve in a short, time-boxed sprint and is framed honestly. Red: eligibility fails, a core must-have is unknown, claims lack provenance, or no complete technical/customer story survives probing. Red means stop or select a different role, not hide the issue.
Full-loop simulation
- Recruiter pitch and eligibility confirmation.
- Resume deep dive with metric and ownership probes.
- Role-depth technical discussion.
- Coding, SQL, API, or debugging exercise as appropriate.
- System design or FDE discovery/architecture case.
- Cross-functional disagreement or customer objection.
- Leadership and mentoring round.
- Written follow-up or application response, following the employer’s instructions.
Score each round on clarification, correctness, depth, evidence, trade-offs, and communication. One evaluator should interrupt, change a constraint, and challenge a metric. Improvement between mocks matters more than the first score.
7-day final sprint before interview loop
When an active role interview is scheduled, use a short, high-signal sprint. This is not a full relearn. It is a focused conversion of your existing evidence into fast retrieval, safer answers, and cleaner execution under pressure.
Daily sequence
- Day 1: role and source lockRe-check the official posting and form. Freeze role-specific requirements, eligibility constraints, and likely interview rounds in one page.
- Day 2: architecture replayRun one timed system-design rehearsal with a changed constraint at minute 20. Capture the adaptation path, not just the final diagram.
- Day 3: coding and SQL pressure setDo one mixed coding block and one SQL block under timer. Grade contract clarity, edge handling, and explanation quality.
- Day 4: reliability and security drillInject one dependency fault and one AI-risk fault. Explain detection, containment, customer communication, and verified recovery.
- Day 5: leadership and writing packRehearse six stories, then write one concise design memo and one outcome-first async status update.
- Day 6: full-loop mockSimulate recruiter, technical depth, coding, design, and leadership rounds with interruption and follow-up probes.
- Day 7: repair and restFix only top failure patterns from the mock, then stop adding new scope. Enter interviews with stable narratives and clear uncertainty labels.
What to carry into each round
- One-page role map with active constraints and must-have capabilities.
- Two architecture templates you can adapt quickly under changing requirements.
- Three evidence snippets with metric definition, baseline, and personal ownership.
- One failure narrative that demonstrates reliability and security judgment together.
- One uncertainty sentence you can use when data is missing without losing authority.
Finish with integrity and a go/no-go decision
Application quality is consistency under scrutiny. Dates, titles, technologies, metrics, ownership, and links must tell the same story in the resume, form, portfolio, and interview.
Submission preflight
- Re-open the official posting and confirm status, requisition, locations, overlap, travel, and deadline.
- Read every application field before drafting; identify own-words, AI-use, confidentiality, and format instructions.
- Check resume links in a logged-out browser and remove private, broken, or misleading artifacts.
- Match every role-specific bullet to the claim ledger and disclosure boundary.
- Use plain, personal language. If assistance is prohibited, do not use generated application content. If assistance is allowed, remain the author and verify every statement.
- Prepare concise questions about outcomes, constraints, team ownership, evaluation, operational responsibility, and interview process.
Decide, do not drift
Apply now
Active and eligible; core evidence is ready; remaining gaps are honest and non-fatal.
Short sprint
Active and eligible; one high-value gap has a specific artifact or mock that can be completed promptly.
Ask first
Location, schedule, travel, or role status is ambiguous and could invalidate the application.
Stop
Posting is unavailable, eligibility fails, or core depth cannot be represented truthfully.
Interview playbook
Use a final answer structure that makes provenance and relevance easy to inspect.
Claim → Context → Action → Evidence → Limits → Role relevance
- Claim: make one bounded assertion.
- Context: users, system, stakes, and your mandate.
- Action: decisions and work you personally owned.
- Evidence: result and measurement provenance.
- Limits: attribution, uncertainty, disclosure, or gap.
- Relevance: why it transfers to the active role’s outcome.
When the source changed
Say, “I prepared from the current posting checked on [date]. I noticed [specific change], so I adjusted [artifact/story/question].” If the interviewer describes a newer scope, accept the correction, ask clarifying questions, and reason from the new constraints. Do not defend stale notes.
Common traps
- Calling a role active because the URL loads or a search cache has a description.
- Quoting obsolete eligibility or compensation from a closed listing.
- Giving precise impact without metric provenance or personal attribution.
- Presenting a practice lab as a customer deployment or an injected failure as a production incident.
- Using generated application content where the employer requires own words.
- Submitting because preparation time was invested, despite a red eligibility or evidence gate.
Question bank
These drills test source judgment and final readiness. Cite the dated official record when answering status questions.
Q1What evidence is sufficient to call an official role active?
Strong answer outline
- A company career page or company-linked ATS exposes the named role and substantive description.
- Current location/arrangement and a usable application path are visible.
- The check is dated and cross-checked against the current careers directory when practical.
Follow-up probes
- Is structured JobPosting metadata enough?
- What if the Apply button errors?
HTTP status or search snippet alone must not be treated as proof.
Q2A supplied ATS URL returns 200 and a page titled “Jobs,” but no role metadata. What status do you record?
Strong answer outline
- Record unavailable or ambiguous, not active, and state the missing named posting evidence.
- Check the employer’s current official careers directory and final redirect.
- Do not submit or quote requirements until a current requisition is found.
Follow-up probes
- How long should you keep checking?
- What if a search cache shows the old description?
Distinguish a live web application shell from a live vacancy.
Q3Why is the supplied Automattic Applied AI role marked unavailable?
Strong answer outline
- The supplied detail URL redirected to Automattic’s current jobs directory.
- The named role was absent from the directory’s live embedded listing on 2026-08-04.
- Preserve historical preparation themes, but do not state an opening or current requirements.
Follow-up probes
- What would change the status to active?
- Can a cached result override the directory?
State observed evidence and date without speculating about why it closed or moved.
Q4How should you use the unavailable Deel role in preparation?
Strong answer outline
- Treat Node.js, PostgreSQL, AI API, ETL, messy data, and document parsing as historical syllabus signals.
- Do not transfer prior tenure, location, salary, or application terms to a future role.
- Find and compile a new official requisition before tailoring or deciding eligibility.
Follow-up probes
- Which artifact remains reusable?
- What if the same title reappears with a new ID?
Durable skill reuse is acceptable; stale job claims are not.
Q5What is the most important non-technical gate in the active Sourcegraph role?
Strong answer outline
- The posting prefers Europe/North America and requires at least 20 hours per week of EST overlap.
- Evaluate both formal eligibility and sustainable working hours before deep preparation.
- Ask recruiting if location interpretation is unclear; do not promise an unworkable routine.
Follow-up probes
- Why is “remote” insufficient?
- How do you record the decision?
Use the exact current constraint and connect it to a real go/no-go choice.
Q6Why must the Supabase careers reference be recompiled into a specific role?
Strong answer outline
- A careers hub is a changing set, not a single requirement profile.
- The active AI Platform role emphasizes governed agents, while Postgres Engineer requires database internals and C/Rust extensions.
- Evidence, gaps, and likely interviews differ materially, so generic preparation misleads.
Follow-up probes
- Which shared topics remain?
- How do you choose between the two?
Name at least one hard depth difference; “both are platform roles” is not enough.
Q7Canonical’s long URL and shorter canonical URL both load. Is the role changed or active?
Strong answer outline
- It is active because the same named requisition, description, and application form persist.
- Record the canonical shorter URL as a low-impact URL change.
- Still re-check substantive constraints, including travel and application instructions.
Follow-up probes
- How do you establish requisition identity?
- Which change would force a new role map?
Separate URL canonicalization from a role-scope change.
Q8The role description and application form disagree about eligible countries. What do you do?
Strong answer outline
- Mark eligibility ambiguous and capture both official observations.
- Ask the recruiter or company hiring channel for clarification before investing or submitting.
- Do not select an inaccurate country or assume broad “remote” language overrides the form.
Follow-up probes
- Should you still prepare?
- How do you phrase the question?
Conservative verification and truthful form completion are mandatory.
Q9How do you audit a claim that latency improved by a percentage?
Strong answer outline
- Recover metric definition, percentile, workload, time window, baseline, measurement system, and exact calculation.
- Check other simultaneous changes and personal contribution.
- If provenance is incomplete, weaken or remove the precise claim and explain what is known.
Follow-up probes
- What if only an old slide remains?
- How do you discuss causality?
Precision must decrease when evidence quality decreases.
Q10How do you distinguish a sandbox failure drill from a production incident story?
Strong answer outline
- Label the sandbox as deliberate practice and describe its synthetic data/load and planned fault.
- Use it to demonstrate debugging method, controls, and evidence—not real customer impact.
- Reserve production claims for verified events you personally handled and can disclose.
Follow-up probes
- Can the same postmortem format be used?
- What value does a sandbox drill prove?
The setting, impact, and evidence category must be unmistakable.
Q11One core domain is amber after week 12. Should you apply?
Strong answer outline
- Check whether it is a hard must-have and how likely/deep the interview probe is.
- Define a short artifact or mock that can materially improve evidence; do not pretend it creates production history.
- Apply if eligibility and core bar remain credible, or select a better-matched role if the gap is fundamental.
Follow-up probes
- What makes an amber become red?
- How long should the sprint be?
The decision must depend on role criticality and evidence, not fear of imperfection or sunk cost.
Q12An application form says generated content is disqualifying. How do you proceed?
Strong answer outline
- Do not use generated content for those responses; follow the employer’s instruction literally.
- Write from personal records in your own words and verify all facts.
- If policy scope is unclear, choose the conservative interpretation or ask the employer.
Follow-up probes
- Can you use a spellchecker?
- What about preparation done earlier with tools?
Never advise evasion. The employer’s current instruction controls the submission.
Q13The interviewer describes responsibilities that differ from the posting you prepared. What do you do?
Strong answer outline
- Acknowledge the newer information and ask which outcomes, constraints, and ownership are now central.
- Adapt relevant evidence while naming any new gap honestly.
- Ask for the updated description after the conversation and reconsider mutual fit.
Follow-up probes
- Would you challenge the inconsistency?
- How do you avoid forcing a prepared story?
Demonstrate curiosity and adaptability; do not defend stale source material.
Q14What are the final conditions for an “apply now” decision?
Strong answer outline
- The named role is active, eligibility is resolved, and instructions can be followed.
- Core requirements have truthful proof, major claims pass provenance/disclosure audit, and gaps are bounded.
- A role-specific full loop has been simulated and the application packet is internally consistent.
Follow-up probes
- Which condition is non-negotiable?
- When should you stop despite strong technical fit?
Availability, eligibility, integrity, evidence, and simulation must all appear.
Proof artifact: the final readiness audit binder
Create a private, versioned binder for one active target. It should prove that the opportunity, claims, evidence, and rehearsal were independently checked. Keep confidential source material out of any public portfolio.
Steps
- Create the dated source register: official URL, final redirect, named content, application path, constraints, instructions, and a semantic diff from the prior capture.
- Build the atomic requirement matrix and mark each item demonstrated, practised, conceptual, or unknown.
- Run the claim ledger across resume, portfolio, four selected stories, and application draft. Remove or qualify unsupported precision.
- Score the nine readiness domains and run two role-specific full-loop mocks with different interviewers or question sets.
- Record the go/no-go decision, unresolved recruiter questions, and exact re-check immediately before submission.
Metrics
- 100% of eligibility fields resolved as pass, fail, or explicit question.
- 100% of consequential quantitative claims have provenance and personal-attribution notes.
- Every core must-have has at least one evidence link or a visible gap.
- Every mock round is scored on clarification, correctness, depth, evidence, trade-offs, and communication.
- The second mock corrects at least one named weakness; record the observed change without turning it into a career metric.
Deliberate failure injection
Test the binder by replacing the role page with a generic 200-response shell, changing an eligible location, breaking a portfolio link, removing the source behind one metric, relabeling a sandbox result as production, and adding a prohibited AI-use instruction. The audit should stop the apply-now decision and name every reason.
What to present
The binder itself is mainly private. Present the role-relevant artifact index, clean public proofs, and concise answers with evidence and limits. If asked about preparation, show the readiness matrix and change discipline without exposing confidential records, other applications, or private employer material.
Chapter review
The handbook ends where a real application begins: with a current source, a truthful evidence base, a clear readiness decision, and the ability to adapt when the employer’s needs change.
Glossary
- Active
- A dated status supported by a named official posting, substantive content, and application path.
- Changed
- A role or source whose URL, title, scope, location, or target requisition differs materially from the syllabus.
- Unavailable
- A supplied role that cannot be verified as a current named opening.
- Semantic diff
- A comparison of meaningful requirements and constraints rather than raw text changes.
- Claim ledger
- A private register connecting each assertion to context, personal action, evidence, provenance, attribution, and disclosure.
- Evidence ladder
- Define, design, build, debug, and judge: five progressively stronger readiness signals.
- Go/no-go gate
- An explicit decision based on availability, eligibility, evidence, integrity, and simulation.
Mastery checklist
- I can justify active, changed, unavailable, or ambiguous status with dated official evidence.
- I do not treat HTTP 200, a cached snippet, or a careers hub as a verified named opening.
- I re-evaluate location, overlap, travel, and application instructions before submission.
- Every resume metric and project claim has provenance, personal scope, and a disclosure decision.
- I label production, practice, and hypothetical evidence without ambiguity.
- I meet the active role’s core evidence gates or have made a clear stop/short-sprint decision.
- I have completed two full-loop mocks and can adapt when a constraint changes.
- I obey role-specific own-words and AI-use instructions.
Official job-description sources
- Qdrant — Forward Deployed Engineer, India: active when checked.
- Remote — Senior Forward Deployed Engineer: active when checked.
- Sourcegraph — Agent Engineer IC4: active when checked.
- Canonical — Cloud Solutions Architect, Alliances: active when checked; canonical URL differs from the supplied long slug.
- Supabase careers, AI Platform Engineer, and Postgres Engineer: changed from a generic role-family reference to specific active openings.
- Automattic — supplied Applied AI Engineer URL and current jobs directory: supplied role unavailable when checked.
- Deel — supplied Senior Backend Engineer, AI focus URL: named posting unavailable when checked.
Checked: 2026-08-04. These statuses are dated observations, not guarantees. Re-open the official role, application form, and location terms immediately before applying.