Learnastra AI SYSTEM DESIGNAnup Rai

Interview toolkit

Research Reading for AI System Design

By Anup Rai22 min readReviewed September 2026

Research is useful in an interview when it helps you explain a mechanism, question an assumption or design a better experiment. A paper's result is evidence about its tested setting. It is not automatically a product recommendation, an industry consensus or a guarantee for a different workload.

This Learnastra reading guide covers fourteen research themes and the engineering questions they raise. Research and specification references were checked on September 24, 2026. A paper's original publication year and its later revisions are different dates; a recent revision does not make an older idea newly established.

How to use this page

  1. Start with the linked foundational lesson if the mechanism is unfamiliar.
  2. Read the paper's task definition, baseline, datasets, model configuration and evaluation protocol.
  3. Identify what the result actually measures: accuracy, success, throughput, memory, cost or another property.
  4. Read limitations, ablations and negative results before proposing adoption.
  5. Design a small comparison using your own permitted data and realistic constraints.
  6. Keep the simpler baseline if the new technique does not justify its additional cost and failure modes.
Evidence type What it can support What it cannot establish on its own
Mathematical result A conclusion under explicitly stated assumptions Behavior outside that model or those assumptions
Controlled experiment A measured difference in the tested setting Universal superiority or causality when the comparison has confounds
Benchmark report Performance under a dataset, harness and budget Production performance for every user or a fair comparison with a different harness
Provider engineering report A documented implementation and observed results Independent replication or the same outcome on another provider
Specification A contract for conforming implementations Universal adoption, identical host behavior or application security
Practitioner terminology Convenient language for a recurring implementation concern A new theorem, a formal standard or agreement across the industry

Questions worth investigating

Theme Concrete design question
Context and memory Which information must remain exact, and what may be summarized or retrieved later?
Reasoning compute and adaptation Does the task need more computation, new evidence or a changed model?
Efficient architectures Which resource actually limits this workload: memory, arithmetic, communication or queueing?
Agent reliability and security What happens after a wrong decision, a repeated attempt or a malicious tool result?
Evaluation Does the score measure the behavior the product needs, including unknown and abstained outcomes?
Tools, skills and orchestration Which responsibilities belong to the runtime, the protocol, the package and the model?
Architecture / visual model
flowchart LR P[Research claim] --> C[Check assumptions and comparison] C --> H[Form a local hypothesis] H --> B[Baseline and candidate experiment] B --> Q[Quality, latency, cost and failure evidence] Q --> D{Meets required constraints?} D -->|Yes| R[Limited rollout with rollback] D -->|No| K[Keep baseline and record findings]
Read diagram source
flowchart LR
  P[Research claim] --> C[Check assumptions and comparison]
  C --> H[Form a local hypothesis]
  H --> B[Baseline and candidate experiment]
  B --> Q[Quality, latency, cost and failure evidence]
  Q --> D{Meets required constraints?}
  D -->|Yes| R[Limited rollout with rollback]
  D -->|No| K[Keep baseline and record findings]

1. Context selection and compression

Definition: context management selects, arranges, retains and removes the information supplied to model calls. Compression replaces some information with a shorter representation. Both can lose details needed later.

The AdaCoM paper studies a separately trained context manager for a frozen agent. It reports different useful retention strategies for agents with different baseline capabilities. That is a reason to evaluate a context policy with the actual agent, rather than assume one summary strategy transfers everywhere.

Parallel Context Compaction evaluates splitting compaction work across blocks on HotpotQA and LoCoMo with several model backbones. Its wall-time comparison uses matched compaction decode volume. Parallelizing a summary is not the same as making the operation free or removing every dependency from the critical path.

Additional direction Mechanism Question to test
Demand-paged context Evict content and retrieve it again when needed, using a memory-management analogy. Can the agent identify what it needs again, and does repeated eviction cause excessive reloading? The paper's L1/L2/L3 labels are its analogy, not standard AI-memory levels.
Learned latent compression Train an encoder to represent a longer sequence using fewer latent vectors for a compatible decoder. What must change in training and serving? This is not a generic text-summary plug-in for any closed model API.

Interview application: for a long-running support case, retain exact operation IDs, permissions, unresolved questions and authoritative outcomes. Summarize discussion only where reconstruction is acceptable. Test recall of an early exception after compaction, as well as average answer quality. See context engineering and memory/state.

2. More inference compute: useful, limited and task-dependent

Definition: test-time compute is computation allocated while producing a prediction. In a frozen-model comparison, this can include longer reasoning, repeated candidates, search or verification without parameter updates.

When More Thinking Hurts reports diminishing returns and cases in which longer reasoning abandons a previously correct answer. Its practical lesson is to examine the quality–cost curve. It does not establish one stopping budget for all models and tasks.

The revised knowledge-intensive task study evaluates closed-book reasoning and finds that more computation does not consistently improve factual accuracy. Its information-theoretic argument concerns fixed-model post-processing without new information. A system that retrieves an authoritative document has changed that information boundary.

Experiment: compare supported reasoning settings at equal task coverage. Report correct answers, wrong answers, abstentions, latency and complete cost. If a missing fact causes the failure, compare adding evidence with increasing the budget. If a long proof causes the failure, additional reasoning or a verifier may be more appropriate.

Tip: do not confuse a provider's qualitative effort setting with a hard token or money limit. Adaptive modes differ by model. See model selection and cost optimization.

3. What reinforcement learning and distillation improve

Definition: post-training modifies a pretrained model using additional objectives and data. RL with verifiable rewards uses checkable outcomes; distillation trains a student using a teacher's behavior or distributions. Their data and optimization procedures differ.

The Limit of RLVR study reports improved low-k success but narrower high-k solution coverage in its experiments. A counter-study using CoT-Pass@K assesses intermediate reasoning as well as final answers and reports gains in reasoning boundaries. These are different operational definitions of capability; citing either as the final answer to “can RL teach anything new?” overstates the evidence.

A controlled study of pre-training, mid-training and RL uses synthetic reasoning tasks with controlled distributions. Its results depend on prior exposure and the difficulty of RL training examples. That controlled setting clarifies a mechanism but does not reproduce the opaque training history of every frontier model.

Method to examine Why it may help Limit to retain
On-policy distillation Teacher feedback is applied to trajectories the student itself visits, potentially helping it recover from its own errors. Teacher cost, access to distributions, task coverage and data rights matter; it is not a universally cheapest default.
ThinkPRM A generative process verifier evaluates intermediate solution steps. A verifier also needs evaluation; lower label requirements in a study do not mean zero labeling or infallible checking.

Interview application: compare prompting, supervised adaptation, distillation and RL on the target task. Include training, data preparation, teacher calls and serving in the cost comparison. Evaluate held-out task families and not only the reward used for optimization. See preference optimization, distillation and RLVR.

4. Latent reasoning and diffusion

Definition: latent reasoning performs intermediate computation in internal representations rather than requiring every intermediate step to be emitted as text. Diffusion language models iteratively refine a sequence representation; these are related research directions, not interchangeable terms.

The recurrent-depth study evaluates repeating a shared computational block at inference. Its proof-of-concept model has 3.5 billion parameters. More block iterations increase computation even when the output-token count does not increase.

PoE-Bridge combines diffusion proposals with an autoregressive target through an intermediate distribution and sampling corrections. Its speed and quality results belong to that algorithm and comparison. They do not establish that every diffusion model is faster than every autoregressive model, or that its reported benchmark accuracy is an exact distribution-preservation guarantee.

Interview application: separate output-token billing, internal computation and end-to-end latency. Ask whether your provider exposes this mechanism or whether adopting it requires different model weights and a serving implementation. See diffusion LLMs and inference fundamentals.

5. Efficient attention, compression and mixture of experts

Definition: sparse attention limits which token pairs interact; recurrent or linear-attention systems maintain alternative sequence state; mixture-of-experts models route computation among expert subnetworks. Each changes a different part of the cost model.

Native Sparse Attention combines compressed context, selected tokens and hardware-aware computation, with training designed for that attention structure. It is not an instruction to replace an arbitrary pretrained model's dense attention and assume unchanged quality.

The MoE architecture scaling study examines active parameters, total parameters and compute together. Its fitted relationships depend on its experimental design. They do not provide a hardware-independent optimal configuration for every deployment budget.

Resource Why the distinction matters
Attention arithmetic Restricting attended pairs can reduce work, but selection and indexing have their own costs.
Persistent model weights An MoE may activate a subset per token while still needing access to all experts' weights.
Sequence state KV compression, learned latent context and recurrent state preserve different information.
Hardware utilization Kernel efficiency, memory traffic and device communication can outweigh a FLOP-only estimate.

Interview application: compare the complete serving configuration on long and short requests. Include exact retrieval of rare details, not only perplexity or a long-context headline. Do not claim that all subquadratic hybrids match dense attention on every task. See attention, model taxonomy and serving infrastructure.

6. Repeated success and error propagation

Definition: agent reliability concerns correct task completion under a stated workload and operating conditions, including repeated attempts and failures. One successful demonstration provides little evidence about that distribution.

τ-bench uses repeated trials to examine consistency. pass@1 is single-attempt success, not a best-case score. pass@k asks whether any attempt succeeds; pass^k asks whether every attempt succeeds under the evaluation protocol. With independent attempts at a fixed 80% success probability, at k = 4 these illustrative probabilities are 99.84% and 40.96%. Heterogeneous tasks and correlated attempts require the benchmark's estimator, not exponentiation of an aggregate score.

From Spark to Fire studies error propagation through messages in multi-agent systems. It offers a dependency-based analysis and a tested mitigation. A graph-derived risk signal is still a detector to validate, not a universal early-warning guarantee for every orchestration topology.

Interview application: record each agent's inputs, output provenance and confirmed state changes. Inject a wrong intermediate result and observe whether another component independently catches it or merely repeats it. If parallel computer-use workers share a browser session, filesystem or credentials, concurrency adds interference and authority concerns that context separation does not solve. See multi-agent orchestration and agent evaluation.

7. Prompt injection and memory poisoning

Definition: indirect prompt injection places adversarial instructions in content an agent reads. Memory poisoning inserts malicious or false retained information that can affect later tasks. A malicious instruction can cross several sessions if the application stores it as trusted memory.

AI Agents May Always Fall for Prompt Injections analyzes attacks through contextual integrity and an ambiguity between legitimate and prohibited information flows. Its impossibility argument depends on its framing and assumptions. It does not prove that server-side authorization, fixed permitted operations or data isolation are impossible.

Hidden in Memory studies a write → later retrieval → action attack chain. Those stages have different denominators. An attack success rate conditioned on successful retrieval is not the overall probability that an arbitrary initial attempt causes harm.

Interview application: preserve provenance and writer identity, limit what may be stored, and validate scope when memory is read. Make deletion and correction propagate through summaries, indexes and caches. Treat generated training data and teacher outputs as another untrusted input surface; apparently harmless examples can still alter behavior during adaptation. Prevention, detection and containment complement each other. See prompt injection, agent security and long-term memory.

8. Control protocols, evaluation awareness and monitoring

Definition: AI control studies deployment protocols under a threat model that may include an untrusted or strategically behaving model. Evaluation awareness is the ability to recognize an evaluation setting. Recognizing a test and deliberately behaving deceptively are separate claims.

Evaluating Control Protocols for Untrusted AI Agents tests monitoring, resampling and deferral in SHADE-Arena. The paper also reports that stronger adaptive attacks substantially weaken one initially promising resampling approach. Carrying only the favorable safety number into a design would omit the result most relevant to an adaptive adversary.

Evaluation-awareness research measures whether models classify transcripts as evaluation or deployment. Use realistic held-out scenarios and examine behavior under relevant conditions; do not conclude that every evaluation is invalid because some models can identify its setting.

Evaluation direction Research evidence Design implication
Abstention incentives Why Language Models Hallucinate argues that scoring systems rewarding guesses can encourage false answers. Measure wrong answers and abstentions separately, with task-appropriate costs. Rewarding abstention everywhere would make an unhelpful system look good.
Reasoning monitoring The CoT monitorability position paper describes a potentially useful but incomplete safety signal. Combine available reasoning signals with action/outcome monitoring; do not treat a transcript as a proof.
Explanation faithfulness A hint-influence study compares acknowledgments in reasoning tokens and visible answers in open-weight models. Its findings concern the tested hints and detection method. Visible explanations need not reveal every influence.
Internal probes Gemini probe research studies distribution shifts including long context and adaptive attacks. A probe's generalization must be tested; an ordinary API customer may not have access to model activations.
Classifier cascades Constitutional Classifiers++ combines cheaper screening with more expensive checks. Include false positives, missed attacks and all cascade costs. A successful red-team exercise is not proof against every attack.

See evaluation, guardrails and governance.

9. Retrieval tools, search training and forgetting

Definition: an agentic retriever selects searches based on observations. Memory management decides what information persists and when it is updated or removed. Retrieval and memory share data concerns but serve different lifecycle roles.

Direction What the cited work studies Limitation to carry into an interview
A-RAG Exposes keyword search, semantic search and chunk-reading interfaces to an agent. Open-domain QA results do not establish tenant isolation, fresh deletions or bounded enterprise latency.
Search-R1 Trains interleaved reasoning/search trajectories with RL, including masking retrieved tokens in the training objective. The reward, retrieval environment and training distribution affect what the model learns.
BrowseComp-Plus Uses a fixed corpus and supporting documents to make deep-research experiments more controlled. Corpus control helps isolate retrieval contributions; it does not reproduce the changing live web.
SleepGate Learns retention and consolidation to reduce interference from obsolete associations. Its reported experiment uses a four-layer, 793K-parameter Transformer. Do not generalize that result into a production guarantee for frontier agents.

Interview application: distinguish retaining information from using the correct current information. Test a changed address, superseded policy, withdrawn consent and deleted source. Include both successful recall and stale-memory errors. A benchmark that rewards remembering everything can conflict with the product's correction and deletion requirements.

Use explicit version and validity metadata where possible before asking a model to infer which of two conflicting memories is current. Learned forgetting may be a research candidate; deterministic lifecycle rules are still necessary for business records. See agentic RAG, memory architectures and retrieval evaluation.

10. Multimodal models, world models and physical actions

Definitions: a multimodal model processes or generates more than one modality. A world model represents or predicts aspects of an environment and its evolution. A vision-language-action model maps visual and language information into actions. None of these terms means support for every modality or safe control of any robot.

Emu3.5 studies native vision–language next-state prediction and multimodal generation. Its stated vision–language interfaces do not justify relabeling it as a universal text/image/audio/action model. DreamX-World studies interactive generated worlds, including camera control and scene persistence. Generated visual plausibility does not establish physically accurate simulation.

The Gemini Robotics 1.5 report separates vision-language-action capabilities from embodied reasoning and discusses transfer across embodiments. Fast control loops and slower planning can have distinct timing and safety requirements. Transferring a learned policy to new hardware still needs validation.

Two other directions clarify the design space. V-JEPA 2 learns predictive representations and studies action-conditioned planning; predicting a representation differs from rendering future pixels. Thinking with Video investigates generated frames as intermediate reasoning on a defined benchmark. Neither approach establishes that a visually convincing imagined sequence is a correct physical prediction or a verified proof.

Interview application: distinguish perception, prediction, planning and actuation. Ask which state is observed versus generated. For video generation, measure temporal consistency, identity and conditioning fidelity. For physical control, model accuracy is only one part of the permitted operating envelope. See multimodal RAG, multimodal generation and computer-use agents.

11. Small models, quantization and speculative decoding

Definition: model size, numerical precision and decoding strategy are separate levers. Reducing one cost does not automatically reduce the complete application's cost at a required quality level.

The VibeThinker-3B report reports strong results on selected verifiable reasoning tasks from a small dense model. That is a candidate for task-specific evaluation, not evidence of equivalent broad knowledge, safety or production behavior across all workloads.

Reasoning-QAT studies low-bit quantization-aware training and the interaction of calibration, distillation and RL. It does not support a general rule that 4-bit is always lossless or that a particular bit width always gives a speedup. Hardware kernels and task error rates matter.

OnlineSPEC adapts speculative draft models using verification feedback. Even when feedback is already available, updating and operating the draft model is not necessarily free. Distinguish acceptance length, single-stream latency and throughput under load.

Interview application: include base weights, adapters, sequence state, runtime memory and concurrency when sizing a small or quantized deployment. For vision workloads, compare input resolution or token reduction against loss of small text and spatial detail. Low-precision training and low-precision inference have different numerical and hardware requirements. See quantization, speculative decoding and edge deployment.

12. Test-time training: three different adaptations

Definition: test-time training adapts trainable parameters or learned state using the current input or a related objective at inference time. Identify exactly what changes, how it is trained and how long that change persists.

Form What changes Primary reference Engineering question
Sequence-model TTT layers A learned hidden-state model is updated while processing the sequence. Learning to Learn at Test Time How do recurrent updates, memory I/O and batching behave on the actual hardware?
Per-task adaptation Model parameters are temporarily updated using examples associated with a task. TTT for Few-Shot Learning Which examples are permitted, how is leakage avoided and how is task state isolated?
Long-context adaptation Context is incorporated through continued learning, with initialization trained for that process. End-to-End TTT for Long Context Which information can be recalled exactly, and which is only reflected indirectly in adapted parameters?

The TTT-layer paper evaluates models from 125M to 1.3B parameters. The few-shot paper reports task-specific improvements on ARC and BBH, with different combinations of adaptation and ensembling. The long-context paper studies a sliding-window Transformer with test-time learning. These experiments concern different systems; combining their strongest numbers into one imagined product would be misleading.

Test-Time Reinforcement Learning explores adaptation using estimated rewards such as majority-vote agreement when explicit labels are unavailable. That changes the learning objective, not the need to validate outcomes: agreement-based rewards can reinforce a shared mistake. The protocol must say whether adaptation may use task-support examples or unlabeled test inputs. Keep hidden answer labels out of adaptation and method selection, and evaluate on separate tasks when claiming generalization beyond the adaptation set.

Property Additional compute with frozen parameters Test-time adaptation
Parameters Fixed during the comparison Some parameters or learned state change
Work Additional forward computation, candidates, tools or verification Adaptation work plus prediction; exact update method depends on the system
State Can still include caches, conversation, tools and a workflow Also includes adapted state with an explicit lifetime
Reproducibility Requires fixed inputs, versions, settings and relevant runtime state Also requires the adaptation data, order, objective and update configuration
Main risk Additional cost without adequate quality gain Overfitting, poisoning, state leakage, forgotten details and extra serving complexity

Key correction: frozen parameters do not make an entire agent stateless or a deterministic pure function. Conversely, TTT does not always mean an ephemeral LoRA that is discarded after one answer.

Architecture / visual model
flowchart LR I[Request and permitted adaptation data] --> S[Isolated adaptation state] B[Versioned base parameters] --> S S --> U[Bounded update procedure] U --> P[Predict and evaluate] P --> O[Return result with provenance] P --> L[Apply retention or reset policy] L --> X[Verify cleanup and tenant boundary]
Read diagram source
flowchart LR
  I[Request and permitted adaptation data] --> S[Isolated adaptation state]
  B[Versioned base parameters] --> S
  S --> U[Bounded update procedure]
  U --> P[Predict and evaluate]
  P --> O[Return result with provenance]
  P --> L[Apply retention or reset policy]
  L --> X[Verify cleanup and tenant boundary]

Interview application: use a non-adapting model as the baseline. Measure quality, exact recall, update latency, peak memory, concurrency and cleanup. If a user's input changes state that influences another user, the design has an isolation problem regardless of its benchmark score. See fine-tuning and state management.

13. Skills and package portability

Definition: Agent Skills is a format for packaging instructions and supporting resources around a SKILL.md file. Agent Plugins is a separate package format that can distribute skills and MCP configurations. A package's portability and its runtime authority are different properties.

The Agent Skills specification defines metadata and Markdown instructions, with optional supporting files. Progressive disclosure aims to load detailed material when needed. The Agent Plugins specification defines a portable package and client-specific extension boundaries. Neither specification implies identical behavior in every host.

Layer What it describes What still needs enforcement
Skill Instructions, procedures and supporting resources Whether the host loads it and whether its actions are allowed
Plugin package Distribution and discovery of supported components Dependency integrity, host compatibility and execution policy
MCP Client–server interactions for tools, resources and prompts Authentication, authorization and business rules
A2A Communication and task coordination between agent applications Identity, delegation authority and data handling across organizations

Interview application: pin and review third-party packages, identify executable scripts and server declarations, and test the procedure after model or tool changes. A static scanner can miss harmful behavior expressed through ordinary commands or prose. Evaluate duplicate skills, vague descriptions and conflicting procedures as maintainability and selection problems. Do not present a scan percentage as a universal security guarantee. See tool use and MCP and agent security.

14. Workflow graphs and managed agent runtimes

Definition: a workflow graph represents operations as nodes and possible transitions as edges. “Graph engineering” is practitioner shorthand for designing those relationships and their state. It is not evidence that the single-agent versus multi-agent decision has been settled.

An explicit graph can make possible transitions easier to inspect. It does not automatically provide replay, determinism or crash recovery. A dynamic router can still operate within a bounded, testable graph. The LangGraph overview is one concrete implementation reference; its runtime semantics must be distinguished from the general graph abstraction.

Choice Possible benefit What to compare
Fixed workflow Predictable transitions and explicit approval states Exceptional cases, maintenance and human escalation
Single agent with tools Flexible next-step selection with a smaller integration surface Tool selection, context growth, budgets and recovery
Orchestrator with specialists Separate context and bounded independent work Shared credentials/state, coordination cost, merge quality and correlated mistakes
Advisor / executor Consultation at selected difficult decisions Consult rate, quality improvement and total latency/cost versus one model
Managed runtime Provider operates some execution infrastructure Persistence, sandbox, identity, observability, export, pricing and provider failure behavior

Interview application: write down the workload before choosing a topology. If two steps modify the same resource, their contexts being separate does not make them independent. If a managed service limits delegation or roster size, treat that as a versioned service constraint, not an optimal universal architecture. See orchestration, loop design, durable execution and agent architecture options.

Worked research decision: adopt context compaction?

The following is an original, hypothetical experiment for interview practice. It is not a result from the cited papers.

Functional requirements

  1. Continue a long-running support investigation across many tool responses.
  2. Preserve exact permissions, operation IDs, confirmed outcomes and unresolved actions.
  3. Reopen a source passage when a later answer needs its detail.

Non-functional requirements

  1. Maintain the established task-success and critical-invariant criteria.
  2. Stay within the permitted context and end-to-end latency budgets.
  3. Reduce full operating cost, with no cross-tenant state reuse.

Baseline: retain recent raw tool results and retrieve earlier sources explicitly. Candidate: summarize older discussion while storing exact operation state separately.

Experiment and failures

  1. Use 500 held-out cases, with the same source snapshot and tool behavior for both candidates.
  2. Include long histories, changed policies, contradictory facts, missing sources and requests to revisit early details.
  3. Predefine the rubric, important slices, retry budget and how timeouts or unjudged results are counted.
  4. Blind reviewers to the configuration where practical. Record paired outcomes and review disagreements.
  5. Measure compaction calls, total tokens, latency and any additional retrieval/review work.

Suppose the baseline succeeds on 350 cases and the candidate on 370. Of the pairs, 40 improve and 20 regress; 330 succeed under both and 110 fail under both. The gain is 4 percentage points, from 70% to 74%, not a 4% relative improvement. The relative increase is about 5.7%. Investigate the twenty regressions and uncertainty before approving a rollout; a better average can hide the loss of an important invariant.

Flaw found Repair Cost of the repair
Summary omits an old exception. Keep a source pointer and retrieve relevant original evidence. More retrieval calls and latency.
Summary treats a proposed action as completed. Keep operation status in an authoritative structured record. State storage and reconciliation logic.
Permissions changed after the summary was made. Recheck access when source content is used. Authorization work and invalidation rules.
Compaction blocks the next response. Evaluate bounded parallel work or earlier compaction. Additional concurrency, potential contention and more complex scheduling.

Full monthly economics

Assume 100,000 cases per month, a loaded reviewer rate of $45/hour and four minutes per review. These are illustrative accounting assumptions.

Cost Baseline Candidate
Main-model calls $10,000 $6,000
Compaction calls $0 $1,000
Human review 2,000 × $3 = $6,000 2,200 × $3 = $6,600
Storage/retrieval $800 $1,000
Operations $1,200 $1,800
Implementation amortization $0 $500
Common platform/support $2,000 $2,000
Total $20,000 $18,900

The candidate saves $1,100 per month under these assumptions. Its non-review cost is $12,300, so break-even is about 2,567 reviews, or 2.57% of cases. If the true review rate becomes 3%, total cost reaches $21,300 and exceeds the baseline by $1,300. The quality comparison above does not itself prove these production cost assumptions; collect both kinds of evidence.

Closing: retain exact business state, validate the important failure cases and introduce compaction only where the quality and cost evidence support it. Keep a baseline route for rollback. Record versions so a later model or summary-policy change triggers a fresh comparison.

A staged research practice plan

Stage Reading focus Deliverable before moving on
1 Context, retrieval and memory A baseline with evidence provenance and a measured failure taxonomy
2 Inference compute and efficiency A matched quality/latency/full-cost comparison
3 Agent reliability and security Repeated-trial results and an injected-failure recovery demonstration
4 Evaluation and control A rubric, judge validation and clearly stated threat model
5 Adaptation and distillation A controlled experiment with separated training and test data
6 Architecture and orchestration A defended decision to adopt, limit or reject one additional mechanism

These stages can fit a personal 90-day study schedule, but the calendar does not establish competence. Select the themes relevant to the role and revisit foundations when a result cannot be explained clearly.

Interview questions

  1. Does a larger context window remove the need for retrieval? No. Cost, effective use, freshness and authorization still matter.
  2. Can more reasoning recover a fact absent from a closed-book setting? It may help use existing information, but it does not introduce an external fact. Compare retrieval when missing evidence is the issue.
  3. Does high pass@k show reliable single-attempt service? No. It assumes several attempts and a success-selection procedure.
  4. Does a prompt-injection impossibility argument invalidate access control? No. Read its assumptions and enforce allowed operations independently of model decisions.
  5. Does a small-model benchmark match imply broad model equivalence? No. Examine task coverage, harness, budget and untested capabilities.
  6. Does sparse attention imply low total memory? No. Count weights, sequence state, activations and runtime overhead.
  7. Why might a learned memory policy fail in a new product? Different data, task objectives, model behavior and lifecycle requirements can change its performance.
  8. Are thinking tokens proof of faithful reasoning? No. They are an observable signal whose relationship to behavior needs evaluation.
  9. Does a frozen model make an agent stateless? No. Conversation, tools, caches and workflow records still carry state.
  10. Can a task-specific adaptation leak across tenants? Yes, if the runtime reuses adapted state without a correct isolation/reset policy.
  11. Does a skill grant authority? Its instructions can request actions; the host's execution and authorization controls determine what is allowed.
  12. Does a static graph provide durable execution? Not by itself. Recovery depends on persisted state and runtime semantics.
  13. Why preserve negative results from a paper? They often expose the assumptions most likely to fail in deployment.
  14. Why report regressions alongside average gains? Important slices and invariants can worsen despite a higher overall score.
  15. When should a promising paper stay out of the first design? When a simpler design meets the requirements or the new technique lacks evidence under the relevant constraints.

Final notes and reading map

Use the glossary for definitions, the pattern reference for design choices and the benchmark lesson for score interpretation. The linked chapters throughout this page provide the implementation foundations.

Before citing a research result in an interview, be able to state what changed, compared with what, on which tasks, at what cost, and with which limitations. If you have only read an abstract, describe the high-level finding and say which implementation or experimental details still need checking. Evidence should make your design more precise, not replace its requirements.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Preparing for an AI Engineering Role
NEXT LESSONDesign an Adaptive AI Learning Tutor →

Explore the diagram