Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Model Taxonomy

By Anup Rai32 min readReviewed September 2026

A taxonomy is a systematic classification using stated characteristics. In model selection, those characteristics include the training objective, input/output contract, architecture, available weights, license and deployment. There is no single universally standardized tier system for AI models.

Example: a downloadable 27B vision-language checkpoint can be instruction-tuned, dense in its feed-forward layers, hybrid in its attention layers, and served privately with a particular quantization. These describe different properties of the same model. None establishes its accuracy on your workload.

Current examples checked September 24, 2026. Definitions are stable; provider IDs, model availability and API limits change. The snapshot below includes GPT-6 Sol/Luna, Claude Opus 5.5 and Grok 4.7. Record the exact deployment and recheck its linked contract before changing a production dependency.

In an interview, use classification to narrow a decision. Explain the mechanism, state a requirement, and identify the measurement that would justify the choice.

Table of Contents


The Taxonomy in One Table

No model belongs to only one box. Classify it independently on each axis:

Axis Useful values Why it matters
Training stage Pretrained/base, instruction-tuned, preference- or safety-tuned, reasoning-tuned, task-fine-tuned Predicts prompting behavior and whether the checkpoint is ready for user-facing work
Purpose General-purpose, code, embedding, reranking, moderation, OCR, transcription, speech, image, video, robotics/control Determines the correct endpoint and evaluation metric
Modalities Text, image, audio, video; supported document formats and action interfaces recorded separately “Multimodal” alone does not state what can be generated
Inference compute Fixed/default response, tunable reasoning effort, adaptive thinking, explicit thinking budget, best-of-N/search Changes quality, latency, token usage, and reproducibility
Tool interface No tools, client-side function calling, provider-hosted tools, computer use, code execution, search Defines the model/system boundary and security surface
Weights and license API-only, open-weight permissive, open-weight restricted, open-source system Determines hosting freedom, auditability, and obligations
Architecture Dense or MoE feed-forward layers; full, sparse or recurrent attention; encoder/decoder information flow Matters mainly for self-hosted memory, throughput, and parallelism
Deployment role Capability ceiling, balanced, high-volume economy, small/on-device Helps build an evaluation shortlist; it is not a quality guarantee
Lifecycle Experimental, preview, stable/GA, deprecated, retired Determines production and migration risk
Versioning Rolling alias, family ID, dated snapshot, provider-specific deployment ID Determines whether behavior can change without a code change
Serving path First-party API, cloud marketplace, managed dedicated endpoint, self-hosted weights, on-device Changes availability, pricing, compliance, and operational ownership
Data geography Global routing, regional processing, data-residency endpoint, on-premises A deployment property—not an intelligence category

This prevents category errors such as comparing “open source” with “reasoning,” or treating “long context” as a capability tier. A model can be an open-weight, multimodal, MoE, instruction-tuned generalist with tunable reasoning and a 256K context window—all at once.

Architecture / visual model
flowchart TB M["One candidate checkpoint + serving endpoint"] M --> A["Task: generation, embedding, ranking or classification"] M --> B["Contract: inputs, outputs, limits and tools"] M --> C["Mechanism: objective, architecture, adaptation"] M --> D["Access: weights, license and deployment"] M --> E["Operations: lifecycle, region, measured performance"] A --> F["Requirements filter, then workload evaluation"] B --> F C --> F D --> F E --> F
Read diagram source
flowchart TB
    M["One candidate checkpoint + serving endpoint"]
    M --> A["Task: generation, embedding, ranking or classification"]
    M --> B["Contract: inputs, outputs, limits and tools"]
    M --> C["Mechanism: objective, architecture, adaptation"]
    M --> D["Access: weights, license and deployment"]
    M --> E["Operations: lifecycle, region, measured performance"]
    A --> F["Requirements filter, then workload evaluation"]
    B --> F
    C --> F
    D --> F
    E --> F

This is a classification map: its branches are independent questions, not stages executed by the model.

Terms that are useful but informal

Term Safe interpretation What it does not prove
Frontier Near the current capability ceiling on some broad evaluations A shared industry threshold, universal superiority, or production readiness
Flagship The provider's leading or default high-capability offering Best quality for every task
Mini / Flash / Haiku / Luna / Small A provider-specific efficiency tier Equivalent capability, size, or latency across providers
Agentic Positioned or trained for multi-step tool-using workflows That the model is itself a complete, safe, autonomous agent
Long context A large accepted token capacity Accurate recall or reasoning across the full window
Realtime An endpoint optimized for streaming interaction A universal latency guarantee under every region and load

Use these labels to form a shortlist, then verify exact model specifications and evaluate the complete application.


Foundation, Base, Instruction, and Reasoning Models

Foundation model

A foundation model is trained on broad data at scale and can be adapted to a wide range of downstream tasks. This is the definition introduced by the Stanford foundation-model report, expressed in practical terms. It covers language, vision and other modalities. An LLM is a language-oriented model; a multimodal foundation model extends the input/output or representation scope. These categories can overlap.

Base or pretrained model

A base checkpoint primarily learns to predict or reconstruct training data. It is useful for research and further training, but it may not reliably follow instructions, refuse unsafe requests, call tools, or maintain a chat contract.

Instruction or chat model

An instruction-tuned model has additional training to respond to instructions and conversational roles. API generalists such as GPT-6 Astra/Sol/Luna, current Claude, Gemini 3.8 Flash, Grok 4.7, and DeepSeek V4.1 Flash expose instruction-following rather than raw next-token-completion behavior.

Reasoning-tuned model

A reasoning-tuned model is optimized to spend additional inference compute on multi-step tasks. In 2026, “reasoning model” is often a mode or configurable capability inside a general model family, not a separate species:

  • GPT-6 Astra supports low, medium, high, xhigh, and max; it does not support none or minimal. GPT-6 Sol and Luna also support none; their documented efforts are none, low, medium, high, xhigh, max.
  • Gemini 3.8 Flash supports low, medium, and high; minimal returns an error.
  • Claude Fable 5.1 and Opus 5.5 use always-on adaptive thinking; Sonnet 5 supports adaptive thinking; Haiku 4.5 supports manual extended thinking.
  • Grok 4.7 supports low, medium, high, and xhigh; reasoning cannot be disabled.
  • DeepSeek V4.1 Flash and the still-served V4 Pro support thinking and non-thinking modes.

Therefore “standard model versus reasoning model” is no longer a reliable top-level split. Record the model and its reasoning configuration.

Fine-tuned, distilled, and quantized variants

These labels describe how a model was adapted or served:

  • Fine-tuned: additional training changes weights for a domain, task, style, or policy.
  • Distilled: a student learns from a teacher’s outputs or internal representations. The student is often smaller, but smaller size is not required by the definition. See distillation.
  • Quantized: weights, activations or caches use reduced-precision representations. This can reduce storage and memory traffic; compute speedups depend on kernels and hardware. See quantization.
  • Adapter-tuned: a small trainable component modifies behavior while the base weights remain frozen in the usual setup. A LoRA adapter is one example; it is a form of parameter-efficient fine-tuning, not a separate alternative to all fine-tuning.

These are not guarantees of quality. Keep the base checkpoint, adaptation method, quantization, tokenizer, and serving engine in the model identity.

Do not assume every hosted model allows fine-tuning. OpenAI's self-serve fine-tuning is being wound down: access is restricted for new/inactive customers, and new jobs end for active existing customers on January 6, 2027. Existing fine-tuned inference has a separate base-model lifecycle. Check the official transition notice.


Generalists and Specialists

General-purpose generative models

Generalists handle a wide task distribution: writing, analysis, coding, structured extraction, vision understanding, and tool selection. Current examples include GPT-6 Astra/Sol/Luna, the current Claude lineup, Gemini 3.8 Flash, Grok 4.7, Mistral Large 3, Mistral Medium 3.5, Mistral Small 4, and DeepSeek V4.1 Flash.

Use a generalist when the workflow requires flexible instruction following or cross-domain reasoning. Do not use one by default for every ML operation.

Specialist model classes

Class Input → output Correct evaluation unit Current examples
Embedding Text/code/image → dense vector Retrieval recall, ranking quality, dimensions, latency OpenAI text-embedding-3-*; Mistral Embed and Codestral Embed
Reranker Query + candidates → relevance scores/order NDCG, MRR, recall after reranking Provider- or deployment-specific reranking models
Moderation / safety Content → labels and scores Policy-specific false-positive and false-negative rates OpenAI omni-moderation-latest; Mistral Moderation 2 and Shieldstral 1.0
OCR / document parsing Page/image/PDF → text, layout, or structured blocks Character/word error, table and layout fidelity Mistral OCR 4.1
Transcription Audio → text/timestamps/speakers Word error rate, diarization, latency OpenAI GPT-Transcribe; Google Gemini 3.5 Transcribe; Mistral Voxtral Mini Transcribe 2
Speech generation Text → audio Naturalness, speaker similarity, latency, safety Provider TTS or realtime voice models
Realtime speech Streaming audio/text → streaming audio/text End-to-end latency, interruption handling, turn detection OpenAI GPT-Realtime-2.1; Gemini 3.8 Live; Grok Voice API
Image generation/editing Text/image → image Prompt adherence, edit fidelity, visual quality, safety OpenAI GPT-Image-2.5 Sunburst/Flare; Google Nano Banana 2 family; Grok Imagine Image 2.0
Video generation/editing Text/image/video → video, sometimes audio Temporal consistency, control, duration, resolution Google Veo/Gemini Omni; Grok Imagine Video 1.5. The Sora 2/Videos API shutdown date is September 24, 2026.
Code completion Code prefix/suffix → code Acceptance rate, exact edit quality, latency Mistral Codestral and provider coding models
Computer use/control Screen/state + goal → actions End-to-end task completion and unsafe-action rate Built-in computer-use capabilities or specialized preview endpoints

See embeddings, reranking, voice and document parsing for the corresponding pipeline and evaluation methods.

An embedding model does not generate prose. A vision-capable generalist is not automatically an OCR system. A text model discussing audio is not a speech model. Choose the class before choosing the brand.


Modalities Are Directional

Write modality support as a mapping:

{accepted inputs}→{native outputs} \{\text{accepted inputs}\} \rightarrow \{\text{native outputs}\}

Examples as of this snapshot:

  • GPT-6 general models: text and image input → text output. Calling an image-generation tool does not make image output native to the general model. Audio uses separate models; Sora's Videos API has a September 24, 2026 shutdown date.
  • Current Claude models: text and image input → text output.
  • Gemini 3.8 Flash: text, image, video, audio, and PDF input → text output.
  • DeepSeek V4.1 Flash: text and image input → text output; V4 Pro is text-only.
  • GPT-Image-2.5 Sunburst and Flare: text/image conditioning → generated or edited image.
  • Realtime voice models: streaming audio/text → streaming audio/text, according to the endpoint contract.

Three distinctions matter:

  1. Understanding is not generation. Image input support does not imply image output.
  2. A platform is not one model. A provider may expose text, image, speech, and video through separate model families.
  3. Native is not preprocessed. A pipeline that transcribes audio and sends text to an LLM is multimodal at the system level, not at the language-model level.

PDF support also varies: one API may extract text, another may preserve page images, and another may use a separate file service. Evaluate the actual upload and tokenization path.


Reasoning and Inference-Time Compute

Reasoning behavior should be described with a configuration, not a binary badge.

Pattern Mechanism visible to developer Primary trade-off
No/minimal reasoning Lowest effort or thinking disabled Usually less per-call compute; whole-task latency and quality still need measurement
Tunable reasoning Effort/level enum such as low, medium, high, or max One model spans several quality/latency points
Adaptive reasoning Model decides when and how much internal reasoning to use Easier defaults, less deterministic cost
Explicit budget Developer sets a thinking-token budget where supported More direct cap, but API and model specific
Sampling/search over answers Multiple candidates, self-consistency, verifier, tree/search loop Higher system-level compute and orchestration cost
Tool-augmented reasoning Model interleaves reasoning with search, code, files, or other tools Better grounded action; larger security and failure surface

Reasoning tokens may be billed or included in an output category even when the full internal chain of thought is not returned. Do not design observability or trust around access to private reasoning. Record model/configuration IDs, usage, tool outcomes and externally checkable evidence. Retain or redact prompt/output content according to the application’s data policy; ordinary logs must not automatically collect every private input.

A fair evaluation holds the reasoning setting constant—or explicitly compares the cost-quality frontier across settings.

Identically named effort levels are not equal compute budgets across models. Compare measured quality, billed tokens, and latency; do not copy none or minimal into a model that rejects it.


Tools, Agents, and Model Boundaries

Function calling

The model emits a structured request to a tool supplied by the application. The application validates arguments, authorizes the call, executes it, and returns the result. The model has proposed an action; it has not executed one by itself.

Provider-hosted tools

The provider may execute search, retrieval, code, computer use, or media tools inside its platform. This changes billing, data flow, latency, and auditability. Record the exact tool version and service contract separately from the model.

Agent

An AI agent is a system that selects actions toward a goal using observations of its environment. For this guide, an LLM agent uses a model to guide some of those actions in a control loop. State, tools, budgets, permissions, recovery and stopping conditions are implementation responsibilities. No equation adding software components defines agency or guarantees reliability. See agent fundamentals.

Architecture / visual model
sequenceDiagram participant App as Application participant Model as Model endpoint participant Gate as Authorization and validation participant Tool as External tool App->>Model: Instructions + authorized context Model-->>App: Proposed tool name and arguments App->>Gate: Identity + exact proposed action alt Authorized and valid Gate-->>App: Permit App->>Tool: Execute with scoped credentials Tool-->>App: Actual result or explicit uncertainty App->>Model: Result, request state and remaining budget Model-->>App: Answer or next proposal else Rejected Gate-->>App: Deny without executing end
Read diagram source
sequenceDiagram
    participant App as Application
    participant Model as Model endpoint
    participant Gate as Authorization and validation
    participant Tool as External tool
    App->>Model: Instructions + authorized context
    Model-->>App: Proposed tool name and arguments
    App->>Gate: Identity + exact proposed action
    alt Authorized and valid
      Gate-->>App: Permit
      App->>Tool: Execute with scoped credentials
      Tool-->>App: Actual result or explicit uncertainty
      App->>Model: Result, request state and remaining budget
      Model-->>App: Answer or next proposal
    else Rejected
      Gate-->>App: Deny without executing
    end

Calling a model “agentic” normally means it is effective at planning and tool use. It does not supply least privilege, idempotency, sandboxing, approvals, memory integrity, or reliable termination. Those are system responsibilities.

GPT-6 supports asynchronous tool calling and mid-turn steering: the model can continue while an application-run tool is pending, and the user can redirect an ongoing turn. Those are interface capabilities, not permission to run unchecked actions. The application still owns authorization, cancellation, and duplicate-action prevention. See the Astra guide.

Structured output

Structured output and function calling are related but different:

  • Structured output constrains a response to a schema.
  • Function calling represents a request to invoke an operation.

A model may support one, both, or neither on a particular endpoint. Test schema validity and tool semantics instead of inferring support from a family name.


Open Weight Is Not Automatically Open Source

Use precise terms:

Term Meaning
Closed-weight / API-only The provider serves inference but does not release weights for independent hosting
Open-weight Model parameters are downloadable; the license determines allowed use, modification, and redistribution
Source-available Some code or weights are visible, but terms may restrict fields of use, redistribution, or commercial deployment
Open-source AI system The system grants freedoms to use, study, modify, and share and supplies the preferred form for modification, including the required code, parameters, and training-data information under the applicable definition

The Open Source Initiative’s Open Source AI Definition 1.0 distinguishes an open-source AI system from weights alone. In architecture discussions, say open-weight unless the full release and licenses justify the stronger term.

For every downloadable model, audit:

  • weight, code, tokenizer, and dataset-information licenses;
  • commercial, geographic, user-count, and field-of-use restrictions;
  • redistribution and derivative-model obligations;
  • acceptable-use policy and downstream notice requirements;
  • whether base and instruction checkpoints are both available;
  • availability of training code, data provenance, optimizer state, and evals.

Current Mistral examples show why the license must be attached to the exact model: Mistral Large 3 and Small 4 are listed under Apache 2.0, while Mistral Medium 3.5 uses a Modified MIT license. “Mistral model” is not a license.

A permissive license still has conditions. For example, Apache 2.0 section 4 requires specified notices and a license copy when distributing covered work. Downloadable weights, approved data handling and permission to redistribute are separate checks.

Open-weight deployment trade-offs

Open weights can provide placement control, custom serving, quantization, adapter training, and deeper inspection. They also transfer responsibility for capacity, patching, abuse prevention, safety layers, monitoring, upgrades, and incident response to the operator.

Closed APIs can provide faster access to current capability, elasticity, managed safety features, and reduced serving work. They add provider dependency and may limit weight-level customization and deployment placement.

Neither is inherently cheaper, safer, or more private. Compare the exact license, data policy, deployment, workload, and total cost.


Architecture and Size

Information flow and generation objective

Family or mechanism What it computes Suitable interview use / limitation
Encoder-only Bidirectional representations of an available input Embeddings or classification after suitable training; no autoregressive output loop by default
Decoder-only causal Next-token distributions from permitted preceding positions Interactive generation; exact attention/caching rules still vary
Encoder–decoder Encodes a source, then a decoder generates conditioned on source states Translation and other sequence-to-sequence tasks
Text diffusion Iteratively predicts/refines masked or noisy token states, sometimes in blocks Potential parallel refinement; compare quality and latency at the same workload and step budget
Recurrent/state-space or hybrid attention Carries a recurrent state, possibly alongside ordinary attention layers Different memory growth; inspect state size, retained history and kernel support

These categories can overlap: a causal decoder may contain both recurrent and full-attention layers. Read Transformer architecture and diffusion LLMs for calculations. A robotics policy or vision-language-action model also needs an explicit action representation, control frequency and safety boundary; fluent text alone does not establish a physical control contract.

Dense versus mixture of experts

  • Dense feed-forward model: each token follows the shared feed-forward layers rather than a router selecting a subset of experts. Embedding lookups still select rows; “dense” does not literally mean every stored parameter is read for every token.
  • Mixture of experts (MoE): a router activates a subset of expert parameters for a token or layer.
  • Hybrid architecture: combines mechanisms, for example full-attention and recurrent/linear-attention layers. A provider may also use “hybrid” for switchable thinking behavior. Name the mechanism instead of treating those meanings as equivalent.

For an MoE, report both total and active parameters when disclosed. Mistral Small 4, for example, is documented as 119B total parameters with 6.5B active; Mistral Large 3 is documented as 675B total with 41B active.

Parameter count is useful for estimating self-hosted memory and compute, but it is not a cross-family quality score. Training data, architecture, tokenizer, post-training, inference compute, quantization, and serving stack all matter.

For API-only models, internal architecture and parameter count may be undisclosed or may change behind an alias. Do not invent them. Treat the API's documented behavior and version contract as the interface.

The Qwen3.8-27B configuration illustrates independent axes: dense feed-forward computation and a mix of linear/full attention. For memory, 27 billion BF16 parameters alone are approximately 27e9 × 2 = 54 GB, about 50.3 GiB. A theoretical 4-bit payload is 13.5 GB before scales, metadata, caches, runtime and media processing. Neither figure proves that a complete service fits a particular GPU. See Transformer architecture.

Small, edge, and on-device

“Small” can refer to parameter count, active parameters, memory footprint, latency tier, or price tier. These are not interchangeable. An inexpensive hosted model may still be large, while a small local model can be slow on unsupported hardware.

Record:

  • checkpoint and quantization;
  • memory required for weights and KV cache;
  • prompt and decode throughput on target hardware;
  • context and batch-size limits;
  • power, cold-start, and thermal constraints;
  • quality after quantization on the real task.

Lifecycle, IDs, and Versioning

Lifecycle state Production interpretation
Experimental / research Behavior, access, and API may change; use for exploration
Preview / beta Usable for evaluation and explicitly risk-tolerant workloads; migration may be required quickly
Stable / GA Provider declares a supported production contract; still monitor deprecations
Deprecated Still callable for a transition period; replacement work should be scheduled
Retired / shut down Endpoint no longer works on that serving path

Do not equate “newest” with “GA.” Google currently lists Gemini 3.8 Flash as stable while Gemini 3.1 Pro remains preview. Preview may be more capable on a particular task but carries a different lifecycle contract.

Alias, family ID, and snapshot

  • A rolling alias can move to a newer backing model or configuration.
  • A family/model ID identifies a named tier but may still follow provider version policy.
  • A dated snapshot is intended to pin behavior more tightly.
  • A cloud deployment ID may map to a provider model version through a marketplace-specific lifecycle.

Record the resolved model when the provider exposes it, requested ID, API version, reasoning setting, tools and prompt version in production telemetry. If no immutable model revision is exposed, record that limitation instead of inventing a stronger pin.

A live alias example: the current DeepSeek catalog maps deepseek-v4-flash and deepseek-v4-flash-vision-exp to V4.1 Flash, while deepseek-v4-pro is still listed as the distinct V4-Pro-0813 text model. Do not infer a completed Pro migration from an earlier announcement. A valid request name alone does not establish the backing model.

Shutdown example: OpenAI lists September 24, 2026 as the shutdown date for Sora 2 and the Videos API, with no replacement in that notice. A price appearing on a rate card does not override the lifecycle notice. OpenAI deprecations.

For each model dependency, maintain:

  • an owner and deprecation feed;
  • a replacement candidate and migration runbook;
  • contract tests for tool calls, schemas, streaming, and usage fields;
  • held-out evals rerun before changing aliases or snapshots;
  • canary thresholds and rollback.

Context, Output, and Knowledge

Context window

The context window is the total token capacity available to the request under the provider's accounting rules. It may include instructions, conversation, images or other media converted to tokens, tool schemas, tool results, reasoning blocks, and generated output.

Always distinguish:

  • maximum input tokens;
  • total context window;
  • maximum output tokens;
  • thresholds that change price or availability;
  • effective accuracy across position and length.

A 1M-token window means the request can fit under documented conditions. It does not guarantee reliable needle retrieval, global reasoning, citations, or acceptable latency at 1M tokens.

Knowledge cutoff

A knowledge cutoff describes training knowledge, not what the model knows at request time. Current information can come from search, retrieval, databases, or tools. Tool access does not silently update the model's weights.

Separate:

parametric knowledgefromretrieved or tool-provided evidence \text{parametric knowledge} \quad\text{from}\quad \text{retrieved or tool-provided evidence}

For current or regulated facts, require authoritative retrieval and citations rather than relying on a family label or cutoff date.

Memory

Long context is not long-term memory. Cross-session memory is an application or platform feature that stores and retrieves state. It needs provenance, retention, deletion, access control, and poisoning defenses independent of the model.


Deployment and Data Control

The same or related model can be available through different serving paths:

Serving path Operator controls Operator inherits
First-party shared API Prompting, tools, application policy Provider inference stack, quotas, regions, lifecycle
Cloud marketplace / managed AI platform Cloud account, region/deployment configuration Cloud-specific model version, price, quota, and retirement schedule
Dedicated or provisioned endpoint Capacity commitment and some placement controls Provider serving software and model contract
Self-hosted weights Hardware, runtime, network, versions, safety stack Full operational and security burden
On-device Local placement and offline behavior Device constraints, update distribution, local threat model

“Sovereign,” “private,” “regional,” and “on-premises” describe deployment and governance, not model intelligence. Verify:

  • where inference and storage occur;
  • retention and training-use terms;
  • whether routing can leave the chosen geography;
  • encryption, keys, logs, and subprocessors;
  • support for private networking and customer-managed controls;
  • whether the regional endpoint changes price, capacity, or latency.

A model offered by the same vendor through two clouds can have different IDs, versions, features, prices, and retirement dates. Treat them as separate deployments in the model registry.


Current Model-Family Snapshot

This is a compact map of representative current first-party general models, not a leaderboard. Prices belong in Pricing and Costs, and production choice belongs in the Model Selection Guide.

Provider / candidate Input → output Reasoning configuration Documented limits Deployment distinction
OpenAI GPT-6 Astra Text, image → text Low through max; no none/minimal 1.05M total, 922K max input, 128K max output Tool calling requires Responses
OpenAI GPT-6 Sol Text, image → text None, low, medium, high, xhigh, max 1.05M total, 922K max input, 128K max output Chat Completions function calling only with none; use Responses with reasoning
OpenAI GPT-6 Luna Text, image → text Same supported effort enum as Sol 1.05M total, 922K max input, 128K max output Efficiency candidate; endpoint restrictions still apply
Claude Fable 5.1 Text, image → text Adaptive, always on 1M context, 128K output Evaluate when cheaper candidates miss quality constraints
Claude Opus 5.5 (claude-opus-5-5) Text, image → text Adaptive, always on 1M context, 128K output Current Opus; separate first-party and marketplace versions
Claude Sonnet 5 (claude-sonnet-5) Text, image → text Adaptive 1M context, 128K output Balanced candidate, not a universal quality ranking
Claude Haiku 4.5 Text, image → text Manual extended thinking 200K context, 64K output Efficiency candidate; own lifecycle
Gemini 3.8 Flash Text, image, video, audio, PDF → text Low, medium, high 1,048,576 input, 65,536 output Stable text-output API; Live uses separate models
Gemini 3.1 Pro Multimodal → text Model-specific thinking See linked catalog Still preview in the current catalog
Grok 4.7 Text, image → text Low, medium, high, xhigh 500K context; verify output cap Current SpaceXAI catalog flagship; search is separate
Mistral Medium 3.5 Multimodal → text Check deployed configuration 256K context Open weights; Modified MIT
Mistral Large 3 Multimodal → text Check deployed configuration 256K context Open weights; Apache 2.0
Mistral Small 4 Multimodal → text Instruct/reasoning modes 256K context Open weights; Apache 2.0
DeepSeek V4.1 Flash Text, image → text Thinking or non-thinking 1M context, 384K max output deepseek-flash; legacy Flash names alias here
DeepSeek V4 Pro Text → text Thinking or non-thinking 1M context, 384K max output deepseek-v4-pro remains separately listed

Important boundaries:

  • Google publishes an input-token limit, while OpenAI specifies separate input, output and combined limits. Check each endpoint’s accounting rules rather than adding or equating those caps.
  • OpenAI's GPT-6 general models accept images but do not natively accept or generate audio/video; use specialist endpoints for those modalities.
  • Gemini 3.8 Flash is natively multimodal on input but produces text, not native image or audio output.
  • Grok does not gain live information merely from being a current model; xAI's documentation says search tools are required for realtime events.
  • DeepSeek's old experimental vision ID is now a compatibility alias for V4.1 Flash, not a separate experimental model to shortlist.
  • Model availability and feature parity can differ on Bedrock, Google Cloud, Microsoft Foundry, or other marketplace paths.

Capability lanes, not universal tiers

For shortlisting, use operational lanes:

Lane Meaning Representative candidates
Capability ceiling Hardest reasoning, coding, planning, and long-running agent steps GPT-6 Astra, Claude Fable 5.1; also evaluate GPT-6 Sol and Claude Opus 5.5
Balanced production Strong quality with lower latency or cost GPT-6 Sol, Claude Sonnet 5, Gemini 3.8 Flash, Grok 4.7, Mistral Medium 3.5
High-volume economy Extraction, classification, routing, and simpler sub-tasks GPT-6 Luna, Gemini 3.5 Flash-Lite, Mistral Small 4, DeepSeek V4.1 Flash
Open-weight / controlled deployment Independent hosting, placement, quantization, or adapter control Mistral Large 3, Medium 3.5, Small 4, and other license-compatible evaluated checkpoints
Specialist OCR, retrieval, moderation, transcription, realtime voice, image, or video Purpose-built endpoint for the task

These lanes are hypotheses for evaluation. They do not establish a cross-provider ordering.


The Model Passport

Store a structured record for every candidate and production deployment:

model:
  provider: openai
  requested_id: gpt-6-luna
  resolved_version: gpt-6-luna
  lifecycle: ga
  checked_at: 2026-09-24

contract:
  inputs: [text, image]
  outputs: [text]
  context_tokens: 1050000
  max_input_tokens: 922000
  max_output_tokens: 128000
  reasoning: [none, low, medium, high, xhigh, max]
  endpoint: responses
  reasoning_effort: medium
  features: [streaming, function_calling, structured_output]
  enabled_tools: []  # Application choice, not every supported tool

deployment:
  serving_path: first_party_api
  region: global  # Illustrative; not a residency commitment
  weights_available: false
  license: provider_api_terms
  data_policy_version: review-link-or-contract-id

operations:
  pricing_card_date: 2026-09-24
  rate_limit_tier: production-account-specific
  prompt_version: support-agent-v12
  eval_suite: support-agent-heldout-v7
  rollback_target: previous-pinned-deployment

Extend it with:

  • tokenizer and media-token rules;
  • cache, batch, flex, or fast-service support;
  • exact regional deployment ID;
  • safety and access restrictions;
  • measured P50/P95 latency and throughput;
  • quality, reliability, and safety results;
  • cost per successful task;
  • deprecation owner and migration deadline.

Do not fill unknown fields from inference or naming conventions. Write unknown, link to the provider contract, and test what can be tested.


Common Taxonomy Mistakes

  1. Treating provider suffixes as universal sizes. Flash, Haiku, Luna, Small, Mini, and Pro are not standardized across vendors.
  2. Calling all downloadable weights open source. Inspect code, data information, weights, and every applicable license.
  3. Equating context capacity with long-context quality. Test recall, reasoning, citations, latency, and price at the real length distribution.
  4. Equating image input with image generation. Record input and output modalities separately.
  5. Calling a model an agent. The model proposes; the harness owns tools, state, permissions, budgets, recovery, and stopping.
  6. Treating reasoning as a permanent model class. Many current models expose effort levels or adaptive thinking inside the same model ID.
  7. Mixing model tier with service tier. Faster/priority serving can use the same model at a different latency and price.
  8. Assuming a knowledge cutoff means current facts. Search and retrieval are separate evidence paths.
  9. Assuming the same model name means cross-cloud parity. Check exact IDs, regions, features, versions, pricing, and retirement dates.
  10. Publishing a giant current-model catalog without a date. Keep the taxonomy stable and the current snapshot small, dated, and sourced.

A small semantic router you can explain

Suppose incoming tasks are shipping FAQs, policy interpretation, and code analysis. Build a labeled sample containing the input, each candidate model's verified outcome, latency, and total cost. Embed the task input, inspect clusters to discover useful categories, and train or calibrate a router against those measured candidate outcomes. A nearest-centroid classifier is a baseline: compare the query vector to FAQ, policy, and code centroids, then choose a permitted model only when confidence and the quality gate allow it.

For example, “Where is order 482?” may route to an authenticated order lookup and response template, while “Do these two exceptions conflict?” routes to the policy model. An uncertain or out-of-distribution input takes an evaluated fallback. Clusters alone do not tell you which model is competent, and a model calling its own answer “easy” is not a calibrated routing label. Split by customer/source/time where needed to avoid near-duplicate leakage. Recheck routing after model or traffic changes.

Open-weight examples for interview breadth

These include current and historical examples whose model cards can be studied, not a ranking of the newest releases or a purchase recommendation. Pin the exact checkpoint and inspect its license, template, modality, and deployment requirements.

Example What it illustrates Interview implication
Qwen3.8-27B A dense feed-forward, hybrid-attention vision-language model with thinking controls Evaluate the selected mode and serving parser; the family name alone does not identify the operating configuration
Llama 3.3 70B Instruct A larger instruction-tuned text model with its own community license Open weights do not remove license obligations or the need to size memory and operations
Mistral Small 3.1 24B Instruct Image-and-text input capability in a deployable checkpoint Test image preprocessing, template compatibility, and actual visual tasks rather than assuming every text server supports it

Recall the dimensions of the comparison: weights, license, modalities, architecture, adaptation, serving, evidence. A memorized vendor list is less useful than classifying a newly encountered checkpoint correctly.

Interview case: choose model classes for a support platform

Prompt: design model selection and routing for an account-support service. These are interview assumptions, not vendor benchmarks.

Functional requirements

  1. Answer order-status requests using authenticated order data.
  2. Answer policy questions with cited, authorized policy versions.
  3. Extract fields from uploaded return-label images and request confirmation when uncertain.
  4. Escalate unsupported requests without fabricating an answer or performing account writes.

Non-functional requirements

  1. Meet a measured p95 end-to-end latency of two seconds for order lookup and eight seconds for generated policy answers at the proposed peak load.
  2. Achieve at least 95% accepted outcomes on the held-out workload and investigate every critical authorization failure separately.
  3. Keep restricted tenants on approved regional deployments, including fallbacks, OCR and telemetry.
  4. Retain the exact routing, prompt, evidence and model versions needed to investigate an outcome, within the data-retention policy.
  5. Bound retries, execution time and spending; define an explicit unavailable outcome when no eligible dependency remains.

Basic design: send all inputs to one general-purpose model. It simplifies the first prototype but does not supply current order state, document authorization, image parsing accuracy or predictable costs. A large context window does not fix these gaps.

Detailed design: filter deployments by hard requirements first. A calibrated task router then selects among eligible paths. It cannot grant a disallowed model access to sensitive data.

Architecture / visual model
flowchart TB C["Authenticated request: tenant, task, attachments"] --> G["Policy gate: data class, region, modality, lifecycle"] V[("Versioned registry: approved model and tool deployments")] --> G G --> R{"Eligible task route"} R -->|Order status| D["Scoped order lookup + deterministic template"] R -->|Policy question| E["Authorized retrieval + evaluated text generator"] R -->|Return-label image| I["Evaluated OCR/vision extractor + field checks"] R -->|Unknown or unavailable| H["Explicit escalation"] E --> O["Evidence and output validation"] I --> O D --> O O --> A["Answer or request for confirmation"] O --> T["Versioned outcome, latency, usage and failure record"] T -.-> Review["Offline review and approval"] Review -.-> V
Read diagram source
flowchart TB
    C["Authenticated request: tenant, task, attachments"] --> G["Policy gate: data class, region, modality, lifecycle"]
    V[("Versioned registry: approved model and tool deployments")] --> G
    G --> R{"Eligible task route"}
    R -->|Order status| D["Scoped order lookup + deterministic template"]
    R -->|Policy question| E["Authorized retrieval + evaluated text generator"]
    R -->|Return-label image| I["Evaluated OCR/vision extractor + field checks"]
    R -->|Unknown or unavailable| H["Explicit escalation"]
    E --> O["Evidence and output validation"]
    I --> O
    D --> O
    O --> A["Answer or request for confirmation"]
    O --> T["Versioned outcome, latency, usage and failure record"]
    T -.-> Review["Offline review and approval"]
    Review -.-> V

The dotted path represents an offline review that may approve a new registry version. Production telemetry must not silently change routing policy or deploy an unreviewed model. An image parser's output is evidence to validate, not an instruction to tools. See document processing and RAG.

State and request contract: a request carries an immutable request ID, authenticated tenant, task input, attachment references and deadline. Persist the selected deployment and policy revision, then record running, completed, escalated or failed. A retry checks existing state and approved routes. Fallback is another explicitly allowed deployment; it cannot silently send data to a different region. Tool writes are outside this exercise, so retrying lookup is read-only, while billed model executions still need separate accounting.

Decision Benefit Cost / flaw to test
Deterministic order lookup Fresh authoritative state and no generative arithmetic Backend dependency; authorization and stale-cache risks remain
Separate OCR/vision path Evaluate field accuracy and image handling directly More components; extraction errors can propagate
Smaller generator for routine policy answers Potential savings and lower latency Task misrouting; test rare policy exceptions
Higher-capability fallback May recover hard cases Additional latency/cost; only permitted deployments qualify
Self-hosted restricted-data route Control over placement and version GPU capacity, operations, patching and license obligations
Pin registry and rollout versions Traceable behavior and reversible changes More release work; aliases can still change upstream

Cost/benefit calculation: suppose 100,000 monthly requests comprise 50% order lookups, 30% policy answers and 20% images. Assume a one-model baseline costs USD 0.020 per request, including model attempts, and achieves 96% accepted outcomes. A proposed router costs USD 0.001 per request; its paths cost USD 0.001, 0.012 and 0.008 per request respectively, including their measured retries. The routed variable cost is:

100000 × (0.001 + 0.50×0.001 + 0.30×0.012 + 0.20×0.008) = USD 670/month.

The baseline costs USD 2,000/month. If routing needs USD 500/month in additional operations and amortized implementation, its full incremental comparison is USD 1,170 versus USD 2,000, saving USD 830. At 95.5% accepted outcomes, that is about USD 12.25 per 1,000 accepted results, compared with USD 20.83 for the baseline. Common infrastructure costs are excluded equally; a full budget must add them to both options. The routed quality is 0.5 percentage points lower, so confirm the required threshold and important slices before accepting the saving. Aggregate cost improvement cannot excuse unauthorized disclosure.

Failures and recovery: a router outage uses a preapproved fixed route or escalation; provider failure uses only an eligible fallback; a registry outage may retain the last approved version only within its explicit freshness/expiry policy; an expired or revoked policy means no inference. An alias change triggers contract/evaluation checks and canary review. A model's self-reported confidence is not the fallback threshold. Calibrate routing decisions on measured outcomes and reevaluate when traffic changes.

Closing remarks: “I would first identify the operation each request needs. Current account state comes from an authorized lookup; language synthesis and image extraction use independently evaluated models. I would choose the least costly eligible path that meets quality and latency requirements, record every version and keep fallback within the same privacy rules. The next evidence I need is peak-load performance, rare-case routing errors and complete cost per accepted outcome.”

Interview Questions

Q: How would you classify an unfamiliar model?

Explain first, then compare your answer

“I classify it on independent axes: training stage, purpose, input and output modalities, reasoning controls, tools, weights and license, architecture, lifecycle, version contract, context/output limits, serving path, and data geography. Then I verify the model ID and deployment instead of inferring capability from names like Pro or Flash. Finally, I attach workload-specific quality, latency, reliability, safety, and cost measurements.”

Q: What is the difference between a reasoning model and a normal LLM?

Explain first, then compare your answer

“The distinction is now a spectrum. Reasoning-tuned models spend additional inference compute on multi-step work, but many 2026 general model families expose reasoning as an effort level, thinking mode, or adaptive behavior within the same model. I record the exact configuration because it changes quality, latency, and billed tokens. I evaluate each setting rather than assuming the highest effort is always best.”

Q: What is the difference between open-weight and open-source AI?

Explain first, then compare your answer

“Open-weight means the learned parameters are available under some license. It does not by itself provide training code, data information, unrestricted use, or redistribution rights. Open-source AI is a stronger claim about the system and the preferred form for modification. For deployment I audit the exact weight, code, tokenizer, and data-information licenses and avoid using the provider family name as a license.”

Q: Is an agentic model an agent?

Explain first, then compare your answer

“No. An agentic model may be good at planning and tool selection. An agent is the surrounding system: model, loop, tools, state, permissions, guardrails, budgets, error recovery, and stopping rules. Security and reliability live at that system boundary, not in the adjective attached to the model.”

Q: Does a one-million-token context remove the need for RAG?

Explain first, then compare your answer

“No. Context size is a capacity limit, not a guarantee of retrieval accuracy, freshness, authorization, citation quality, or acceptable cost and latency. RAG can provide current evidence, source attribution, access filtering, and a smaller working set. I compare long-context and retrieval designs on the actual document distribution.”

Q: How should lifecycle affect model selection?

Explain first, then compare your answer

“I separate capability from release status. Preview models can be evaluated but need explicit acceptance of change and shutdown risk. Production dependencies need an owner, versioning policy, deprecation monitoring, contract tests, held-out evals, a canary, and a rollback target. A stable lower-tier model may be the better production choice than a stronger preview.”

Q: A model has 100B total and 5B active parameters. Can it fit wherever a 5B dense model fits?

Explain first, then compare your answer

No. The active count describes computation selected per token, not necessarily stored weights. The full expert set, shared weights, precision, cache and runtime state determine storage. Expert offloading or parallelism changes placement and latency. For example, a hypothetical 100B BF16 parameter payload is 200 GB before runtime allocations.

Interview tip: Ask whether the number is total parameters, active parameters or measured device memory.

Q: Can a model be both dense and hybrid?

Explain first, then compare your answer

Yes. A dense feed-forward model can mix full-attention and recurrent/linear-attention layers. Those are different architectural axes. Qwen3.8-27B is a current example. A provider may also use hybrid to describe thinking modes, which is a behavioral setting rather than that layer architecture.

Interview tip: Name the layer types and state representation when sizing memory.

Q: A vision model accepts invoices. Does that establish production OCR accuracy?

Explain first, then compare your answer

No. Image acceptance establishes an input contract. Evaluate character/field accuracy, reading order, tables, low-resolution scans and missing-field handling on representative invoices. The correct architecture may include a dedicated parser and deterministic field checks.

Interview tip: Measure downstream task errors as well as OCR output.

Q: An Apache-licensed checkpoint is downloadable. Can we omit notices when redistributing a modified version?

Explain first, then compare your answer

No. Permissive licensing still has conditions. Apache 2.0 includes license-copy, modification-notice and applicable attribution/NOTICE requirements for distribution. Inspect the complete release and exact license; API terms and training-data permissions are separate.

Interview tip: Do not label a whole provider’s catalog with one license.

Q: The API accepts the same ID after an upgrade. What might have changed?

Explain first, then compare your answer

The resolved weights, tokenizer, prompt interpretation, supported features, defaults, safety behavior or serving configuration may have changed. Record the provider version information available, contract-test schemas and tools, rerun held-out tasks, canary and retain a supported rollback path. An opaque alias cannot create a stronger pin than the provider offers.

Interview tip: A successful HTTP response is not a compatibility evaluation.

Q: Does disabling thinking always make a workflow cheaper?

Explain first, then compare your answer

It may reduce one call’s compute but lower accuracy enough to cause additional calls, retries or human review. Compare complete task cost and elapsed time under the same acceptance criteria. Some models reject disabled thinking entirely.

Interview tip: Reasoning effort names are not standardized compute budgets.

Q: What belongs in a data-residency decision beyond the model name?

Explain first, then compare your answer

The serving region, storage, logs, caches, tool services, support/subprocessors, fallback routes and retention/training-use terms. Validate the exact account and endpoint contract. A regional primary with a global fallback can violate the intended restriction.

Interview tip: Attach policy eligibility to each deployment, including specialists.

Q: Two embedding models return 1,024 numbers. Can we mix their vectors?

Explain first, then compare your answer

Equal dimension does not establish a shared learned space. Query/document encoder compatibility, model version, normalization and similarity metric must match the index contract. A migration normally needs re-embedding or parallel versioned indexes.

Interview tip: See the embedding migration lesson.

Q: Does a semantic cluster prove which model should receive the query?

Explain first, then compare your answer

No. Similar input language can contain easy and hard tasks, private data or a high-impact action. Train or calibrate routing against observed candidate outcomes and cost, enforce hard policy first and handle out-of-distribution inputs. Test the selected traffic slice and fallbacks together.

Interview tip: A routing model is another evaluated dependency, not a source of authorization.

Final summary and notes

Recall card One-sentence answer Common trap
Classify State purpose, contract, mechanism, access and operations independently. Treating every model as one brand tier
Measure Attach task quality, tail latency, reliability and complete cost. Turning a vendor benchmark into a production guarantee
Pin Record model, endpoint, tokenizer/template, tools and policy versions. Assuming a stable-looking alias is immutable
Authorize Enforce data and action policy outside model selection. Letting a fallback bypass residency or permissions
Review Reevaluate changed models and traffic before expanding deployment. Assuming yesterday’s best route remains optimal

60-second answer: “I start by identifying the operation: lookup, generation, ranking, extraction or action. I classify each candidate by its inputs and outputs, training/adaptation, architecture, available weights and deployment contract. I filter out candidates that fail privacy, modality or lifecycle requirements. Then I compare measured quality, latency and full cost on representative tasks. Model names form a shortlist; the evidence decides the deployment. I document the selected version, fallback rules and migration owner.”

Keep the capability-assessment, pricing and selection records together. A model catalog answers what exists; an assessment answers what works for this application.


Official Sources

Checked September 24, 2026:


Next: Capability Assessment

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Inference Pipeline
NEXT LESSONCapability Assessment →

Explore the diagram