Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Case Study: Customer-Specific Fine-Tuning Platform

By Anup Rai17 min readReviewed September 2026

This is a hypothetical interview scenario. Tenant counts, workloads, latency, training times and costs are planning assumptions, not measured results. Model, framework and research references are linked to their primary sources.

Interview focus: justify learned customization, isolate each customer's data and behavior, budget shared serving, and promote a compatible evaluated release.

60-second interview answer

I would first test whether prompting and retrieval satisfy each customer's needs. Where persistent behavior still needs adaptation, I would train a tenant-specific LoRA adapter against a pinned base model. Training, evaluation, artifacts and serving routes would all carry the verified tenant identity. The gateway would select an approved release, enforce token and concurrency limits, and keep caches within the required privacy boundary. Each release would bind the base, tokenizer, adapter, data and evaluation results. I would measure cold and warm latency separately, evaluate every tenant's critical cases, and retain a compatible rollback pair while planning deletion and complete operating costs.

Remember: Justify training → Isolate data → Evaluate release → Route correctly → Monitor and retire.

Interview problem and scope

A legal-technology vendor serves 280 enterprise tenants, with the tenant count expected to double annually. Each provides 1,000–250,000 input/preferred-edit pairs and wants consistent clause formats and drafting style. Customers also need evidence that customization works on their own documents.

Separate learned behavior from current facts. Fine-tuning can help with stable structure or phrasing; retrieval supplies changing policies, contract facts and citations. An adapter is not an access-control mechanism or a reliable database of current clauses. Training data rights, client confidentiality and retention must be agreed before use.

Functional requirements

  1. Upload and validate tenant-scoped training datasets with immutable manifests and permitted-use records.
  2. Run reproducible training jobs with quotas, checkpoints and bounded retries.
  3. Evaluate candidates against tenant holdouts and platform critical-failure tests before approval.
  4. Route inference to the tenant's approved base/adapter release; callers cannot load arbitrary artifacts.
  5. Support warm/cold adapter loading, usage attribution, rollback and base-model migration.
  6. Offboard tenants across data, artifacts, running jobs, caches, evaluation records and backups under the approved schedule.

Nonfunctional requirements

  1. Isolation: prevent cross-tenant reads, writes, adapter routing and cache reuse outside the permitted boundary.
  2. Latency: clarify the proposed 1.2-second p99 target before accepting it. Use separate first-token and full-completion SLOs with defined input/output lengths.
  3. Fairness: bound each tenant's concurrent requests, queued work, input/output tokens and training GPU-hours.
  4. Quality: no promotion with a known critical isolation regression; assess tenant-specific correctness separately from average style scores.
  5. Reproducibility: pin data, base weights, tokenizer/template, training code, adapter configuration and serving settings.
  6. Operations: retain tested rollback releases, expose cold-start behavior and report deletion/retention exceptions accurately.

If the user expects an 800-token answer, an assumed 60 decoded tokens/second already takes about 13.3 seconds after the first token. A 1.2-second full answer target is infeasible under that assumption. Negotiate a first-token target, shorter task or different serving requirement explicitly; do not silently redefine the metric.

Traffic and throughput envelope

Assume four million drafting requests per 30-day month, averaging 4,000 input and 800 output tokens. That is about 1.54 requests/second averaged across the month, 16 billion input tokens and 3.2 billion generated tokens. Provision for an explicit 40-request/second burst rather than the average alone.

Capacity step Planning calculation What to verify
Peak output demand 40 requests/s × 800 tokens = 32,000 output tokens/s Real output distribution and reasoning overhead
Illustrative replica throughput 1,000 aggregate output tokens/s at the required workload/SLO Benchmark input prefill, concurrent adapters and tail latency together
Peak serving replicas 32,000 / 1,000 = 32 A replica can use one or more GPUs
With 25% headroom 40 replicas Failure-domain capacity and load redistribution

Aggregate replica throughput is different from an individual request's decode speed. The 1,000-token/second value is an assumption to test, not a model-card guarantee. Prefill, cold adapter loading and long inputs may be the bottleneck first; the fleet cost worksheet must ultimately use measured replicas, GPUs per replica, utilization and redundancy.

Why adapters, and when to stop at the baseline

LoRA means low-rank adaptation. Instead of updating every base weight, it learns a small set of added weight adjustments stored as an adapter. For a weight matrix W, a common form is W + (α/r)BA, where A and B have inner rank r. The base is frozen during this adaptation; the target matrices, rank and scaling are configuration choices. LoRA paper.

QLoRA trains adapters while keeping the frozen base in a quantized representation to reduce memory. Quantization does not remove activation, optimizer, temporary dequantization or working-state memory. Training and serving must qualify their actual precision and kernel combinations. QLoRA paper.

Approach Use when Main cost or limitation
Prompt/template plus retrieval Behavior is specified clearly and current facts dominate Prompt/context cost and possible inconsistent style
Tenant LoRA adapter Repeated behavior still fails despite a strong baseline Data curation, per-tenant evaluation and serving complexity
Full fine-tuning Adapter capacity cannot meet demonstrated needs Much larger trainable state and separate weight artifacts
Dedicated hosted or local deployment Isolation, support or workload requires it Reserved capacity and potentially higher unit cost

Start with a small tenant cohort. Compare the same unseen tasks using prompting, retrieval and candidate adapters. Stop training if the improvement does not justify extra maintenance or if source rights and evaluation quality are unresolved.

A current open-weight candidate is Qwen3.8-27B. Its language model uses a hybrid attention layout, and it also has vision support. A serving framework supporting inference for a base does not prove every proposed LoRA target module, quantization and multimodal path is supported. Pin and test that combination. Compare against current hosted alternatives where the customer's terms permit them; hosted inference access does not imply downloadable adapter support.

Separate control, training and serving planes

Architecture / visual model
flowchart TB TENANT[Tenant administrator] --> CONTROL[Authenticated control plane] CONTROL --> DATA[(Scoped immutable dataset manifests)] DATA --> TRAIN[Isolated training jobs and checkpoints] TRAIN --> CAND[(Candidate adapter registry)] CAND --> EVAL[Tenant holdout and platform critical tests] EVAL --> APPROVE{Quality passes and release approved?} APPROVE -->|No| HOLD[Hold candidate and preserve active release] APPROVE -->|Yes| MAP[(Approved tenant release mapping)] USER[Inference request] --> AUTH[Authenticate and enforce quotas] AUTH --> ROUTE[Resolve server-selected release] MAP --> ROUTE ROUTE --> LOAD[Trusted artifact verification and bounded loading] CAND --> LOAD LOAD --> POOL[Shared or dedicated serving pool] POOL --> RESULT[Response and tenant usage record]
Read diagram source
flowchart TB
    TENANT[Tenant administrator] --> CONTROL[Authenticated control plane]
    CONTROL --> DATA[(Scoped immutable dataset manifests)]
    DATA --> TRAIN[Isolated training jobs and checkpoints]
    TRAIN --> CAND[(Candidate adapter registry)]
    CAND --> EVAL[Tenant holdout and platform critical tests]
    EVAL --> APPROVE{Quality passes and release approved?}
    APPROVE -->|No| HOLD[Hold candidate and preserve active release]
    APPROVE -->|Yes| MAP[(Approved tenant release mapping)]
    USER[Inference request] --> AUTH[Authenticate and enforce quotas]
    AUTH --> ROUTE[Resolve server-selected release]
    MAP --> ROUTE
    ROUTE --> LOAD[Trusted artifact verification and bounded loading]
    CAND --> LOAD
    LOAD --> POOL[Shared or dedicated serving pool]
    POOL --> RESULT[Response and tenant usage record]

The control plane validates tenant identities, data rights, quotas and release approvals. Training workers and evaluation workers receive scoped credentials. Serving workers read only approved artifacts through a trusted loader. Keep artifact-management endpoints and arbitrary remote resolver paths away from customer inference traffic.

APIs and records

Operation Contract
POST /datasets Finalize tenant-owned object references and immutable manifest; reject cross-tenant paths
POST /training-jobs Reference approved dataset/base/configuration with budget and region
GET /training-jobs/{id} Authorized progress, checkpoint and cost status
POST /releases/{id}/promote Verify current rights, evaluation, approval and expected active-release revision
POST /inference Gateway resolves release from authenticated tenant; no arbitrary adapter URL/path
POST /tenants/{id}/retirement Stop new access/jobs and start tracked data/artifact retirement
Record Fields and purpose
Dataset manifest Tenant, content hashes, source rights, split definitions, retention and region
Training job Dataset/base/config digests, code image, seed, GPU-hour budget and checkpoint state
Adapter artifact Immutable digest, tenant, target modules/rank, tensor shapes/dtype and base dependency
Evaluation report Holdout/rubric version, per-slice metrics, critical failures and adjudication
Release manifest Base/tokenizer/template/adapter digests, quantization/engine contract, signer and approval
Active mapping Tenant, pool/region, approved release, revision and rollback release
Usage and deletion ledger Attributed resources plus completion/pending/exception status for each artifact class

A signature verifies an attestation to particular bytes under a trust policy. It does not establish that the training examples were lawful, the labels accurate or the adapter safe. Verify artifact checksum before deserialization/loading, use supported safe formats and keep registry publishing privileges separate from serving reads.

Training data and isolation

  1. Validate data ownership and allowed training use before scheduling. A tenant upload is not proof it may train on every included client document.
  2. Split by document family, client matter or time, as appropriate, before generating near-duplicate training pairs. Random row splits can leak almost identical clauses into the test set.
  3. Keep immutable training/development/holdout manifests. Do not repeatedly tune against the final holdout.
  4. Give each job only its tenant's data and artifact permissions, approved network destinations and bounded compute resources.
  5. Checkpoint model/optimizer/scheduler and data progress as required for the trainer's resume semantics, not just adapter weights.
  6. Apply the same isolation to evaluators, metrics, traces and caches; a shared evaluation service is another data-processing boundary.

A Kubernetes namespace alone is not an isolation boundary against a privileged workload. Combine identity/RBAC, storage permissions, network policy, pod security, quotas and appropriate scheduling isolation. A node selector selects nodes; it does not prevent another tenant being scheduled there. Use taints/placement controls or dedicated clusters/hosts where the required threat model demands them. Kubernetes multi-tenancy.

Tenant-provided training code or model loaders can execute code. Restrict the job image and formats, minimize credentials and isolate the runtime. NetworkPolicy does not by itself enforce object-store prefix authorization; storage IAM and application validation do that work.

Compute adapter size instead of guessing

For a target matrix with input dimension d_in and output dimension d_out, a rank-r LoRA update adds r × (d_in + d_out) parameters, before optional biases or other trainable modules.

Use an intentionally simplified model with 32 layers and two adapted 4,096 × 4,096 matrices per layer. This is a sizing exercise, not the Qwen3.8 architecture.

Rank Adapter parameters BF16 weight bytes
8 32 × 2 × 8 × 8,192 = 4,194,304 8 MiB
16 8,388,608 16 MiB
32 16,777,216 32 MiB

Real artifacts depend on layer count, asymmetric projection dimensions, attention/MLP targets, expert layout and saved modules. Training state and GPU serving allocations exceed the raw artifact bytes. There is no universal “120 MB adapter” or rank that reliably captures a fixed fraction of quality.

Sweep a small set such as 8/16/32 with comparable training budgets and independent evaluations. More rank may improve fit or overfit; adding MLP targets may matter more than doubling attention rank. Report quality, memory, training time and latency together.

GPU memory, caches and fair serving

Nominal 27B weights require approximately 54 GB at two bytes/weight, 27 GB at one byte or 13.5 GB at four bits, before quantization metadata and any components excluded from the parameter count. These are decimal byte estimates, not whole-model memory guarantees. For comparison, 70B int8 weights alone need about 70 GB, not 40 GB.

An illustrative nominal 80 GB budget might assign 54 GB to weights, 12 GB to request state, 4 GB to adapters, 6 GB to runtime and 4 GB headroom. Replace every line with measured allocations for the pinned model/engine; actual device capacity, vision components, hybrid recurrent state, KV layout and kernels can invalidate the worksheet. Serving memory is not training memory.

Layer Capacity/control to measure Common mistake
GPU adapter slots Active batch diversity, rank, runtime allocations Dividing free memory by artifact size and declaring a supported adapter count
CPU adapter cache Host RAM and load/eviction behavior Treating CPU-resident adapters as already GPU-ready
Local/remote storage Transfer, verification and deserialization latency Calling object download time the whole cold start
Request state Input length, generation length, concurrency and model architecture Assuming weight quantization equally shrinks all caches
Scheduler Per-tenant concurrency, tokens, queue age and fairness A requests/second limit alone allows long requests to monopolize capacity

vLLM distinguishes max_loras, the maximum LoRAs in one batch, from max_cpu_loras, the host-side cache bound. These are not interchangeable with a claim that 200 adapters can concurrently serve on one GPU. Pin the release and configure the supported rank/module combinations. vLLM LoRA configuration.

Reserve warm slots or dedicated pools for contracted latency classes. Use load coalescing so many misses for one adapter trigger one download, bounded preload concurrency and admission queues. Never evict an adapter still used by in-flight requests. Warmup and fairness policies need per-tenant metrics; a high global hit rate can conceal a consistently slow small tenant.

Cache isolation and the serving boundary

The gateway derives the tenant and approved release from trusted identity. Request IDs and output caches include the full base/adapter version and required access scope. Never accept a customer-supplied arbitrary filesystem path, remote URL or another tenant's adapter name.

PagedAttention improves attention-memory management; it is not a tenant-privacy proof. Prefix caches can intentionally reuse computation across requests, introducing timing and identity concerns. Inject a server-controlled unpredictable cache salt for the required privacy boundary where supported, or disable sharing/use separate pools. A predictable tenant ID is not a secret salt. Media-cache identities also require isolation; do not assume a prefix salt fixes every cache layer. vLLM security guidance.

Dynamic adapter loading belongs behind the trusted control plane. vLLM warns about runtime updates and remote resolver use in untrusted production exposure. Restrict management APIs, artifact origins, paths and credentials; ordinary inference callers must not choose what the loader downloads. vLLM LoRA loading.

Shared GPU serving may be appropriate for some tenants, while others require dedicated infrastructure or a different authorized deployment. Test wrong-adapter routing, reused prefixes/media, worker reuse, artifacts after offboarding and output/trace access. A red-team sample with zero failures does not prove absolute isolation.

Release gating and base migration

Each tenant needs a rubric covering its formats, languages and critical errors. Keep platform-level isolation and safety checks alongside that rubric. Measure factual support, semantic edits, preservation of exceptions, prohibited disclosures and task completion; a composite style score must not hide a new critical failure.

This pure function checks a validated server-created release record against a trusted serving contract. Signature verification and approval authenticity happen before the function; its result still requires an atomic update of the tenant's active mapping.

def release_gate(tenant_id, release, serving_contract):
    if release["tenant_id"] != tenant_id:
        return "WRONG_TENANT"
    if not release["rights_current"] or not release["signature_verified"]:
        return "UNTRUSTED_RELEASE"
    for key in ("base_digest", "tokenizer_digest", "engine_contract"):
        if release[key] != serving_contract[key]:
            return "INCOMPATIBLE_RELEASE"
    if release["critical_failures"] or not release["tenant_quality_passed"]:
        return "EVALUATION_HOLD"
    if not release["approved"]:
        return "APPROVAL_HOLD"
    return "ELIGIBLE_FOR_PROMOTION"

Bind all approval and evaluation records to the immutable candidate digest. A changed adapter must not inherit a passing report from an earlier artifact. Critical checks are release gates, not promises that every possible failure was tested.

A base-model refresh creates a new compatibility and behavior question even when tensor shapes still match. Retrain or explicitly validate a supported conversion, evaluate each tenant and promote the base/adapter/tokenizer/serving contract together. Do not automatically route old adapters to the newest model name.

Architecture / visual model
sequenceDiagram participant T as Tenant owner participant C as Control plane participant J as Training job participant E as Evaluation participant S as Serving pool T->>C: Approve dataset scope and acceptance rubric C->>J: Pin dataset base and training configuration J->>E: Submit immutable candidate and provenance E-->>C: Per-tenant results and critical checks alt Candidate passes and release approved C->>S: Verify and prewarm compatible release S-->>C: Ready with measured capacity C->>C: Atomically update tenant route revision else Candidate fails or approval missing C->>C: Keep current release and record hold end Note over C,S: A new base follows the same release process T->>C: Offboard or revoke training rights C->>S: Stop new access and drain affected work C->>C: Track artifact and backup retirement
Read diagram source
sequenceDiagram
    participant T as Tenant owner
    participant C as Control plane
    participant J as Training job
    participant E as Evaluation
    participant S as Serving pool
    T->>C: Approve dataset scope and acceptance rubric
    C->>J: Pin dataset base and training configuration
    J->>E: Submit immutable candidate and provenance
    E-->>C: Per-tenant results and critical checks
    alt Candidate passes and release approved
        C->>S: Verify and prewarm compatible release
        S-->>C: Ready with measured capacity
        C->>C: Atomically update tenant route revision
    else Candidate fails or approval missing
        C->>C: Keep current release and record hold
    end
    Note over C,S: A new base follows the same release process
    T->>C: Offboard or revoke training rights
    C->>S: Stop new access and drain affected work
    C->>C: Track artifact and backup retirement

Keep old and new pools while affected tenants migrate, subject to actual capacity. Rollback returns a tenant to a compatible approved release, not a cold archived adapter that may take hours to restore. Revalidate rights and retirement state before rollback; old approval cannot revive an offboarded tenant.

Training capacity and recovery economics

Assume a measured pilot job uses two GPUs for four hours: eight GPU-hours per candidate. At a hypothetical $3/GPU-hour, the compute allowance is $24, before evaluation, data preparation and retries. An eight-GPU job lasting six hours instead costs 48 GPU-hours or $144 at the same rate. The tenant's 1,000 versus 250,000 examples can change these assumptions substantially.

If all 280 tenants require one eight-GPU-hour retrain, that is 2,240 GPU-hours. A 16-GPU pool at 70% useful utilization supplies 268.8 GPU-hours/day, for an ideal minimum around 8.3 days. Three candidate runs per tenant require about 25 days at that rate, before evaluation and customer approval. Do not promise a three-week fleet refresh without reconciling trials, concurrency and shared capacity.

Checkpoint frequency balances storage/write overhead against lost work. With uniformly distributed independent failures and 30-minute checkpoints, mean lost work is approximately 15 minutes per interrupted job; restart and checkpoint reload add more. Test trainer resume semantics and cap repeated failures. A loss plateau may be convergence rather than a hung job; monitor completed steps, data progress and checkpoint age too.

Failure modes and mitigations

F1: Adapter quality regresses

Hold the candidate, inspect failing slices and keep the current approved release. A better average score cannot override a known privacy or critical semantic regression.

F2: Training reads another tenant's data

Use scoped workload identity, immutable object manifests and storage authorization. Test both list and read denial, guessed object paths and shared preprocessing caches. Namespace labels are insufficient.

F3: Adapter cache thrashes

Bound batch adapter diversity, concurrent loading and per-tenant token/concurrency demand. Coalesce identical loads, reserve contracted warm slots and monitor evictions by tenant. Do not evict active adapters to satisfy another load.

F4: Uploaded data poisons or contaminates training

Review provenance, permitted use, duplicates, sensitive data and malformed/adversarial examples. Isolate tenant effects and evaluate relevant jurisdictions and languages. A PII detector cannot certify the dataset clean.

F5: A new base breaks adapters

Reject incompatible manifests, tensor shapes and serving contracts before loading. Even compatible shapes require behavioral evaluation against the new weights and tokenizer/template.

F6: A training job runs away

Reserve GPU-hour budget, cap duration and retries, and detect lack of step/checkpoint progress. Stop on the approved limit; do not infer a stuck job from loss alone.

F7: A GPU fails during training

Resume from a durable tenant-scoped checkpoint with the required trainer state. Verify dataset position and record retried compute. Repeated failures enter a held state instead of consuming an unlimited retry budget.

F8: Signing-key rotation rejects valid releases

Distribute the new verification key before use and define an overlap/historical-verification policy. Distinguish planned retirement from a compromised key requiring revocation and artifact re-attestation. Never disable signature checks to restore service.

F9: Shared evaluation leaks results

Scope evaluator input, caches, output stores, telemetry and result-publishing credentials. Check the destination tenant as well as the training-job tenant. A private training job does not automatically make its evaluator private.

F10: Artifact versions accumulate

Retain active and tested rollback releases in a ready tier; archive or delete older ones under policy. Index full lineage efficiently rather than deleting metadata needed to locate historical artifacts. Cold archives have restore delays that belong in the rollback contract.

F11: Artifact bytes fail verification

Reject before loading, retain the prior approved route and investigate the source/transfer. A checksum detects mismatch; a trusted signature and manifest establish the approved identity. Neither validates model quality.

Operational Considerations

SLOs and tenant onboarding

Metric What to separate
First-token latency Warm requests, cold requests and queueing, by tenant/pool
Full completion Defined input/output lengths and reasoning settings
Training completion Infrastructure success, evaluation pass and business promotion are different outcomes
Quality Tenant-specific task slices and platform critical failures
Cache behavior CPU/GPU hits, coalesced loads, evictions and active adapter count
Economics Direct training/serving usage and allocated standby/support costs

An initial 4–6 week onboarding plan may allocate week 1 to data/processing agreements, week 2 to the rubric and holdout, weeks 3–4 to baseline/training, and later weeks to canary and handover. These are planning allowances. A 60–90 minute workshop begins rubric construction; it cannot finish representative evaluation by itself. Advance on evidence and approvals, not elapsed time.

For latency spikes, inspect tenant queues, input/output lengths, warm/cold mix and cache churn before adding GPUs. For a single-tenant regression, restore its approved compatible release. For an isolation failure, stop the affected path and investigate across training, evaluation and serving. Track provider/model retirement notices and test upgrades before the deadline.

Fleet cost worksheet

Use the same month and cost scope throughout. Suppose the business allocates $180,000/month to compute; the following is a hypothetical allocation, not a cloud-provider quote.

Category Fleet/month Mean per tenant/month
Serving compute $100,000 $357.14
Training/retraining compute $40,000 $142.86
Evaluation compute $15,000 $53.57
Standby and burst compute $25,000 $89.29
Compute total $180,000 $642.86
Data/artifact/backup storage $8,000 $28.57
Networking $7,000 $25.00
Security/monitoring tools $5,000 $17.86
Support/operations cost of revenue $40,000 $142.86
Included operating total $240,000 $857.14

Compute alone averages $1,928.57 per tenant per quarter. An $80–$800 quarterly total cannot describe the same workload and scope. At $720,000 monthly revenue, the remainder after compute is 75%, while the margin after the included $240,000 cost is 66.67%. Finance must classify remaining cost-of-revenue items before calling this final gross margin.

Calculate totals from unrounded fleet values. Attribute direct job/GPU/token usage to tenants, then state how shared standby and staff costs are allocated. The mean is a reconciliation check, not a suitable flat price for every tenant. A hosted per-token price is not a quote for self-hosting at the required utilization and redundancy.

Offboarding and compliance evidence

Stop new requests and training jobs, revoke mappings and drain or terminate in-flight work according to policy. Retire adapters from GPU/CPU caches and local disks as well as the registry. Track datasets, generated pairs, checkpoints, evaluation artifacts, traces and backup copies separately. Restoring a backup must replay deletion/retirement state before serving.

Do not report complete erasure while a 90-day archive or legal hold still retains data. Report pending deadlines and approved exceptions; key revocation does not remove plaintext copies already in memory or other systems. Training solely into a private adapter can simplify retirement compared with mixing tenant data into a shared base, but both data lineage and actual copies still need accounting.

Provide the scoped controls and independent evidence the customer actually requires. SOC 2 reports, ISO 27001 certification, HIPAA obligations and FedRAMP authorization are distinct; neither encryption nor a Kubernetes topology confers all of them. Avoid claiming certifications or authorizations the hypothetical platform has not obtained.

Interview follow-ups

1. Why not train the shared base on every tenant's data?

That mixes customer contributions and complicates rights, isolation, deletion and attribution. Begin with tenant-scoped adaptation and an unchanged base unless broader training is explicitly permitted and separately governed.

2. Does rank 16 guarantee a small, high-quality adapter?

No. Size depends on target matrices and saved modules; quality depends on data and task. Calculate parameters and compare ranks/targets on unseen examples with memory and latency measurements.

3. Can 280 adapters all share one GPU?

Sharing the base can reduce duplication, but active batch limits, device memory, request state, workload and isolation constraints determine capacity. CPU-cached artifacts are not all GPU-ready, and one busy tenant can dominate tokens.

4. Why does an adapter need the exact base identity?

It modifies particular weights under a particular architecture and preprocessing contract. Matching shapes alone does not preserve behavior. Promote and roll back a compatible evaluated release bundle.

5. Is a valid signature sufficient for promotion?

No. It proves an approved signer attested to bytes under the trust policy. Current data rights, tenant identity, compatibility, independent quality results and release approval are separate checks.

6. How do you make cold starts invisible?

You cannot guarantee that for arbitrary idle tenants without paying for reserved capacity. Prewarm predictable demand, coalesce loads, keep CPU/local caches and offer explicit warm-capacity classes. Report the remaining cold latency.

7. What is the hardest offboarding problem?

Finding and preventing reuse of every derived copy, including adapters, checkpoints, evaluation caches and restored backups. Track completion and exceptions honestly; removing the public route is only the first step.

Closing notes

The central release unit is a tenant-approved, compatible base-and-adapter bundle with traceable data and evaluation. Shared serving saves weight duplication but introduces scheduling, loading and privacy obligations. Close with the evidence that training beats the baseline, the actual warm/cold capacity budget, and the tested migration, rollback and retirement paths.

Related: LoRA and QLoRA, Access control, Inference fundamentals.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Case Study: Expense Operations with a Computer-Use Agent
NEXT LESSONDesign an Evaluation Gate for AI Releases →

Explore the diagram