Multimodal generation produces content in one or more media types—such as images, video or audio—conditioned on text or other media. A text-to-image model is one example. A system that generates a video with dialogue from a script and reference pictures is another. Multimodal understanding interprets existing media; generation creates or edits media. A product can do both.
In an interview, the central question is: How does a requested asset become a usable, authorized and affordable result, even when generation is slow or fails? A successful model call is only one step.
This chapter develops the concepts first, then designs a 30-second video service from requirements through recovery, evaluation and cost. Workload and cost figures in the design are explicit interview assumptions. Product availability was checked on September 24, 2026.
Learn the vocabulary first
| Term | Standard meaning | Concrete example |
|---|---|---|
| Modality | A type of information or representation | Text, image, audio and video |
| Conditioning | Input information that guides a generation | A product photograph and a description of the camera movement |
| Text-to-image / image-to-video | Generation whose name describes its input and output | Animate a supplied reference image into a short clip |
| Joint generation | A model produces related modalities together | Dialogue and video generated as one coordinated output |
| Cascade | Separate stages produce or transform successive outputs | Generate narration, generate video, then align and compose them |
| Latent representation | A learned internal encoding, often more compact than raw media | An image encoded before iterative generation in latent space |
| Seed | An initial value for a pseudorandom process | Hold a randomness setting constant while comparing prompts |
| Inference steps | Iterations used by an iterative generation procedure | A draft uses fewer denoising or integration steps, if supported |
| Keyframe | A specified frame used to guide part of a video | Require the shot to begin on the product photograph |
| Rendition | A particular encoded or resized version of an asset | A vertical MP4 and a square MP4 derived from one approved edit |
| Muxing | Combining encoded media streams in a container | Put video and audio streams into an MP4 without necessarily re-encoding them |
| Provenance | Information about an asset's origin and changes | Which source image, model and editing operation produced a rendition |
Memory card: conditioning defines the requested result; generation proposes pixels or samples; validation determines whether the result can be used.
How generation works at a useful interview depth
- Encode the inputs. Text, images or audio become representations the model can process. Inputs still need file, size, authorization and content checks.
- Generate a representation. A diffusion model learns to reverse a noising process. A flow-based model learns a vector field that transports a simple distribution toward a data distribution. An autoregressive model predicts the next element conditional on preceding elements. These approaches can appear in hybrid systems.
- Decode and post-process. Convert the result to pixels, waveform samples or another usable representation; resize, encode or compose when required.
- Check the actual output. A valid prompt does not guarantee legible text, correct product geometry, safe imagery, an accurate voice or usable timing.
“Transformer” and “diffusion” do not describe mutually exclusive model categories: a transformer can be the neural network used inside a diffusion or flow-based generator. Similarly, native audio-video support does not remove the need for editing, evaluation or rights checks. The original latent diffusion paper and flow matching paper explain the training distinctions.
How to control an image or shot
| Technique | What it controls | What it does not guarantee |
|---|---|---|
| Text and, where supported, negative prompts | Desired content and discouraged attributes | Exact compliance or supported negative-prompt semantics on every API |
| ControlNet-style conditioning | Structural signals such as pose, edges or depth | Identity, ownership or compatibility with every base model |
| Reference-image adapter | Visual guidance from example images | Exact preservation of a face, brand mark or product |
| Inpainting | Regenerates a selected region using surrounding context | Perfectly unchanged pixels outside the region on every implementation |
| Outpainting | Extends the image beyond its original boundaries | Physically correct continuation of a scene |
| Regional prompting | Applies different instructions to areas of a composition | Sharp isolation when the implementation blends conditions |
| LoRA | Learns a low-rank update for adaptation | A universally best personalization method or permission to train on a subject |
| Keyframes and continuation | Guide shot boundaries or extend a clip | Exact frame continuity, consistent physics or available support on all models |
Start with reference conditioning when it meets the requirement. Add adaptation only after an evaluation shows a persistent gap and the training rights are established. Keep adapter versions, base-model compatibility and permitted uses with the release. See LoRA, QLoRA and PEFT. ControlNet and reference adapters are described in their original ControlNet and IP-Adapter papers.
Production Pipeline Patterns
Start with the smallest workable design
A small image editor can send an authenticated request to one provider and return the resulting image if it fits the application's latency and timeout budget. Streaming previews may improve that interaction. Synchronous generation is not inherently invalid.
For longer work, separate acceptance from completion:
Read diagram source
flowchart LR
U[Authenticated client] --> A[API: validate scope and reserve budget]
A --> D[(Job database)]
D --> Q[Durable work queue]
Q --> W[Generation worker]
W --> P[Provider or owned GPU service]
P --> W
W --> S[(Private output storage)]
W --> D
U --> R[Read status and authorized result]
R --> D
R --> S
The queue may be fed through a transactional outbox: save the job and an enqueue record in one database transaction, then let a dispatcher publish the work. This avoids losing accepted jobs between a database commit and queue submission. A queue notification is permission to inspect a job, not permission to charge or execute it again.
Tip: state the distinction between an HTTP request timeout, a job deadline and a provider cancellation. They are different events.
Separate the identities
| Identity | Lifetime | Purpose |
|---|---|---|
| Project and revision | User's editable creative work | Changing a script creates a new revision |
| Logical generation ID | One requested generation or deliberate new variant | An ordinary client retry retrieves the same job |
| Stage attempt ID | One attempt to produce an intermediate output | Investigate failed image, narration or encoding work |
| Provider job ID | Provider's record of an accepted operation | Reconcile completion, cancellation and billed usage |
| Asset ID and content hash | One immutable stored output | Reuse exact bytes and verify approval matches those bytes |
| Publication ID | One approved delivery package | Track the released files and their disclosures |
Scope a client idempotency key to the authenticated tenant and operation. Save a canonical request hash with it. Reusing the key with different inputs returns a conflict; pressing “generate another variant” creates a new logical request. Do not treat identical prompts as a universal request to reuse the same image.
Model each stage as a recoverable operation
Read diagram source
stateDiagram-v2
[*] --> Queued
Queued --> Submitting: lease and budget reservation
Submitting --> Running: provider ID recorded
Submitting --> OutcomeUnknown: response lost
OutcomeUnknown --> Running: accepted job found
OutcomeUnknown --> Queued: nonacceptance established
OutcomeUnknown --> Held: cannot reconcile safely
Running --> Validating: output copied and checked
Running --> Failed: confirmed terminal failure
Running --> CancelRequested: user cancels
CancelRequested --> Cancelled: cancellation confirmed
CancelRequested --> Quarantined: completion wins the race
Validating --> Ready: required checks pass
Validating --> Quarantined: invalid or prohibited output
Ready --> [*]
Failed --> [*]
Cancelled --> [*]
Quarantined --> [*]
Held --> [*]
This is a logical operation state machine, not a provider-specific API schema. Persist transitions with a version check so a stale worker cannot overwrite a newer state. Use renewable leases and stop dispatching when a worker loses its lease.
A provider may accept a costly job before the connection fails. Your database's idempotency key alone cannot prevent a duplicate provider charge. Use provider-supported deduplication within its documented scope and retention window, or query the accepted operation. If neither is possible, hold the ambiguous operation for reconciliation instead of blindly submitting again. See durable execution.
Verify webhook signatures where supported, reject replays that would change settled state, and deduplicate events. Treat callbacks as notifications; reconcile authoritative provider status where necessary. A periodic poller handles missed callbacks and jobs stuck past their deadline. An event arriving twice must not settle the same usage record twice.
Keep a private production manifest
Node editors such as ComfyUI express generation as connected operations and can save the workflow as JSON. Version that graph alongside model and custom-node dependencies. ComfyUI can also embed workflow data in generated files; inspect exported metadata before public delivery so private prompts and paths do not escape with an image. A visual graph still needs application-level permissions, budgets and durable recovery.
Store the information needed to explain a generation, subject to access and retention policy:
- Authenticated owner, project revision, purpose and approved usage scope.
- Input asset IDs and hashes, consent/rights references and prompt-template version.
- Encrypted prompt content where retention is permitted; never authentication secrets.
- Provider, requested and returned model versions, seed if supported, dimensions, duration, sampling settings and adapters.
- Workflow version, attempt IDs, provider IDs, timestamps and settled costs.
- Output hashes, evaluations, approval identity and final publication references.
A seed controls a source of randomness; it is not a complete reproducibility contract. Numerical kernels, hardware, precision, library versions and batching can affect results. Even local execution with fixed inputs needs explicit determinism controls, which can cost performance. Hosted reproducibility depends on the provider's documented contract. PyTorch reproducibility guidance describes these limits.
Preserving an approved asset is simpler than regenerating it exactly: retain its immutable bytes and manifest. Reproducibility is useful for debugging, but should not be the only way a customer can recover an approved file.
Control cost without breaking the product
| Decision | Benefit | Cost or failure to manage |
|---|---|---|
| Reuse an authorized immutable asset | Avoid regeneration and preserve the accepted result | Recheck access, retention and permitted use; never share tenant-private cache entries |
| Generate cheap drafts first | Reduce expensive renders for discarded ideas | Draft and final quality can differ; validate the final rendition |
| Regenerate only an invalidated stage | Preserve successful work | Track dependencies; a changed narration can invalidate timing and lip-sync |
| Use a faster model or fewer steps | Lower latency or cost where supported | Measure prompt adherence and downstream rejection, not only call price |
| Batch compatible work | Improve throughput or use provider discounts | Longer waiting time; respect deadlines and tenant fairness |
| Keep workers warm | Avoid repeated model-loading delays | Pay idle capacity; reserve headroom for peaks |
| Bound retries and reserve spend | Prevent runaway generation | Some legitimate work waits or fails when its allowance is exhausted |
| Separate generation and encoding pools | Scale the actual constrained stage | More queues and operational complexity |
Cache lookup uses tenant scope, immutable input hashes, model/workflow settings and the product's reuse policy. A cache hit is usable only while access and rights remain valid. Refresh an expired delivery URL for a stored asset; do not rerun the generator merely to get another URL.
For owned workers, size from arrival rate, measured service time, memory constraints and target utilization. Watch queue age, not just depth. For a hosted API, adding local workers does not increase provider concurrency or rate quotas. Apply backpressure before filling an unbounded queue.
Draft arithmetic: assume 100 requests, four candidate clips each, $0.80 per full render and $0.04 per draft. Rendering every candidate at full quality costs $320. Drafting all 400 and fully rendering 40 selected candidates costs $16 + $32 = $48, an 85% reduction under these assumptions. Selection rates, provider enhancement charges and rework can change that result. These are illustrative rates, not a quote for a named model.
Provenance and Safety
Separate origin, truth and permission
| Question | Evidence to collect | Insufficient evidence |
|---|---|---|
| Where did this file come from? | Validated creation/edit history and asset bindings | A filename or a claim in the prompt |
| Has the bound content changed? | Validate the signature and applicable content binding | Merely displaying a credentials icon |
| Does it depict a true event? | Independent factual evidence | A valid provenance signature |
| May we use the subject or source material? | Applicable rights, consent, license and usage scope | Possession of an uploaded photograph |
| May we publish this rendition? | Current policy checks and approval for the actual output | Approval of an earlier draft |
C2PA / Content Credentials is a standard for signed provenance records associated with digital assets. A valid record supports claims about the signing source and bound content; it does not establish that every assertion is true or that a photographed event happened. Its absence also does not prove that an asset is fake. See the C2PA explainer; the specification index currently lists version 2.4.
A hard binding uses cryptographic information to associate the manifest with specified asset content, following the format's binding rules. It need not mean a naive hash of every file byte: embedded manifests require appropriate exclusions. A soft binding, such as a watermark or fingerprint, can help discover associated provenance after some transformations. Recovery depends on the binding, transformation, detector and repository; it is not guaranteed to survive every edit. The C2PA technical specification defines these mechanisms.
For this design, keep the detailed production manifest private. Publish only appropriate provenance assertions and source references. Do not expose private prompts, customer IDs, credentials or source documents merely because a format supports metadata. Preserve provider credentials where possible and create a correctly linked new record for edited or transcoded derivatives. Protect signing keys and support revocation; a signed false statement remains false.
Watermarks and detection are additional evidence
A watermark embeds a detectable signal in content. Some systems are designed to tolerate common resizing, recompression or other edits. For example, SynthID covers several media types. It is not a universal detector for content produced by unrelated systems.
Robustness depends on the watermark and threat model. Research such as the NeurIPS 2024 watermark-removal study demonstrates attacks under specified conditions; it does not justify claiming every possible watermark always fails. Likewise, a classifier's “AI-generated” score is probabilistic evidence, not proof of origin or a substitute for consent records.
Put controls at input, output and publication boundaries
- Authenticate and authorize. Restrict source files, projects, collaborators and output access to their permitted scope.
- Validate uploaded files. Check decoded media, dimensions, duration, size and parser behavior. Use bounded processing and isolated converters.
- Establish rights and consent. Record permissions for voice cloning, likeness, music and brand assets with their intended use. A model license and a person's consent are separate requirements.
- Screen prompts and outputs. Apply the product's policies to sexual content, impersonation, violence and other prohibited uses. Evaluate false positives and missed violations; classifiers are fallible.
- Review consequential or ambiguous work. Route uncertain cases to qualified people and offer an appeal or correction path.
- Approve final bytes. Enforce exact revision, rendition hashes, policy version and unexpired authorization before release.
- Respond to abuse. Receive reports, preserve appropriate evidence, revoke delivery access and remove prohibited copies according to applicable obligations.
Laws depend on the system's role, jurisdiction, content and exceptions. Under EU Article 50, provider marking/detection duties and deployer deepfake disclosures are distinct obligations; standard editing and creative works have specific treatment. Article 50 must be read with the amended timetable: transparency applies from August 2, 2026, while qualifying systems already on the market have until December 2, 2026 for Article 50(2) marking/detection compliance. The Commission's current enforcement FAQ explains that transition.
In the US, the TAKE IT DOWN Act's platform provisions concern covered platforms and qualifying intimate imagery. The FTC describes a 48-hour removal duty after a valid request, with reasonable efforts to remove known identical copies. Do not generalize this into a blanket rule for every media service or every report. Use the FTC's compliance guide and the broader governance chapter when assigning requirements.
Evaluating Generative Quality
Quality is a collection of requirements. A beautiful image can contain the wrong product; an intelligible voice can say the wrong price. A preference vote cannot settle factual accuracy or ownership.
Use the right measure for the requirement
| Measure | What it measures | Practical limit |
|---|---|---|
| FID: Fréchet Inception Distance | Distance between Gaussian approximations to real and generated image feature distributions, conventionally using Inception features | Depends on sample size, preprocessing, reference set and feature space; not a per-image truth score |
| CLIPScore | Image-text compatibility using CLIP representations | Can miss fine detail, counting, facts or context; alignment is not overall quality |
| FVD: Fréchet Video Distance | Distributional distance in learned video-feature space | Depends on temporal sampling, features and reference data; one low score does not certify a clip |
| FAD: Fréchet Audio Distance | Distributional distance in audio-embedding space | Embedding choice and reference distribution affect conclusions; not a speech-transcript correctness metric |
| MOS: mean opinion score | Average ratings from listeners under a defined subjective test | Specify the scale, question, listeners and conditions; scores from different protocols need not compare |
| OCR / transcription checks | Text or speech content relative to required wording | Recognizers can also err; verify critical facts directly |
| Human rubric and blind preference | Task-specific quality judgments | Requires representative cases, qualified raters and disagreement handling |
| VLM or audio-model judge | Automated rubric assessment of media | Can overlook defects or favor particular styles; calibrate against human review |
| pHash / SSIM or embedding similarity | Particular forms of similarity between outputs | A valid creative alternative may look different; similarity does not prove correctness |
The original CLIPScore, FVD and FAD papers define different evaluation targets. Rethinking FID studies image-metric shortcomings and an alternative based on CLIP embeddings; FAD embedding research shows why encoder choice matters. No metric in this table establishes legal permission.
Build a release evaluation, not a beauty contest
- Version a test set covering actual content, languages, aspect ratios, difficult inputs and prohibited requests. Keep a held-out set for release decisions.
- Define hard constraints first: correct product identity, exact required text, supported format/duration, allowed source use and prohibited-content checks.
- Define quality rubrics: prompt adherence, artifacts, temporal consistency, speech intelligibility, lip synchronization and usefulness to the customer.
- Compare candidate and current systems on matched cases, with repeated generations when output variability matters. Blind and randomize subjective comparisons.
- Measure rejection, regeneration, review time and cost per accepted asset, alongside latency and preference.
- Apply practical regression margins and uncertainty intervals by important slice. One severe policy failure can block release regardless of an average score or statistical significance.
- Canary the change, watch drift and keep a rollback path for model, prompt, adapter and workflow versions.
A statistically significant improvement can be too small to matter. An important loss on a small language slice can lack statistical significance because the sample is insufficient. Report that uncertainty rather than calling it safe. Changes in input mix, judges or preprocessing can also cause score drift; do not immediately blame an unannounced provider update.
For byte-preserving operations such as retrieving a stored asset, exact equality is appropriate. For creative generation, evaluate requirements and distributions rather than requiring every new output to resemble a single “golden” picture. Public rankings are a candidate-selection aid; verify their current methods and test on your own distribution. See benchmarks and leaderboards.
The Model Landscape
Use this as a procurement checklist, not a ranking. Select from current API contracts and a task-specific evaluation, then pin the model or record the version returned by the service.
| Requirement | Current options to investigate | Decision to verify |
|---|---|---|
| Hosted image generation/editing | OpenAI's image guide lists gpt-image-2.5-sunburst; BFL provides FLUX image endpoints |
Reference support, edit fidelity, dimensions, version pinning, price and retention |
| Hosted multimodal video | Google's video guide recommends Gemini Omni Flash for general generation; Veo 3.1 supports specified workflows such as extension and last-frame control | The applicable API, supported inputs/durations, native audio and editing limitations |
| Another current video implementation | BFL's FLUX 3 Video supports video generation/editing and optional native audio, with draft/enhancement workflows | Supported duration/resolution, quota, actual draft-plus-enhancement billing and final quality |
| Owned image inference | FLUX.2 klein 4B is one Apache-2.0 model; licenses differ for other variants | The exact checkpoint license, hardware, safety requirements and adaptation compatibility |
| Speech and music | Task-specific speech/TTS/music services or licensed source tracks | Voice consent, music rights, allowed distribution, language quality and retention |
| Retired integration | OpenAI Sora 2 models and the Videos API have a documented shutdown date of September 24, 2026 | Migrate existing work; do not select the retired endpoint for a new service |
Primary checks: OpenAI images, Google video generation, BFL FLUX 3 Video, FLUX.2 klein 4B model card, and OpenAI deprecations. The voice chapter covers speech stack selection.
Do not infer that an entire model family is open source or commercially usable from one permissive checkpoint. For example, BFL's non-commercial and self-hosted commercial terms differ, and the FLUX 3 page describes separate rollout stages for video, image, action and open weights. Model-use permission, output-use terms, source-material rights and likeness consent require separate checks.
An adapter must implement the actual provider contract. BFL's image quickstart returns a polling URL and a temporary output URL; it documents a ten-minute expiry for the signed result URL. Follow the documented polling destination after validating its provider origin, copy the output to authorized durable storage, and avoid giving customers a provider URL as their permanent asset record. Never send a provider credential to an arbitrary user-supplied URL.
Joint audio-video or a cascade?
| Choice | Good fit | Tradeoff |
|---|---|---|
| Joint generation | A short scene where motion, speech and ambient sound should be coordinated | Changing one component may require regenerating more of the scene; evaluate synchronization rather than assuming it is perfect |
| Separate narration and video | Exact approved wording, reusable voices, dubbing and independently edited tracks | Must manage duration, transitions, loudness and synchronization explicitly |
| Hybrid | Native ambient sound with a separately approved narration track | More control, but careful mixing and prevention of conflicting speech are needed |
In a cascade with visible speech, lip-sync depends on both the video and the final speech track. Putting lip-sync before its speech input is a dependency error. A voice-over on a product shot usually needs no lip-sync stage at all.
Interview design: a branded 30-second video service
Prompt: “Design a service where a business uploads a script and product images, reviews a draft, and downloads a narrated promotional video.”
1. Clarify scope
Ask about duration, required words, output formats, languages, identity/voice permissions, review responsibility and acceptable wait time. Confirm whether the service publishes to advertising accounts or only produces files.
For this interview, agree on the following scope: a private business workspace produces three-shot, 30-second videos with voice-over. It exports a final MP4 and captions. Publishing to social accounts, unrestricted celebrity cloning, live generation and long-form film editing are outside the first release.
2. Functional requirements
- Upload and authorize source pictures, script, brand settings and permitted narration voice.
- Validate inputs and display a cost estimate before paid generation begins.
- Create a storyboard and narration-text preview; let the customer approve or revise them before generating the media.
- Generate shots and compose the selected revision into a 30-second draft.
- Regenerate an individual shot without losing valid work from other stages.
- Show durable progress, failure reasons, cancellation state and actual usage.
- Check and approve the final rendition, then provide an authorized download and captions.
- Keep source, generation, approval and publication records; support retention and deletion policy.
3. Non-functional requirements
- Latency: acknowledge accepted work within one second at p95; target automated completion within ten minutes at p95 for the standard job class, excluding time awaiting customer review. Treat this as a target to validate under load.
- Durability: no acknowledged job is lost after a worker restart; status must explain partial completion or an unresolved outcome.
- Correctness: publish only the approved revision and validated rendition; preserve exact required words and product facts.
- Isolation: keep tenant inputs, outputs, cache entries and review records within authorized scope.
- Cost: enforce concurrent reservations, per-job limits and tenant limits before submitting additional billable work.
- Capacity: support 1,000 projects per working day, with four times average arrival rate at peak.
- Operability: expose stage latency, queue age, retry charges, rejection causes, unresolved provider operations and cost per accepted project.
Interview tip: a deadline is not a promise that every third-party request finishes. Say how the UI and refund/credit policy handle a deadline miss.
4. Estimate the scale
Assume a ten-hour active day, twenty working days per month, three ten-second generated shots per project, and a mean active generation time of 90 seconds per shot. Use a provider that supports the chosen clip contract; otherwise generate supported lengths and trim while accounting for their full charge.
| Calculation | Result | Meaning |
|---|---|---|
| Projects per month | 1,000 × 20 = 20,000 | Denominator before rejection |
| Baseline shot jobs per day | 1,000 × 3 = 3,000 | Other stages need separate sizing |
| Peak shot arrival rate | 3,000 / 36,000 × 4 ≈ 0.333/s | Four times the active-day average |
| Peak rate with 15% additional generation attempts | 0.333 × 1.15 ≈ 0.383/s | Assumed rerender workload |
| Mean active slots at that rate | 0.383 × 90 = 34.5 | Arrival rate × mean service time |
| Slots at 70% target utilization | ceil(34.5 / 0.70) = 50 | A planning estimate, not a p95 queue guarantee |
| Encoded video at 8 Mb/s for 30 seconds | 8 × 30 / 8 = 30 MB | Before audio and container overhead |
| One month's final encoded videos | 20,000 × 30 MB = 600 GB | Excludes drafts, sources, replicas and downloads |
For comparison, uncompressed 1920 × 1080 RGB frames at eight bits per channel, 30 frames/s, for 30 seconds occupy about 5.6 GB per video. Encoding changes the storage requirement dramatically. Do not estimate an encoded MP4 from its raw pixel count.
The 50 slots may be a provider concurrency allocation, not 50 physical GPUs. Owned hardware needs benchmarks for model, resolution, duration, batching, accelerator memory and co-location. Autoscaling cannot compensate for an unavailable provider quota.
5. Draw a baseline, then identify its flaws
The baseline is one worker that generates everything sequentially and saves the final file. It is enough to prove the product flow on a small workload.
| Baseline flaw | Observed consequence | Repair and its cost |
|---|---|---|
| A worker owns all progress in memory | Restart loses successful intermediate work | Persist stage state and immutable assets; more storage and workflow logic |
| Retry the entire video on any failure | Duplicate generation and different previously approved shots | Retry only reconciled failed stages; requires explicit dependency tracking |
| One queue for model calls and encoding | Long video calls block cheap tasks | Separate stage queues and limits; additional scheduling |
| Approval refers only to project ID | A revised or newly rendered file can bypass review | Bind approval to revision and final asset hashes; repeated review when outputs change |
| Final link points to the provider | Download expires or private inputs become accessible | Copy to private storage and authorize delivery; storage/egress costs |
| One average “quality score” | Wrong prices or rights violations pass a beauty test | Hard constraints plus quality rubrics; more validation and qualified review |
6. Develop the detailed design
Read diagram source
flowchart TB
C[Workspace client] --> API[API: identity, scope, quotas]
API --> DB[(Projects, jobs, approvals, cost ledger)]
API --> UP[(Private source storage)]
DB --> O[Outbox and durable orchestrator]
O --> PLAN[Validate script and create storyboard]
PLAN --> REVIEW[Approve storyboard and narration text]
REVIEW --> IQ[Image queue]
REVIEW --> AQ[Audio queue]
IQ --> IMG[Image adapter]
IMG --> VQ[Video queue]
VQ --> VID[Video adapter]
AQ --> TTS[Authorized TTS and licensed music]
VID --> AS[(Immutable intermediate assets)]
TTS --> AS
AS --> COMPOSE[Composition and encoding pool]
COMPOSE --> CHECK[Media, wording, policy and quality checks]
CHECK --> FINAL[Final rendition approval]
FINAL --> PUB[Publication transaction]
PUB --> DELIVERY[Authorized CDN or signed download]
O --> RECON[Webhook inbox and status reconciler]
RECON --> DB
CHECK --> DB
The orchestrator creates versioned stage inputs. Before dispatch it checks the current project revision, scope, policy, budget and dependencies. Workers return immutable outputs; they do not decide which revision is public. Only the publication transaction changes the released package pointer.
The stage graph represents dependencies, so independent work can run concurrently:
Read diagram source
flowchart LR
S[Approved script revision] --> N[Narration track]
S --> B[Storyboard and product references]
B --> V[Three generated shots]
S --> M[Licensed or permitted music]
N --> C[Compose voice-over video]
V --> C
M --> C
C --> R[Encode required renditions]
R --> Q[Check actual media and captions]
Q --> A[Approve final hashes]
If a later version includes a speaking avatar, insert a lip-sync stage after the relevant video and narration outputs. If a script change alters spoken duration, invalidate downstream timing and composition, and regenerate shots only when their content or duration requirements have changed.
Data model: projects, project_revisions, jobs, stage_attempts, assets, asset_dependencies, rights_records, approvals, usage_reservations, usage_settlements and publications. Unique constraints prevent duplicate logical jobs and settlements. Every lookup applies tenant scope. A versioned publication transaction checks required stage completion, asset hashes, current rights and approval before releasing the package.
API sketch:
| Endpoint | Contract |
|---|---|
POST /projects/{id}/generations |
Authorized revision and idempotency key; return a stable job ID and reserved allowance |
GET /jobs/{id} |
Current stage, safe status, cost summary and any action the customer must take |
POST /jobs/{id}/cancel |
Request cancellation; report confirmed versus pending provider work |
POST /projects/{id}/approvals |
Approve a specified revision/rendition set after access and role checks |
POST /projects/{id}/publications |
Atomically validate and release the approved package |
GET /assets/{id}/download |
Check current authorization and issue a short-lived download |
A signed download URL is a temporary bearer capability. Expiry limits its duration; immediate revocation may need an authorization gateway or CDN invalidation strategy. Do not describe an already issued URL as instantly revoked merely because a database flag changed.
7. Walk through failures and repairs
| Failure | Response | Remaining tradeoff |
|---|---|---|
| Third shot fails after two succeed | Keep the first two; reconcile the failed attempt, then retry within budget | A replacement shot may need continuity review |
| Provider accepted a request but response was lost | Reconcile using provider identity/deduplication; hold an unknowable outcome | Customer may wait; blind retry risks duplicate charges |
| Callback is forged, duplicated or late | Verify, deduplicate, and apply versioned transitions | Polling/reconciliation still needs capacity |
| Customer cancels while provider completes | Stop new stages, quarantine late outputs, reconcile actual spend | Cancellation may not reverse provider charges |
| Worker crashes after copying output | Recover asset and stage records using stable IDs and hashes | Clean up unreferenced uploads after a safe retention interval |
| Provider URL expires before copy | Recover the same output through the provider if supported | If unrecoverable, explain the failure before any paid regeneration |
| Final render changes the product label | Fail the exact-content check and require correction/reapproval | A cheap draft approval does not authorize a defective final |
| A customer loses permission to use a voice | Block new use and apply the relevant removal/retention decision to existing assets | Previously downloaded copies cannot be remotely erased |
| Provider outage or retirement | Pause admissions or use a tested compatible fallback | Different model output may require new approval and cost estimate |
| One tenant floods the service | Tenant queues/limits, fair scheduling and global spend admission | Some work is delayed rather than consuming all capacity |
A fallback is a new implementation of the task, not a string substitution in an endpoint URL. Recheck allowed input types, licenses, output behavior, price, retention and evaluation thresholds.
8. Calculate complete operating cost
Assume these illustrative rates for 20,000 projects/month. They are deliberately separate from any vendor's current price card.
| Cost | Calculation | Monthly estimate |
|---|---|---|
| Initial video generation | 20,000 × 3 × 10 seconds × $0.08/s | $48,000 |
| Additional generation attempts | 15% × initial generation cost | $7,200 |
| Storyboard images | 20,000 × 4 × $0.03 | $2,400 |
| Narration and music allowance | 20,000 × $0.05 | $1,000 |
| Composition and automated checks | 20,000 × $0.10 | $2,000 |
| Storage and delivery allowance | Assumed monthly amount | $750 |
| Maintenance and operations | Assumed allocated monthly cost | $4,000 |
| Internal quality review | 20% × 20,000 × 4 minutes / 60 × $45/hour | $12,000 |
| Total | Sum of listed costs | $77,350 |
This is $3.87 per attempted project. If 90% become accepted assets, the denominator is 18,000 and cost becomes $4.30 per accepted project. Internal review alone takes about 267 hours/month. Customer approval time is separate; music licensing, payment fees, taxes, support incidents or source retention may require additional line items.
If extra generation rises from 15% to 30%, add $7,200/month. A provider charging less per second can still be more expensive per accepted asset if its output triggers more rejection and review. Benchmark the full workflow before procurement.
9. Close the interview
“I would launch the scoped voice-over workflow with durable jobs, private immutable assets and approval of the actual final rendition. The main scaling controls are provider-aware admission and separate generation/encoding queues. The main correctness controls are revision-bound dependencies, reconciliation of unknown outcomes and a publication gate. I would validate the ten-minute target and unit economics with realistic traffic and rejection rates before expanding to avatars, more languages or direct social publishing.”
Interview Questions
1. Is every multimodal generation request necessarily asynchronous?
No. A short image operation can complete within the application's request budget. Long or interruption-prone jobs benefit from durable asynchronous execution. Explain the latency target, timeout behavior and recovery contract rather than prescribing one transport universally.
2. What is the difference between a diffusion model and a transformer?
Diffusion describes a generative modeling approach involving a noising process and learned reversal. A transformer is a neural-network architecture. A diffusion or flow-based generator can use a transformer; these are not exclusive categories.
3. Why does an idempotency key not automatically prevent two charges?
It deduplicates only where it is enforced. The local database can recognize the same request, but a provider may have accepted a timed-out submission. Safe recovery requires provider deduplication or reconciliation; otherwise the outcome remains uncertain.
4. The user asks for another image with the same prompt. Should the cache return the previous image?
Only if the product action requests reuse. “Generate another variant” is a new generation intent. Retrieval of an existing result is different from sampling again, even with identical text.
5. Can a fixed seed reproduce an approved video next month?
A seed alone cannot promise that. The execution environment, model version, parameters and provider contract matter. Store the approved bytes so access to the accepted result does not depend on regeneration.
6. How should a changed narration affect the workflow?
Create a new revision and invalidate dependent timing, lip-sync if present, composition and final approval. Reuse an unchanged shot only if its content, duration, rights and compatibility still satisfy the new revision.
7. Does a valid C2PA credential prove that an image is real?
No. It supports verification of signed provenance and bound content. Factual truth and permission to use the depicted subject require separate evidence. Missing credentials do not prove falsity either.
8. How do you scale a hosted video API?
Bound submissions to provider quotas, schedule tenants fairly, measure queue age and service time, and obtain additional capacity or test a fallback when needed. More local workers do not create more remote quota.
9. Why can the cheapest generator have the highest total cost?
Extra attempts, rejected work, internal review, encoding, storage and support affect the denominator. Compare complete cost per accepted asset at the required quality, not only the first-call rate.
10. Should you gate a creative-model release only on statistically significant FID improvement?
No. FID is one distributional measure. Use hard task requirements, safety constraints, representative human evaluation, practical regression margins, uncertainty and cost. Statistical significance does not establish usefulness or permission.
11. What if the draft was approved but the final render changes the text?
Final validation must inspect the delivered rendition. A changed label can fail an exact-content requirement even if the draft was correct. Repair it and obtain the required final approval.
12. How do you cancel a paid generation safely?
Record the request, stop dispatching new work and invoke provider cancellation if available. Reconcile whether generation completed and what was charged. A late result must not become public merely because the worker finished.
13. Why keep private production metadata separate from public provenance?
Debugging may require sensitive prompts, source references and account records. Public credentials should disclose appropriate origin/edit information without leaking private data. Both records can reference the same immutable asset identity.
14. When would you use a cascade instead of native audio-video generation?
When exact narration, separate language tracks or independent editing matter enough to justify synchronization work. Native generation can simplify coordinated scenes, but neither option universally has the best quality or lowest total cost.
15. What would you check before commercial use of open weights?
The exact model license, any commercial-use conditions, derivative/adapter terms, allowed deployment and output-use terms. Also check training/reference material rights and likeness or voice consent. A permissive software or model license does not grant those other rights.
Final revision cards
| Remember | Explain it in the interview |
|---|---|
| Intent → job → stage → asset → publication | Different identities prevent retries, revisions and approvals from being confused |
| Accepted is not completed | Durable status and reconciliation handle long-running provider work |
| Unknown is not failed | Do not repeat a costly operation until retry safety is established |
| Approve the delivered bytes | Draft approval alone cannot certify a changed final rendition |
| Origin ≠ truth ≠ permission | Provenance, factual verification and usage rights answer different questions |
| Rate × time estimates active work | Then add utilization headroom and test actual queue behavior |
| Cost / accepted assets | Include rerenders, review, operations and rejected outputs |
| Model version is part of the workflow | A provider change can require evaluation, migration and new approval |
For practice, draw the workflow in five minutes, explain one ambiguous provider failure, then calculate how a doubled rejection rate changes cost and staffing. Close with the smallest useful release and the evidence needed to expand it.