Interview problem: build a tutor that explains approved course material, gives hints, checks answers and recommends the next exercise. Optimize learning and independent problem solving rather than conversation length.
This is a hypothetical Learnastra-style product exercise, not a description of Learnastra's deployed architecture. Volumes, targets and costs are assumptions. Start with adult learners and one bounded subject; educational use involving children adds separate consent, privacy and safeguarding requirements.
1. Clarify the learning contract
Ask what the learner should be able to do without assistance, which curriculum is authoritative, what counts as acceptable feedback, and whether assessments are practice or consequential grading. Assume formative interview practice; the system does not make hiring or certification decisions.
Functional requirements
- Explain concepts from a versioned, instructor-reviewed curriculum with source links.
- Offer graduated hints before a worked solution when the learner requests practice.
- Check free-text and structured answers against an explicit rubric; allow correction and disagreement.
- Track attempted skills and schedule review exercises with user control.
- Recommend exercises by prerequisites, observed gaps and the learner's stated goals.
- Let learners inspect, edit or delete retained profile information and export progress.
Non-functional requirements
- Target p95 first useful feedback below two seconds and completed ordinary feedback below eight seconds.
- Target 99.9% monthly practice-session availability, with static lessons and answer keys available during model outages.
- Measure delayed independent performance, feedback correctness and harmful/unsupported instruction separately.
- Isolate learner records and make deletion propagate to derived memory and analytics according to the retention contract.
- Bound calls, output lengths and weekly learner spend; avoid unlimited conversational loops.
2. Estimate the workload
Assume 20,000 daily learners, ten tutor turns each: 200,000 turns/day, about 2.31/s averaged over a day. A 20× peak is about 46.3/s. At a six-second mean, approximately 278 requests are in flight. If each turn uses 1,500 input and 250 output tokens, daily totals are 300M input and 50M output tokens.
A hypothetical $1/M input and $5/M output price gives $300 + $250 = $550/day for one model call per turn. Separate grading, hint generation and retries can multiply that. Compare complete practice-session cost and independently solved exercises, not tokens alone.
3. Build a small baseline
Read diagram source
flowchart LR
L[Learner] --> API[Authenticated tutoring API]
API --> C[(Reviewed lessons and exercises)]
API --> M[Model generates a bounded hint]
M --> CHECK[Rubric and source checks]
CHECK --> L
API --> P[(Exercise attempts)]
Start with authored exercises and answer rubrics. A deterministic checker can grade a numeric answer or validated code test; a model may provide a supplementary explanation. Preserve uncertainty for open-ended answers.
4. Find flaws and evolve deliberately
| Flaw | Repair | Benefit | Cost or limit |
|---|---|---|---|
| Tutor reveals answers immediately | Explicit hint stages and learner-controlled reveal | Supports effort before disclosure | Some learners want faster answers |
| Fluent feedback marks a correct alternative wrong | Rubric examples, alternative-answer tests and appeal | Fairer, more reliable feedback | Author/reviewer effort |
| Recommendation rewards easy completion | Delayed unaided assessment and difficulty-aware reporting | Better evidence of learning | Longer experiments |
| Conversation summary invents a learning preference | Provenance, confidence and editable memory | Prevents stale inferred facts driving the path | More state lifecycle work |
| A learner pastes instructions into an answer | Treat answer as untrusted data, enforce tools in code | Reduces authority confusion | Does not guarantee perfect model behavior |
5. Detailed architecture
Read diagram source
flowchart TD
L[Learner web or mobile app] --> A[Identity and session API]
A --> O[Bounded lesson orchestrator]
O --> CUR[(Versioned curriculum and prerequisite graph)]
O --> RET[Authorized lesson retrieval]
RET --> CUR
O --> H[Hint or explanation generator]
O --> G[Deterministic or rubric-based grader]
H --> V[Grounding and response checks]
G --> V
V --> L
A --> E[(Append-only attempt events)]
E --> S[Skill evidence projection]
S --> REC[Prerequisite-aware recommendation]
REC --> O
A --> MEM[(Editable learner goals and preferences)]
MEM --> O
E --> EV[Offline evaluation and delayed learning experiments]
EV --> RELEASE[Versioned release gate]
RELEASE --> O
Keep observed evidence separate from inferred skill estimates. An incorrect answer can indicate a misconception, a typo, an ambiguous question or bad grading; it is not automatically proof of low ability.
6. API and data design
POST /sessions chooses subject and goals. POST /sessions/{id}/turns accepts an idempotency key, exercise version and learner answer. POST /attempts/{id}/appeals records disputed feedback. DELETE /learners/me/memory/{id} revokes a retained preference.
| Record | Fields | Rule |
|---|---|---|
| Exercise | ID, version, skill IDs, prerequisites, prompt, rubric, author approval | Grade against the attempted version |
| Attempt | learner ID, exercise version, answer reference, hint level, grade version | Do not erase the distinction between assisted and unassisted work |
| Skill estimate | learner, skill, supporting attempt IDs, confidence, updated time | Derived estimate, never an immutable fact |
| Memory | source, user-confirmed flag, expiry, permission scope | Retrieved preferences are data, not executable instructions |
Store sensitive free text separately from minimal attempt metadata, with appropriate access and retention. User deletion creates a tombstone used by retrieval and projection workers so delayed events do not recreate removed memory.
7. Trace a practice session
- Authenticate the learner and load their explicit goal and current exercise state.
- Pick an exercise from prerequisite-eligible material; explain why it was suggested.
- Record the answer and assistance level before grading.
- Use a deterministic checker where valid; otherwise apply the pinned rubric with calibrated uncertainty.
- Generate bounded feedback from the rubric and authorized lesson evidence.
- Validate references and return feedback, a next hint or a clear inability to judge.
- Append the event, update skill evidence and schedule a future independent check.
Two concurrent devices may submit different answers. Use attempt IDs and optimistic version checks; do not let the latest arrival silently replace the first attempt. A model timeout does not erase the submitted answer.
8. Failure tests and evaluation
| Test | Expected behavior |
|---|---|
| Model outage | Preserve answer; show authored hint/key or retry state |
| Grader disagreement | Mark uncertain; retain both evidence and rubric; enable review |
| Curriculum update mid-session | Finish against pinned version or explicitly restart |
| Deleted preference arrives from a delayed queue | Tombstone/version check rejects resurrection |
| Learner repeatedly asks for final answers | Respect the selected study mode; report assisted performance honestly |
| Prompt injection in an uploaded exercise | No new tool authority; source treated as untrusted input |
Evaluate feedback correctness on reviewed answer variants, groundedness, age/subject suitability for the supported audience, and accessibility. For learning outcomes, compare matched or randomized groups using delayed unaided tasks; account for baseline ability and attrition. Clicks and time in chat are engagement, not proof of learning.
9. Decisions and cost-benefit
| Choice | Benefit | Cost or risk |
|---|---|---|
| Authored question bank first | Known quality and reviewable rubrics | Limited breadth and authoring cost |
| Generated variations | More practice diversity | Must validate correctness and difficulty |
| Simple skill evidence counters | Explainable and easy to correct | Coarse personalization |
| Probabilistic knowledge tracing | Can model uncertainty and forgetting | Data, calibration and explanation burden |
| Human feedback review | Catches ambiguous grading | Staffing delay and cost |
Launch one subject with a reviewed set, compare to static worked examples, then expand. At $550/day model cost and 20,000 daily learners, the model line is $0.0275/learner-day before other calls, storage and staff. A cheaper model that causes more incorrect feedback is not a saving under the quality requirement.
10. Interview questions
Q1: How would you measure whether personalization works?
Sample answer: Use delayed independent exercises aligned to the intended skills, not the same examples seen during tutoring. Compare outcomes under a controlled study, record hint usage, account for prior ability and attrition, and review incorrect feedback separately from recommendation quality.
Q2: Should every wrong answer reduce a mastery score?
Sample answer: No. Preserve the observation and its context first: exercise difficulty, assistance, grader uncertainty and whether the prompt was ambiguous. Update an estimate with those limitations and let a learner dispute the result.
Q3: Why not use the whole chat history forever?
Sample answer: It adds cost, irrelevant or stale context and privacy exposure. Keep recent task state plus selected, attributable preferences under retention and deletion controls. A compact summary remains fallible and should not become authoritative.
Closing remarks and recall notes
I would start with reviewed exercises, explicit hint stages, auditable feedback and learner-controlled state. Personalization should improve demonstrated understanding. The main tradeoff is adaptation breadth versus feedback reliability and the cost of maintaining high-quality curriculum.
| Remember | Evidence |
|---|---|
| Teach a defined skill | Curriculum and rubric |
| Separate help from mastery | Assistance recorded on attempts |
| Personalize with uncertainty | Editable preferences and attributable estimates |
| Measure later | Independent delayed performance |
Tip: A tutor is a learning system. Show where the learner's progress is measured independently of the language model's enthusiasm.