In-context learning (ICL) is a model's adaptation to a task through information supplied in its context, without updating its parameters. Few-shot prompting supplies a small number of demonstrations of the desired input-to-output behavior. A demonstration is information available to the current inference, not a permanent training update.
The GPT-3 paper demonstrated task performance using instructions and examples without gradient updates. It also reported limitations; ICL does not guarantee that any arbitrary task can be learned from a handful of examples. Brown et al..
Understand what an example teaches
| Component | Purpose | Failure if poorly chosen |
|---|---|---|
| Input | Represents the task's real data | Clean examples conceal production noise |
| Correct output | Demonstrates the required decision | Wrong labels teach the wrong boundary |
| Optional brief explanation | Clarifies a subtle distinction | Plausible but wrong explanations create confusion |
| Consistent format | Makes the contract recognizable | Mixed schemas encourage inconsistent output |
Demonstrations can convey label names, format and input distribution as well as task-specific relationships. Min et al. found that these factors explained substantial gains in the classification and multiple-choice settings they studied. That result is not permission to use incorrect labels in a production example bank. Rethinking the Role of Demonstrations.
A worked decision boundary
Use the routing rubric from prompt fundamentals. The useful examples distinguish requested actions, rather than simply containing keywords.
| Ticket | Label | What the example teaches |
|---|---|---|
| “End my subscription at renewal.” | cancel |
Cancellation without a refund request |
| “Return the duplicate charge; keep my subscription active.” | refund |
A refund does not imply cancellation |
| “Cancel my subscription and return this month's payment.” | multiple |
Both actions are explicitly requested |
| “What is your cancellation policy?” | other |
Asking about a policy is not requesting execution |
For an unseen ticket, “Please refund the charge. Do not cancel the service,” the expected label is refund. This is an original practice example, not evidence that a particular model passes the task.
A brief explanation can describe why a label is correct. Do not require verbose intermediate reasoning for every example. Longer demonstrations spend tokens and can make a concise classification task harder to maintain.
Choose the number of examples experimentally
There is no universal example count determined by parameter size. Task difficulty, demonstration quality, the model's prior training and the input budget all matter.
- Evaluate a zero-shot rubric.
- Compare a small sweep, such as 2, 4 and 8 examples, on a development set.
- Inspect per-label precision/recall and ambiguous cases, not only the mean score.
- Measure token count and latency after rendering the actual chat template.
- Keep the smallest set that meets the chosen quality and operating requirements.
- Confirm the selected version on an untouched test set.
Suppose 1,200 tokens remain after instructions, the current ticket and response allocation. If each serialized example costs about 200 tokens, approximately six examples fit. This is a budget estimate; tokenizer and template overhead must be counted. It is not a recommendation that six examples are optimal.
Needing twenty examples does not establish that fine-tuning is necessary. Compare the cost and quality of longer context, better demonstrations, a different model and fine-tuning using actual traffic and maintenance needs.
Static versus dynamic selection
Read diagram source
flowchart LR
Q[Current request] --> S[Select permitted examples]
B[Reviewed example bank] --> S
S --> D[Deduplicate and cover boundaries]
D --> P[Assemble prompt within budget]
P --> M[Model prediction]
M --> V[Validate and measure]
| Strategy | Advantage | Main risk |
|---|---|---|
| Fixed examples | Easy to inspect, reproduce and cache | May miss a specialized input |
| Similarity retrieval | Can select locally relevant demonstrations | Similar wording may hide a different label |
| Diversity-aware selection | Covers distinct boundaries | More selection logic and tuning |
| A fixed core plus retrieved examples | Stable contract with task-specific coverage | Needs conflict and budget handling |
Dynamic retrieval is an option, not a default improvement. A nearest-neighbor search may return five almost identical cancellation tickets, omitting the refund/cancellation distinction. Compare semantic similarity with boundary coverage, remove near-duplicates, and use only examples the current tenant may access.
Version the bank and selector. Exclude evaluation targets and close duplicates from selection to prevent leakage. Review user-contributed examples before making them part of application instructions.
Label balance, order and distribution
Examples can influence output tendencies. A balanced set helps expose all labels, while a frequency-matched set reflects the workload's prior distribution. These are different choices; neither is universally correct. Use a representative evaluation set and inspect rare-category errors before deciding.
Order can matter too. Test a few permutations when the deployment is sensitive to it, and keep a stable order for diagnosis. Shuffling at runtime without measurement can add variance rather than solve bias.
Incorrect predictions after adding examples are not evidence of catastrophic forgetting of model weights: inference has not changed those weights. Diagnose conflicting conditioning, misplaced emphasis, truncation or a task the model cannot reliably perform.
Interview practice
Q1: Is in-context learning the same as fine-tuning?
No. ICL conditions inference on supplied information. Fine-tuning updates parameters. The application may store demonstrations for future prompts, but storage does not turn those examples into weight updates.
Q2: How would you choose between two and ten demonstrations?
Measure both against the same rubric, model and held-out cases. Compare category errors, cost and latency. Ten may supply missing boundaries, or may repeat information and introduce conflicts. The choice follows measured benefit per additional token.
Q3: Why can the most similar example be the wrong example?
Similarity can follow vocabulary rather than the decision boundary. “Explain how to cancel” and “Cancel now” may be close in embedding space but require different labels. Include counterexamples and test the selector as part of the complete pipeline.
Q4: Should every label have the same number of examples?
Not automatically. Balanced examples can clarify all labels, while actual class frequencies affect calibration and evaluation. I would compare selections on representative traffic and important rare cases, with explicit costs for different mistakes.
Q5: Can a demonstration include an explanation?
Yes, when it helps specify a difficult distinction and has been checked. It is optional. An incorrect explanation can harm quality, and a long explanation adds cost. For factual or arithmetic decisions, use independent evidence or computation to validate the output.
Q6: What would make an example-selection experiment misleading?
Retrieving the test answer or a near-duplicate, repeatedly tuning against the final test set, or comparing different model settings. I would separate datasets by relevant entity/time boundaries and log the selected example IDs for each evaluation request.
Final notes
Recall card: Correct examples, meaningful boundaries, controlled budget, independent evaluation. Examples demonstrate behavior; they do not confer authorization or permanently train the model.
Practice: Remove one row from the worked table. Name the new ambiguity, construct a test ticket for it, and decide whether a clearer rule could replace the missing example.