Chain-of-thought (CoT) prompting elicits intermediate reasoning steps before a final answer. The original few-shot approach included worked reasoning demonstrations in the prompt. It improved results on several reasoning benchmarks in the studied models; it is not a guarantee of correct reasoning on every task. Wei et al..
For system design, separate three questions: how additional computation is requested, how an answer is checked, and what explanation the user should receive. These have different requirements.
Distinguish the techniques
| Technique | What changes | What it does not establish |
|---|---|---|
| Zero-shot CoT | An instruction requests intermediate reasoning | Correctness of the resulting steps |
| Few-shot CoT | Demonstrations include worked steps | Generalization to every new problem |
| Reasoning-model effort setting | The API allocates model-specific inference effort | A universal latency or quality multiplier |
| Explicit workflow | Application code splits work into stages | Correctness of each generated stage |
| Self-consistency | Several sampled solutions are aggregated | Independence of their mistakes |
| External verification | A calculator, test or source checks a claim | Correctness beyond what the check covers |
Kojima et al. studied zero-shot reasoning instructions, and Wang et al. studied sampling multiple reasoning paths and aggregating answers. Treat their results as evidence for the tested methods and workloads, not as timeless performance promises for every contemporary API. Zero-shot reasoning, self-consistency.
A calculation that can be verified
Consider an inventory question:
Starting stock: 240 units
Received: 80 units
Shipped: 95 units
Reserved but not shipped: 30 units
Question: how many units are available to promise?
The required calculation is 240 + 80 − 95 − 30 = 195. The distinction between physical stock and availability matters: physical stock is 225, while 30 of those units are already reserved.
A useful final explanation is: “195 units are available: 225 physically present minus 30 reserved.” An application can verify the arithmetic with code and confirm that the quantities came from the relevant inventory snapshot. A long narrative is neither necessary nor sufficient.
Now change the input to “30 reservations may have expired.” Extra generated reasoning does not resolve that missing fact. Query the authoritative reservation state or report the conditional answers. Reasoning over missing evidence is still missing evidence.
Do not mistake an explanation for a faithful trace
A generated rationale can be plausible while omitting factors that influenced the prediction. Turpin et al. demonstrated unfaithful explanations under controlled input biases. The safe conclusion is that visible reasoning text is not a complete audit of the model's internal computation. Unfaithful explanations in CoT prompting.
Some APIs keep internal reasoning private or expose only summaries. Design for a concise answer, relevant evidence, calculations and verifiable artifacts. Do not depend on access to a private reasoning trace or treat its absence as proof that no reasoning occurred.
For consequential actions, the audit record should contain authenticated inputs, evidence versions, proposed arguments, validation results and actual execution outcomes. It should not rely on the model's assertion that it “carefully checked everything.”
Add verification to the workflow
Read diagram source
flowchart LR
I[Task and permitted evidence] --> P[Generate candidate answer]
P --> C[Check calculations or tests]
C -->|Pass defined checks| A[Answer with concise support]
C -->|Fixable error| R[Bounded revision]
R --> C
C -->|Missing evidence or exhausted budget| E[Clarify or escalate]
The revision path needs a maximum number of attempts and a deadline. A checker that merely asks the same model “are you sure?” may repeat the original mistake. Independent tests, source checks and typed constraints provide different evidence, although they too have coverage limits.
Validate candidates before any external side effect. An agent can calculate and revise a proposed inventory update in a sandbox; it cannot assume that retracting a written explanation reverses a committed stock change.
Measure the benefit of more computation
Suppose one solution costs 800 generated tokens and five sampled solutions each cost the same. Candidate generation now costs 4,000 tokens, before aggregation and tool calls. Parallel requests may shorten wall-clock time but do not remove token cost or rate-limit pressure.
If five answers are 195, 195, 195, 225, 225, voting chooses 195. If all five overlook reservations and answer 225, voting preserves the error. Agreement is not a calibrated probability of correctness.
Compare a direct baseline, additional effort, multiple candidates and external verification on the same tasks. Record success, serious errors, latency percentiles, total tokens and tool use. A router that chooses expensive reasoning only for difficult cases has its own false-negative risk: it may misclassify a hard case as easy.
Interview practice
Q1: When would you use CoT prompting?
I would test it on tasks that benefit from intermediate decomposition, such as multi-step calculations or planning. I would compare it with a direct-answer baseline and validate final outcomes. Simple extraction may gain little and become slower or more verbose.
Q2: Does a coherent explanation prove the answer is correct?
No. It may rationalize a mistaken answer or omit an influential bias. I check the evidence and computation independently. In the inventory example, an explanation that ignores reservations is coherent but answers a different question.
Q3: How does self-consistency differ from asking for one longer answer?
Self-consistency samples multiple solution paths and aggregates their final answers. A longer single answer follows one continuation. Multiple candidates may explore alternatives, but correlated errors and aggregation rules remain important.
Q4: Is “think step by step” an appropriate default for every reasoning model?
No universal phrasing is best. I follow the supported interface, state the task clearly, and test the model's effort controls where available. The application should request the useful final explanation without depending on disclosure of hidden reasoning.
Q5: What is a better critique loop than “check your work”?
Give the checker a specific contract: recompute totals, run tests, verify source references or check constraints. Feed back the failed check and allow a bounded revision. Repeated self-approval is weak evidence, and some failures require new information rather than another attempt.
Q6: How would you justify the extra cost in an interview?
Identify the costly failure the additional computation is meant to reduce. Compare measured quality gain against latency and cost per successful task. Establish an escalation policy and show that the router and verifier meet the requirements too.
Final notes
Recall card: Decompose when useful, verify what matters, explain the result concisely. More generated steps can improve some outcomes; their value must be measured.
For explicit search across alternative intermediate states, continue to Tree of Thoughts. For trained reasoning behavior, see RLVR.