Many Hands — Multi-Agent Workflows, Model Choice and Reasoning Effort
When splitting the work across a drafter, a fresh-context critic and a verifier genuinely beats one good prompt — when it is expensive theatre — and how to choose the model and the effort level per task instead of by faith.
The hook
“Ask a model to check its own work and it marks its own homework — generously. Put the same draft in front of a fresh model that has never seen the conversation, briefed as opposing counsel, and watch the holes appear. The difference is not a better model. It is independence — the thing chambers discovered centuries ago, now available in a second browser tab.”
What you'll be able to do
- Split a legal task across roles — drafter, fresh-context critic, verifier — and explain why the critic's independence from the drafting context is what makes it work.
- Say when multi-agent decomposition genuinely beats one good prompt (independent critique, breadth-first research, parallelisable sub-tasks) and when it is overhead — and account for the token cost honestly.
- Choose a model and a reasoning-effort level per task: what thinking modes actually change, what they cost, where extra effort pays, and where it does nothing — including the one comparison where paying is genuinely the answer.
- Treat every automated critique and checklist as input to human verification, never as verification — the White Paper's Guideline 15, taught as engineering rather than compliance.
On the syllabus
- Will it catch its own mistake? — self-review, sycophancy, and the evidence on self-correction
- The pipeline: drafter, fresh-context critic, verifier — and why independence is the active ingredient
- When many hands beat one prompt — and when they are theatre at fifteen times the tokens
- A critique is not verification: Guideline 15 as engineering
- The 2026 model map: families, free-tier routing, thinking modes and effort tiers — dated, and re-checked each cohort
- Choosing per task: a selection table, and the one place where paying is genuinely the answer
In short
Session 5 teaches orchestration and selection — the two decisions left after the prompt, the context and the harness are right. The evidence on self-correction is blunt: models are poor critics of their own reasoning inside the same conversation, and 'are you sure?' mostly triggers sycophancy; but a fresh context, briefed to attack, reliably finds weaknesses — the vendors themselves now build verification as separate, fresh-context agents. So the class builds the pipeline: a drafter writes the moot argument from the Session 4 workspace; a critic in a second tab, with none of the drafter's context, attacks it as opposing counsel; a verifier reduces every authority and factual claim to a checklist addressed to a human with Indian Kanoon. Then the selection half: the 2026 model landscape as a working map — families, free-versus-paid routing, thinking modes and effort tiers, what each is actually for — closing on Compare C, the course's honest case where the free tier drops a limb of a multi-step statutory computation and a high-effort reasoning model holds it: the model ceiling, real at last, and now diagnosable because everything cheaper has been ruled out.
Why it matters for using AI well
Two upgrades that cost nothing: a critic that actually criticises — because it is fresh, briefed and unattached — and model choices made from a table instead of a mood. You will know before you spend whether a paid tier would help your task, and you will be right, which over a career prices this hour rather well.
What you can do on Monday
Before your next moot practice or submission, run your argument through a fresh-context critic briefed as the other side — second provider if you have one — and take its three weakest points into your revision. Then check its counter-authorities on Indian Kanoon before you fear them: critics hallucinate too.
What they leave with
The skill
Run the pipeline when independence or breadth pays — drafter with context, critic without, verifier producing a human work list — and choose model and effort from the task, not the hype.
The insight
The second model's value is not intelligence, it is innocence: it has not been contaminated by the conversation that produced the mistake. And agreement between two confident predictors is still not evidence.
The moment they remember
The critic's first pass. The drafter's argument — polished in the workspace, contract-clean, entirely plausible — goes into a cold tab briefed as opposing counsel, and the critique comes back naming the assumption everyone in the room had stopped seeing, plus an authority against that half the room notes down. The gasp is not that the machine found holes; it is that the drafting model, asked seconds earlier to review its own work, had found none. Independence, demonstrated in ninety seconds, on free accounts.
In this session
- 01
The opening evidence, felt before it is cited: the room bets on whether the model will catch the planted flaw in its own draft when asked to review it. It mostly does not — and the literature says that is the norm, not the anecdote: models struggle to self-correct reasoning without an external signal, and performance can degrade after self-review. Worse, 'are you sure?' is not a check but a nudge — Session 1's sycophancy fold wearing a quality-assurance costume. What does work, per the same literature and now the vendors' own engineering guidance: a verifier with a fresh context, no stake in the draft, and a concrete brief.
- 02
The pipeline, role by role. The drafter works inside the matter's workspace with the full context and the Session 2 contract. The critic gets the opposite: a clean tab — a different provider works even better — no drafting history, and a brief with teeth: 'You are opposing counsel. Identify the three weakest points, the strongest authority against this argument, and every factual assumption it makes without support.' The verifier's product is not a verdict but a work list: every authority, proposition and quotation in the draft, tabulated with where a human should check it. Independence is the active ingredient in all three separations — the same reason your moot partner, not you, moots your own memorial.
- 03
When many hands genuinely win: independent critique (the fresh context is the point); breadth-first research, where sub-tasks are separable and can run in parallel — the pattern behind the 'deep research' modes, which are research agents wearing a product name; and long pipelines where each stage's output is checkable before the next consumes it. When one good prompt wins: single tightly-coupled reasoning tasks, where splitting adds hand-off loss and no independence benefit; and anything where you cannot say what the second agent adds. The cost is real and the vendors say so — an orchestrated multi-agent research run can burn an order of magnitude more tokens than a single chat — so the honest default is: one good prompt first; add hands only for independence or breadth.
- 04
The line the course will not let blur: a critique is not verification. The critic finds weaknesses; the verifier organises claims; neither has opened Indian Kanoon. The Supreme Court's White Paper puts it as Guideline 15 — one generative tool shall not be used to verify or authenticate another's output — and the self-correction literature explains why the rule is sound engineering, not bureaucratic caution: two confident predictors can agree on the same fabrication. Session 1's rule, upgraded: agreement between models is still not evidence. The pipeline's output is a better draft and a to-do list; Session 6 is where the to-do list meets a database.
- 05
The model map, August 2026 — dated on the slide, re-checked each cohort: OpenAI's GPT-5.6 family (the free tier routes to its smallest member, with a limited 'Think' mode); Anthropic's Claude 5 generation — Fable and Mythos at the frontier on paid plans, Opus, Sonnet and Haiku tiers below, the free tier routed among the smaller ones, and the whole service 18+ with phone verification; Google's Gemini 3.x line (free tier on the Flash models with rotating access to Pro); and the open-weight world — DeepSeek, Qwen, Llama — one line each: capable, cheap, and the reason model choice now includes 'run it where the data never leaves'. What the paid ₹399–₹2,000-a-month tiers actually buy, in order of classroom relevance: bigger windows, higher effort ceilings, more tool quota — not a different species of intelligence for reshaping tasks, as Compare A already proved.
- 06
Thinking modes and effort tiers, honestly: a reasoning mode spends more computation before answering — visible as 'thinking' — and measurably helps on multi-step problems: chains of conditions, computations with traps, constraint-satisfaction across documents. It does not consult a database, and on person-fact benchmarks the vendor's own system card recorded a newer reasoning model hallucinating more than its predecessor. Effort tiers (the low-to-max dials on paid plans) buy depth on exactly the tasks where steps get dropped — and latency and cost everywhere else. The free tiers' single thinking toggle is the same idea at its lowest rung; 'deep research' modes are effort plus browsing, rationed on free plans, and their outputs cite more and still fabricate — every deep-research report is a Session 6 work list, not a finished product.
- 07
The selection table students copy out, tasks in rows: reshaping supplied text → any free model, no thinking needed; long single document → the workspace (free) before the big window (paid); multi-limb statutory computation or a chain with traps → a thinking mode, highest effort available, and Compare C shows why; anything that must be cited → grounded mode plus the four-step check, whatever the model; anything confidential → nothing consumer at all. And the standing diagnostic, now complete: prompt ceiling → fix the brief (Session 2); context ceiling → fix the window (Session 3); tool ceiling → fix the harness (Session 4); model ceiling → now, and only now, pay.
- 08
Agentic modes, previewed with appropriate weight: tools that browse, click and act multiply every duty in this course, because the output is no longer a draft you check but an action already taken — Session 8 returns to them. For now, one working rule: the more autonomous the tool, the earlier the human check must move.
The build-along
Build it with the room. Leave holding it.
The class constructs the thing alongside the presenter — the prompt, the workspace, the pipeline — with the mechanism explained as it is built, and a named skill at the end.
Why this shape
One mirror opens it — the room must bet on 'will it catch its own mistake?' and watch self-review fail before independence means anything — and then the session builds: the class constructs the drafter–critic–verifier pipeline with the presenter, role by role, and runs it on a real moot issue. Orchestration is a workflow you assemble, so the hour assembles one.
Will it catch its own mistake?
Experiment on the classOn the class
A short argument with one planted, findable flaw goes on screen. The room bets: asked to review its own draft, will the model that wrote it find the flaw?
In the model
Self-review happens inside the context that produced the error, under training that rewards agreement — the mistake is part of the story the model is completing. A fresh context has no such loyalty.
Live on the model
The drafting chat is asked to review itself: praise, minor edits, flaw intact. The same draft goes to a cold tab briefed as opposing counsel: the flaw is named in the first three lines.
The skill
Never let the drafter mark its own work. Independence — a fresh context with an adversarial brief — is the cheapest quality control in the building.
The pipeline, assembled
Build-alongThe task
The room writes the critic's brief line by line ('opposing counsel… three weakest points… strongest authority against… every unsupported assumption') and the verifier's table columns, voting on what each role must be denied as much as what it is given.
Why it works
Drafter with full context; critic with none of it; verifier reducing the draft to claims a human can check — three prompts, two tabs, one discipline: separate the doing from the doubting.
Built live
The Session 4 workspace drafts the moot argument; the critic (second provider, clean tab) attacks it; the verifier tabulates every authority with a checkbox. All three outputs on screen, in the order a chambers would produce them.
The skill
Drafter, critic, verifier — with the critic kept innocent and the verifier producing a work list for you, not a verdict for the file.
Effort, where it earns its keep
DemonstrationThe room predicts
The room works the limitation problem by hand first — award received, objection filed, vacation intervening — and commits an answer before any model sees it.
What is going on
Multi-limb computations with traps are where reasoning effort pays: each limb dropped is an answer wrong with confidence. Reshaping never needed the effort; this does.
Shown live
Compare C runs: free default versus high-effort reasoning tier on the identical problem, against the room's own answer.
The skill
Match effort to structure: chains and traps get the thinking mode at full effort; reshaping gets none; and when the free tier keeps dropping a limb, that — finally — is the model's ceiling.
Free versus paid · you watch this one
The limitation computation — where paying is the answer
The task, on both: A s.34 Arbitration Act challenge: signed award received on a stated date, court closed for vacation over the deadline, filing on a stated later date — in time, out of time, or condonable? Work every limb: three months, the thirty-day proviso, 'but not thereafter', and the s.4 Limitation Act interaction.
The presenter runs the same prompt on a free-tier account and then on a stronger model or a higher reasoning-effort tier, side by side. Students watch rather than replicate — nobody needs a paid plan to take this course.
Free tier
The free default answers fluently and, across repeated pre-runs, drops a limb — most often the 'but not thereafter' ceiling or the vacation interaction — reaching a confident conclusion the room, having worked it by hand, catches at once.
Stronger model / higher effort
The high-effort reasoning tier walks every limb in order, states the interaction correctly, and shows the computation — matching the room's hand-worked answer, with its working laid out as a checkable structure.
This is the class of task worth paying for: multi-step chains with traps, where dropped limbs are silent. Everything cheaper was ruled out first — the brief was full, the statute was supplied, the harness was irrelevant — which is what an honest 'buy the better model' diagnosis looks like. And even here: the computation gets checked by a human, because confidence still is not evidence.
The legal thread
Guideline 15 of the Supreme Court's White Paper — no generative tool verifies another's output — is this session's spine, taught as sound engineering rather than compliance. And the pipeline doubles your disclosure surface: two providers now hold pieces of your matter, which is why the roles exchange abstracted fact patterns and published law, never client identity. When the tools become agents that act, every one of these duties arrives earlier — Session 8 takes that up.
Hands-on · 30 minutes · on your own laptop
Drafter, Critic, Verifier
Run the pipeline on your own moot issue. Draft in your Session 4 workspace, hand the draft to a cold critic briefed as opposing counsel, and reduce everything to a verification work list — then feel the difference between a machine's critique and your own reading of it. The critique that survives your judgment goes into your revision; the work list goes to Session 6.
1 · Watch — the instructor demonstrates
On
Two free tabs: the Session 4 workspace (NotebookLM or a Project) as drafter; a second provider's free tier (Gemini or ChatGPT, temporary chat) as critic
The exact prompt
Critic brief (paste into the cold tab with the draft below it): You are opposing counsel in an Indian moot. The argument between the markers is my opponent's. Identify: (1) its three weakest points, ranked, with one line on how you would attack each; (2) the strongest authority or line of reasoning against it — if you are not certain a case exists, describe the argument rather than inventing a citation; (3) every factual or legal assumption it makes without support. Do not soften anything. === ARGUMENT BEGINS === [draft] === ARGUMENT ENDS ===
Point at
Point at the ranked weaknesses — then at clause (2)'s hedge working: where the critic describes an argument instead of naming a case, that is the anti-fabrication instruction earning its place. Then ask the drafting workspace to review the same draft and put the two critiques side by side: the cold one has teeth; the warm one has manners.
Roughly what comes back
Three genuine weaknesses (typically one the room had not seen), a described-not-cited counter-argument or a real authority needing verification, and a list of unsupported assumptions running longer than anyone expects. The warm self-review, by contrast, praises structure and suggests a heading.
If it misbehaves
The pre-captured pair of critiques (cold and warm) on the same draft (fallback folder, S5). If the second provider is rate-limited, the critic runs on the same provider in a temporary chat — weaker isolation, same principle, and say so.
2 · Your turn — a variant, not a copy
Your turn, on the other side: draft the opposing argument on the same moot issue (or your own moot's next issue) in your workspace, then run the critic brief in a cold tab against your draft, and build the verifier's table — every authority and proposition, one row each, with a 'where I will check it' column. Choose one critique point to accept and one to reject, and write a line saying why: the judgment between the two is the assessed skill.
Free tier
Five or six prompts split across two providers — inside both free tiers even at class scale, and the split itself halves the rate-limit risk. Under-18 students run both roles on Gemini/ChatGPT (Claude is 18+); if one provider locks out mid-lab, the pipeline collapses gracefully to one provider with temporary chats as the isolation.
3 · The reveal
Phones answer two questions: did the cold critic find something you had missed (the room's yes-rate goes on screen), and did anyone's critic fabricate an authority while attacking? Fabricating critics — there are usually one or two — are shown with relish: even the doubting machine needs the four-step check, which is exactly where the course goes next.
Deliverable
The argument, the cold critique, your accept/reject judgment with reasons, and the verification work list — the critic-pass component of the assessed Grounded Workspace, and the input Session 6's drill consumes.
Run of show · 30 minutes
- 0–10 min — Watch: draft from the workspace, cold critique beside warm self-review, verifier table assembled.
- 10–16 min — Your turn: draft your side's argument in your workspace, contract and all.
- 16–22 min — Critic: cold tab, second provider if you have one, the brief verbatim; read the critique like a lawyer, not a fan.
- 22–27 min — Verifier: build the work-list table; accept one critique point, reject one, with reasons.
- 27–30 min — Reveal: catch-rate and fabricating-critic tallies; one accepted and one rejected critique read aloud with reasons.
For the instructor · before the session
- Pre-run the demo pipeline that morning; capture the cold and warm critiques side by side as the fallback pair.
- Verify the Compare C fact pattern against the statute texts and pre-run it five times on each tier; keep the transcripts — the free tier's limb-dropping is probabilistic and the capture must be honest about that.
- Stage the two demo accounts (workspace + second provider) signed in, temporary-chat mode confirmed.
- Post the critic brief and the verifier table template as copyable text before the session.
- Launch the Session 5 poll deck: the self-review bet, the hand-worked limitation answer, and the two lab tallies.
- Remind students to bring their Session 4 workspace and their moot issue; have three spare issues from the bundles for anyone without one.
Key sources & cases
Huang et al., 'Large Language Models Cannot Self-Correct Reasoning Yet' (ICLR 2024); Kamoi et al., TACL 2024 (survey)
The evidence under the opening bet: models struggle to self-correct without external feedback and can get worse after self-review; the survey finds no demonstrated success from prompted self-correction outside tasks with checkable criteria. Verified to abstracts 2026-08-27.
Anthropic, model documentation and engineering guidance (read 2026-08-27)
The vendor's own current position: 'separate, fresh-context verifier subagents tend to outperform self-critique'; orchestrator-and-workers as the multi-agent pattern; multi-agent research runs consuming many times a single chat's tokens. Vendor claims, cited as such — and the clearest available statement that independence is engineering, not folklore.
Supreme Court of India, White Paper on Artificial Intelligence and Judiciary (Nov 2025), Guidelines 14 and 15
Guideline 14: AI-derived information independently verified before reliance. Guideline 15: no generative tool used to verify or authenticate another's output. Verified 2026-08-26 against the official PDF. The session teaches 15 as the engineering conclusion the self-correction literature independently reached.
The 2026 model landscape (vendor pages, checked 2026-08-27 — RE-CHECK EACH COHORT)
GPT-5.6 family (free tier on its smallest member); Claude 5 generation — Fable/Mythos on paid plans only, free tier among the smaller models, service 18+ with phone verification; Gemini 3.x (free on Flash tiers); open-weight DeepSeek/Qwen/Llama. Every name and routing on this slide carries its checked date; strand 04 of the research file holds the full map.
OpenAI, o3 and o4-mini System Card (2025), §3.3; 'deep research' documentation across vendors
Reasoning models can hallucinate more on person-fact benchmarks (o3 at 0.33 vs o1 at 0.16, the vendor's own table) — thinking is not knowing. Deep-research modes are rationed on free tiers and their outputs remain unverified research memos. Verified to the system-card PDF 2026-08-27.
Arbitration and Conciliation Act, 1996, s.34(3) and its proviso; Limitation Act, 1963, s.4
Compare C's statutory instrument: the three-month limit, the thirty-day condonable window, 'but not thereafter', and the court-holiday interaction — a real multi-limb computation with a trap at every limb. VERIFY TO SOURCE: confirm the section texts and the leading authority on the computation before teaching; the compare's fact pattern is fixed only after that check.
Readings
- Huang et al., 'LLMs Cannot Self-Correct Reasoning Yet' (ICLR 2024) — abstract
- Anthropic, 'How we built our multi-agent research system' (2025) — the sections on when multi-agent pays and what it costs (vendor account; read as one)
- White Paper on Artificial Intelligence and Judiciary (Nov 2025) — Guidelines 12–15, one page
- OpenAI o3/o4-mini System Card §3.3 — the hallucination table
- Arbitration and Conciliation Act 1996, s.34(3) with proviso — read it cold before the session; you will compute against it
Sixteen hours, one professional discipline.
Using AI well is not a knack — it is craft, competence and verification, practised until they are habits you could defend in court.