Skip to content
Session 02 / 08· Direct it precisely

Prompting Like a Lawyer — Specification, Worked Examples and the Output Contract

You already know how to brief a junior. This hour turns that skill on the machine: a specification, an output contract that forces honesty, worked examples for form — and a diagnostic for whose ceiling you have hit when the answer disappoints.

90 minutes60 taught + 30 hands-onDecode the machine, direct itBuild-along Free vs. paid comparisonYour laptop · free tier

The hook

'Summarise this judgment' is not an instruction; it is a hope. The difference between that and a specification with an output contract is the difference between a corridor shout and a brief a junior can execute — and unlike the junior, the machine will never come back to ask what you meant.

What you'll be able to do

  • Write a specification for a legal task — task, audience and register, context, constraints, format — that says everything that matters exactly once, and nothing else.
  • Attach an output contract that forces the model to show its uncertainty: an authorities table with a 'checked?' column, mandatory flags on anything unverified, and standing permission to say 'I am not certain'.
  • Steer form with worked examples — two or three exemplars where the output must match a house style — and know when examples help, when they constrain, and where the vendors now disagree.
  • Handle refusals and hedges professionally, iterate by changing one thing at a time, know when to abandon a chat and start clean — and diagnose whether a disappointing answer is your prompt's ceiling or the model's.

On the syllabus

  • The specification: task, audience and register, context, constraints, format — everything that matters, once
  • The output contract: authorities tables with a 'checked?' column, mandatory uncertainty flags, permission to abstain
  • Worked examples: steering form with two or three exemplars — and where the vendors now disagree about examples
  • Refusals and hedges: reframing, and making 'it depends' earn its keep
  • Deliberate iteration: change one thing, keep a log, and start fresh when the chat is contaminated
  • The folklore autopsy, and the diagnostic: is this the model's ceiling, or my prompt's?

In short

Session 2 is the first craft session, and it teaches prompt engineering as what it actually is in 2026: briefing, not incantation. The room builds a real internship research-note prompt live, from a one-line ask to a full specification with an output contract, watching the output improve at each step and voting on what to add next. Along the way the session buries the folklore with the evidence that killed it — tips, threats and politeness do nothing measurable; 'think step by step' adds little on models that already think; a persona changes register, not accuracy — and replaces it with what holds: specify once and clearly, show the shape you want with examples, force the uncertainty into the open, restate or restart rather than argue with a derailed chat. The hour closes on the course's running diagnostic, proven by the first free-versus-paid comparison: on a well-specified reshaping task, paying changes nothing, because the ceiling you usually hit is your own prompt's.

Why it matters for using AI well

A reusable template you will actually use this term: a specification plus an output contract that turns any research or drafting ask into checkable work with its uncertainty on the surface. And a habit that saves money and blame in equal measure: before you curse the model or buy a better one, fix the brief.

What you can do on Monday

Take the next real ask you get — a moot issue, a project topic, an internship request — and run it through your saved specification-plus-contract template instead of typing a one-liner. File the output with its 'checked?' column as your work list.

What they leave with

The skill

Brief the machine like a junior: a specification that says everything once, an output contract that forces the uncertainty into the open, and exemplars where the form must match.

The insight

Prompt quality is not a property of magic words; it is a property of how clearly you had already thought about the task — and the contract matters more than the phrasing, because it makes the model's ignorance visible instead of fluent.

The moment they remember

The build-along's third round. The one-line ask produced confident mush; the specification produced structure; and then the output contract lands and the same model, on the same question, suddenly confesses — a table with 'checked?: no' against every authority and a flagged 'I am not certain' where it had previously been smoothest. The room audibly reacts to seeing the confidence stripped off. Students arrive believing prompting is about coaxing better answers and leave having watched it force honest ones.

In this session

  • 01

    The specification, field by field: the task stated as a deliverable; the audience and register ('for a supervising advocate who has not read the file') — which is what a 'role' is actually for: the evidence says personas change tone and focus, not factual accuracy, so cast the role as who the output is for, never as a spell for correctness; the context the model cannot know; the constraints (jurisdiction, date, length, what to exclude); the format, precisely. Then the discipline the 2026 guidance actually stresses: say each thing once and stop — every vendor now warns that bloated, repetitive, shouty prompts degrade output, and the working test is the colleague test: if a colleague with no context would be confused by your prompt, so will the model.

  • 02

    The output contract is the lawyer's move, and the heart of this session: require an authorities table — authority | proposition | paragraph | checked? — with 'checked?' set to no on everything, because the model cannot check; require anything uncertain to be flagged rather than smoothed over; grant standing permission to say 'I am not certain', which vendors report sharply reduces invented detail and which Session 1's mechanism explains: you are changing the scoring the model was trained to optimise. The contract does not make the model honest — it makes its dishonesty visible, and converts fluent output into a work list.

  • 03

    Worked examples: for anything that must match a shape — a clause in the firm's style, a case comment in a journal's format, an issue list the way your supervisor writes them — two or three exemplars beat any description of the shape. Taught with the 2026 caveat: Anthropic recommends a handful of diverse examples, Google says always include them, OpenAI says reasoning models often do better without and examples that conflict with your instructions actively mislead. The rule that survives: use exemplars to steer form; never expect them to supply accuracy; and if the output is slavishly copying the wrong thing, the examples are why.

  • 04

    Refusals and hedges: 'I can't provide legal advice' usually means the ask sounded like a client matter — reframe as an academic fact pattern and it proceeds. The deeper skill is hedge management: an 'it depends' answer is often the correct legal answer wearing a lazy face, so make it earn its keep — demand the best view, the strongest contrary view, and the facts that would flip the answer. And do not bully: the current guidance is explicit that aggressive, capitalised, threatening prompt language now degrades compliance rather than improving it.

  • 05

    Deliberate iteration: change one variable per run and log what changed, or you have learned nothing from the improvement. And know when to stop iterating: the multi-turn evidence is that models get lost in long corrective conversations and rarely recover — an average 39% drop from a clean single ask across 15 models in one 2025 study. The professional move after two failed corrections is not a third correction; it is a fresh chat with everything you have learned folded into a better first prompt.

  • 06

    The folklore autopsy, with the receipts: offering a tip, issuing a threat, and adding 'please' have no reliable aggregate effect (Wharton Prompting Science Reports, 2025); 'think step by step' produces marginal-to-no accuracy gain on models that already reason internally, at real cost in time — and today's free tiers all run thinking models, so the projector demo of that folklore shows, honestly, nothing. What replaced it: ask for a structured argument you can check — IRAC with paragraph pinpoints into supplied text — rather than a narration of the model's hidden reasoning, which research shows is not a faithful transcript anyway.

  • 07

    The diagnostic that runs through the rest of the course: when output disappoints, is it the model's ceiling or your prompt's? The order of operations is fixed — specification first, context next (Session 3), the harness after that (Session 4), and only then the model (Session 5). Today's comparison shows why the order starts there: the same well-specified reshaping task, run on the free tier and on a paid frontier model, comes back equivalent. Almost everything students blame on the model is a prompt or context problem — and the course will show the two places it genuinely is not.

  • 08

    The computational-thinking lens, named: a specification is abstraction — stripping the request to what matters, the same move as extracting a ratio; a staged sequence is decomposition — issue-spotting applied to your own ask. The pillars are not a syllabus in this course; they are the reason a legal education is a head start at this craft.

The build-along

Build it with the room. Leave holding it.

The class constructs the thing alongside the presenter — the prompt, the workspace, the pipeline — with the mechanism explained as it is built, and a named skill at the end.

Why this shape

Craft is learned by building, not by being fooled. The room constructs the internship-note prompt with the presenter, voting on each addition before it runs and watching the output change — and leaves holding a reusable template. An experiment here would only prove students cannot prompt yet, which nobody needs proven.

The internship note, built in four rounds

Build-along

The task

The task: a supervising advocate wants a note on a live question by tomorrow. Round one is the ask most of the room would type today — one line — and its output goes on screen. Before each further round the room votes on which addition will help most, then dictates it: specification, then context and constraints, then the output contract.

Why it works

Each addition changes the output for a reason the mechanism explains: the specification narrows the distribution; the context replaces the statistically average assumption with your facts; the contract changes what the model is optimising toward — visible uncertainty instead of smooth confidence.

Built live

Four runs, side by side on the projector, one addition at a time. The fourth output carries an authorities table with 'checked?: no' down the column — the first honest draft of the day.

The skill

Build the brief before you type it: task, audience, context, constraints, format — then the contract that forces the doubt into the open.

The folklore autopsy

Demonstration

The room predicts

The room votes on which of four prompt 'boosters' will improve a fixed answer: a ₹500 tip, a threat, 'please', and 'think step by step'. Most rooms bet on at least two.

What is going on

The 2025 evidence: no reliable aggregate effect from tips, threats or politeness; marginal-to-nothing from CoT incantations on models that already reason before answering — which is what every 2026 free tier now runs.

Shown live

The same question runs plain, tipped, threatened and step-by-stepped, live. The answers are equivalent — and that null result, predicted on screen before the run, is the finding.

The skill

Spend your effort on the specification and the contract, not on incantations. When a technique claims magic, ask for its evidence and its date.

Making 'it depends' earn its keep

Demonstration

The room predicts

A genuinely contestable question goes up, and the room drafts the follow-up they would send when the model hedges.

What is going on

Hedging is trained-in safety, but it yields to structure: a contract demanding the best view, the strongest contrary view, and the facts that would flip the answer converts a shrug into legal analysis — without bullying, which current models are documented to punish rather than reward.

Shown live

The hedge appears on cue; the structured follow-up runs; the answer comes back as two argued positions with conditions — the shape a supervisor actually wants.

The skill

Never accept mush and never bully. Demand both sides and the switching conditions — that is what 'it depends' owes you.

Free versus paid · you watch this one

The clause rewrite — where paying changes nothing

The task, on both: Rewrite a dense indemnity clause in plain English for a client, to a tight specification with two exemplar rewrites supplied — a pure reshaping task with a full brief.

The presenter runs the same prompt on a free-tier account and then on a stronger model or a higher reasoning-effort tier, side by side. Students watch rather than replicate — nobody needs a paid plan to take this course.

Free tier

The free tier, given the full specification and exemplars, returns a clean, accurate, correctly-registered rewrite. Nothing material to fix.

Stronger model / higher effort

The presenter's paid frontier model, given the identical prompt, returns an equivalent rewrite — differently worded, not better. The room compares them blind and cannot reliably pick the paid one.

Paying changes nothing

A well-specified reshaping task saturates the free tier. When output on a task like this disappoints, you have hit your prompt's ceiling, not the model's — fix the brief before you reach for a wallet. The course will show you, in Sessions 3 and 5, the two places where the ceiling genuinely is the model's.

The legal thread

The output contract is your first professional artefact: a table that says 'checked?: no' against every authority is the honest starting state of every AI-assisted document, and converting those entries to 'yes' by hand is what candour to the tribunal will demand before your name goes on anything (Session 6 drills the check; Session 8 the certification). Notice also what the built prompt never contained: a client's name. Abstraction before prompting starts here and becomes doctrine in Sessions 4 and 7.

Hands-on · 30 minutes · on your own laptop

Technique: Specification + output contract

The Internship Note Prompt

Build the prompt you will reuse all term. Watch the instructor assemble a full specification and output contract for a realistic internship research note; then build the same machinery around a task of your own — your moot issue, your project topic, your last internship ask — run it on a free tier, and file the output with its uncertainty flags as a work list. The prompt itself, saved, is the deliverable that matters.

1 · Watch — the instructor demonstrates

On

Gemini or ChatGPT (free tier)

The exact prompt

You are drafting for a supervising advocate who has not read the file. Task: a research note on whether a licensor can forfeit the full security deposit when a licensee exits a leave-and-licence agreement before the lock-in period. Context: Maharashtra; commercial premises; no negotiated liquidated-damages clause. Constraints: Indian law only; note any point where the position is unsettled; do not invent authority. Format: (1) issue in one sentence; (2) short answer in three; (3) analysis under headings; (4) a table — authority | proposition | paragraph | checked? — with 'checked?' as 'no' for every row, because you cannot verify; (5) a section titled 'What I am not certain of', which must not be empty. If you are not certain a case exists, say so instead of naming one.

Point at

Point at three things when it returns: the table's 'checked?: no' column (the honest starting state of every AI draft); the populated 'not certain' section (permission to abstain, working); and any authority in the table — which the room should already distrust on sight, after Session 1.

Roughly what comes back

A well-structured note with two to four authorities of mixed reality in the table and a genuinely useful uncertainty section. The substance will be plausible and unverified — say so, and leave it unverified: verification is Session 6's drill, and the flags are today's point.

If it misbehaves

The pre-captured note from the same prompt (fallback folder, S2). If the live model refuses the framing as legal advice, prepend 'For a law-school exercise:' — and show the room that reframing move as content, not as a trick.

2 · Your turn — a variant, not a copy

Your turn, on your own task — the moot issue you are actually arguing, the seminar topic you are actually writing, the kind of note your internship actually asked for. Build the specification field by field, attach the contract verbatim (adapt the table columns if your task is drafting rather than research), run it, and save the prompt — not just the output — to a personal prompt bank you will keep all course.

Free tier

One long prompt and one or two refinement runs — trivially inside every free tier's limits, text-only. If you are rate-limited, draft the full prompt offline in your notes app and run it once when the window resets; the drafting, not the running, is the skill being assessed.

3 · The reveal

Two volunteered contracts go on the projector. The room reads each output's 'What I am not certain of' section aloud and votes: which prompt forced more honesty out of the same machine? The best contract clause found in the room is added to everyone's template on the spot.

Deliverable

A saved, reusable specification-plus-contract prompt for a real task of yours, its first output with every unverified claim flagged, and one line on what you changed after seeing the output. First entry in the Prompt & Context Portfolio.

Run of show · 30 minutes

  1. 0–10 min — Watch: the instructor builds the internship-note prompt from one line to full specification and contract, running it at each stage.
  2. 10–14 min — Your turn: choose your real task and draft the specification fields.
  3. 14–22 min — Attach the output contract, run it on your free tier, and read the output against the contract: did every section it promised appear?
  4. 22–26 min — Save the prompt to your prompt bank; mark in the margin every claim you would have to verify before this went anywhere near real work.
  5. 26–30 min — Reveal: two contracts compared on screen; the room's best clause is adopted into the shared template.
For the instructor · before the session
  • Pre-run the four build-along stages that morning and screenshot each output — the improvement arc is the lesson, and models drift.
  • Load the fallback folder: the four staged outputs, the folklore-autopsy runs, and the pre-captured demo note.
  • Have the paid-tier account signed in for Compare A, with the identical clause-rewrite prompt staged in both tiers.
  • Launch the Session 2 poll deck; the build-along votes and the folklore predictions run on phones.
  • Post the shared prompt-bank template (a plain document with fields) to the class channel so nobody builds theirs from a blank page.
  • Keep the Session 1 excerpt set handy — students without a live task of their own borrow a question from it for the practice segment.

Key sources & cases

  • OpenAI, GPT-5-era prompting and reasoning guidance (developers.openai.com, read 2026-08-27)

    The vendor's current positions, cited to the live pages: prompting reasoning models to 'think step by step' is unnecessary; try prompts without examples first on reasoning models; leaner prompts outperformed (vendor-internal figures — quote as vendor claims); contradictory instructions damage the newest models most. Re-check each cohort; these pages moved hosts in 2026 and will move again.

  • Anthropic, prompting best practices and 'reduce hallucinations' documentation (platform.claude.com, read 2026-08-27)

    Worked examples (3–5, diverse) for format and tone; permission to say 'I don't know' as a first-line hallucination reduction; 'dial back' aggressive prompt language on current models; quotes-first grounding for long documents. Vendor guidance, not measured findings — labelled as such in class.

  • Meincke, Mollick, Mollick & Shapiro, Prompting Science Reports 1–3 (Wharton, 2025)

    The folklore autopsy's evidence base: politeness effects are contingent and unpredictable (Report 1); chain-of-thought prompting yields marginal-to-no gains on reasoning models at 20–80% time cost (Report 2); tipping and threatening have no significant aggregate effect (Report 3). Verified to the arXiv abstracts 2026-08-27.

  • Zheng et al., 'Personas in System Prompts Do Not Improve Performances of LLMs' (EMNLP Findings 2024)

    162 personas, four model families, 2,410 factual questions: no accuracy improvement. The verified basis for teaching 'role' as audience-and-register, not as an accuracy lever.

  • Laban et al., 'LLMs Get Lost in Multi-Turn Conversation' (2025)

    An average 39% drop from single-turn to multi-turn performance across 15 models; models that take a wrong turn 'do not recover'. The evidence behind 'after two corrections, start a fresh chat with a better first prompt'. Verified to the abstract 2026-08-27.

  • Chen et al. (Anthropic), 'Reasoning Models Don't Always Say What They Think' (2025)

    Visible chains of thought verbalised the cue the model actually used in often under 20% of cases. Why the course teaches 'demand a checkable argument with pinpoints', not 'read the model's mind'.

  • Jeannette Wing, 'Computational Thinking', CACM 49(3) (2006)

    Decomposition and abstraction as general habits of mind — the lens, retained from the course's first design, for why legal training transfers to this craft.

Readings

  • Anthropic, 'Prompting best practices' — skim the sections on examples, roles and reducing hallucinations (link in the class channel; read the current page, not a cached one)
  • OpenAI, 'Reasoning best practices' — the two paragraphs on step-by-step prompting and few-shot examples
  • Wharton Prompting Science Reports 1–3 (2025) — abstracts only: the evidence that buried the folklore
  • Laban et al., 'LLMs Get Lost in Multi-Turn Conversation' (2025) — abstract
  • Your own last five real prompts, re-read against today's template — bring the worst one to Session 3

Sixteen hours, one professional discipline.

Using AI well is not a knack — it is craft, competence and verification, practised until they are habits you could defend in court.