Context Engineering — What Goes in the Window
The discipline nobody taught you: the model reads only what is in front of it, the free tier's window is small, and choosing, ordering and refreshing what fills it is where most of the quality actually comes from.
The hook
“A model has no memory of Indian law sitting in a drawer — it has a window, and on a free account that window holds perhaps fifty pages. Your judgment is two hundred. Deciding what goes into the window, in what order, with what question at the end, is not a preliminary to the work. It is the work.”
What you'll be able to do
- Explain what the context window is and is not — everything the model can actually read right now, not a memory of your matter — and why a free tier's small window makes excerpt selection a legal skill.
- Supply the document instead of asking from recall, structure long text for the model (document first, question last, tagged sections), and demand paragraph pinpoints so every claim in the output is checkable against the source.
- Manage what does not fit: choose between excerpting, chunking with a running structure, and re-supplying the operative passage — and say what each choice loses.
- Practise context hygiene: one task per chat, fresh context after a wrong turn, memory off for clean experiments, and a long conversation treated as a degrading asset rather than an accumulating one.
On the syllabus
- The window: what the model can actually read right now — and how small it is on a free account
- Recall mode versus grounded mode: the same app, opposite reliability — supplying beats asking
- Structuring long text: document first, question last, tagged sections, pinpoints demanded
- What gets lost: the middle of long contexts, the thread of long chats, the dissent in every summary
- Excerpt, chunk, or re-supply: managing a judgment that does not fit, and what each choice loses
- Context hygiene: one task per chat, fresh windows after wrong turns, memory off for clean runs
In short
Session 3 teaches the discipline students will feel most and have been taught least. The window is everything: a model answering about a judgment it was given is doing a reading task on its reliable ground; a model answering from recall is composing. But the window is finite — small on free tiers — and it degrades: material buried mid-context gets missed, long conversations decay, and a wrong turn early poisons everything after it. So the session teaches the craft that follows: pick the operative text like a lawyer, put the document first and the question last, tag the parts, demand paragraph pinpoints, verify a sample of them, and when the thread goes wrong, fold what you learned into a better first prompt in a clean window. The long-judgment problem gets the course's second free-versus-paid comparison — a genuine model ceiling that technique mostly closes — and the legal-research craft the course has always taught, reading the summary against the judgment for the flattened dissent, lives here as the lab.
Why it matters for using AI well
Speed without the flattening. You get the model's pace on long documents — the brief, the issue table, the comparison — while the pinpoint discipline and the read-back keep you the one person in the room who knows what the summary dropped. And you stop losing hours to degraded chats: the fresh-window habit alone repays the session.
What you can do on Monday
Next time you need a judgment for class or a moot, paste its operative text and build the brief with staged prompts and pinpoints — and open two of the pinpoints before you rely on any of it. Never again ask what a case held when you are holding the case.
What they leave with
The skill
Supply, structure, pinpoint: document first, question last, parts labelled, paragraph numbers demanded, a sample of them verified — and a fresh window the moment a thread degrades.
The insight
The model's answer is only ever as good as the window it read — and the window is yours to engineer. Most of what people call a model problem is a context problem wearing the model's clothes.
The moment they remember
The planted-fact demonstration. A long extract goes into the window with the case-deciding line buried mid-document; asked cold, the model summarises fluently — and misses it. The room, who watched the line go in, protests. Then the presenter re-supplies just that passage above the same question and the model nails it instantly. The machine did not get smarter in ten seconds; the window got better — and the room never again believes that pasting a document means the model has read it the way they would.
In this session
- 01
The window, mechanically: at every turn the model reads one finite sequence — your conversation so far, plus whatever you pasted — and predicts from that. It has no other access to your matter. Everything Session 1 said about recall being composition follows; the corollary is that the single highest-leverage act in AI-assisted legal work is deciding what fills the window. As of August 2026 the free tiers hold roughly 27,000–32,000 tokens — call it fifty pages, minus room for the answer — while paid frontier tiers advertise up to a million. Re-checked each cohort; the numbers move, the principle does not.
- 02
Recall mode versus grounded mode — the reframe that replaces the old 'generative versus grounded tools' line, because the same free app now does both: ask what Anvar held and the model composes from training-data residue; paste Anvar's operative paragraphs and ask, and it reshapes text in front of it. Grounding is not a different product, it is a different practice — and it is measured: even the best models, given the document and told to rely on it, still produce some ungrounded material (roughly one answer in six on Google's FACTS benchmark), which is why grounding changes the failure mode from invention to misreading rather than abolishing the duty to read.
- 03
Structure for the machine: the long material first, the question last — every vendor's current guidance, with one claiming up to 30% better answers on multi-document tasks from ordering alone; label the parts (JUDGMENT / STATUTE / MY QUESTION) so instructions never blur into evidence; ask for word-for-word quotations before analysis; and demand paragraph pinpoints for every proposition. A pinpoint the model gives you is a claim; a pinpoint you open is evidence — the Ladder, applied to your own workflow. One Indian reality check: a scanned PDF has no text layer, so nothing can be quoted or pinpointed from it — get the text version, or you are grounding on fog.
- 04
What the window loses even when text fits: material buried mid-context is found less reliably than material at the edges — the 'lost in the middle' effect, flattened but not abolished on current models — and unrelated material in the window actively distracts. What long chats lose is worse: across 15 models, performance dropped by an average of 39% when the same task arrived through a meandering multi-turn conversation instead of one clean brief, and models that took a wrong turn did not recover. The professional habits follow directly: one task per chat; the operative passage re-supplied right before the question that turns on it; and after two failed corrections, a fresh window with a better first prompt — never a fifth 'no, I meant…'.
- 05
When the judgment does not fit, three honest options: excerpt — choose the operative paragraphs yourself, which is issue-spotting and the best option when you know where the matter lives; chunk — feed the judgment in parts against a fixed structure (issues → holding per issue → what was obiter → only then a summary), carrying a running note forward, accepting that boundary-spanning reasoning can fall through the cracks; re-supply — keep the full text elsewhere and paste back only the passage each question turns on. Rolling summaries are for continuity, never for evidence: compression loses detail by design, and what it loses first is the qualification, the dissent, the distinguishing fact — precisely what your matter turns on.
- 06
The flattening problem, taught as legal-research craft: a summary requested cold compresses toward 'the court held', collapsing concurrences and dissents, dropping per incuriam findings and the narrow fact that made the ruling small. A summary built on a structure you supplied — issues, holdings, obiter, with pinpoints — can be checked against that structure, paragraph by paragraph. The comparison discipline: the AI summary is a hypothesis; the judgment is the evidence; read the operative paragraphs against it and log every divergence, which is also, quietly, how you become a faster reader of judgments.
- 07
Side-by-side work — two clauses, two statutory regimes, a section before and after amendment — is reshaping and therefore strong ground, with one standing blind spot: the model compares only what is in the window. The defined term three pages away, the proviso in another instrument, the amendment after its cutoff — supply them or the comparison silently proceeds without them. The reviewing habit: after any comparison, ask what is not on the page.
- 08
Hygiene, as a checklist the lab enforces: memory and personalisation off — or a temporary chat — for any clean run, because free tiers now carry memory across conversations by default and your last matter can leak into this one; one matter per chat, always; label your chats like files; and treat the window as a courtroom record — everything in it is before the judge, nothing outside it exists.
The build-along
Build it with the room. Leave holding it.
The class constructs the thing alongside the presenter — the prompt, the workspace, the pipeline — with the mechanism explained as it is built, and a named skill at the end.
Why this shape
A build-along with one opening mirror. The recall-versus-reading poll needs the room to bet wrong before the point lands, so that beat stays an experiment; everything after it is craft — the class assembles a checkable judgment brief with the presenter, stage by stage, and leaves knowing how to do it to any judgment. The skill is construction, so the session is construction.
Recall versus reading — the opening bet
Experiment on the classOn the class
The room commits on their phones: which will be more accurate — asking the model what Arjun Panditrao held, or pasting its operative paragraphs and asking the same question? And by how much?
In the model
Asked cold, the model composes from training residue — usually the gist, with the primary/secondary distinction blurred or the certificate rule overstated. Given the text, it is reshaping, with paragraph support on demand.
Live on the model
Both run, side by side. The cold answer is graded by the room against the pasted-text answer's pinpoints — and the errors in the cold version are exactly the subtle kind (a condition dropped, a distinction blurred) that survive a careless read.
The skill
When you have the source, supply it. Turn every recall question you can into a reading question — and grade recall answers as drafts, never as law.
The judgment brief, built in stages
Build-alongThe task
The room drives the staged sequence on a mid-length judgment from the excerpt set: what issues did the court frame → the holding on each, with paragraph numbers → what was obiter → only now, a summary. Each stage's prompt is dictated from the floor before it runs.
Why it works
A summary built on a structure you supplied can be checked against that structure; a summary requested cold can be checked against nothing. Pinpoints convert the output from prose into a map of the judgment.
Built live
The staged brief assembles on screen; then the presenter opens two cited paragraphs live — one checks out, and one, usually, is subtly off, which is the moment the read-back habit is born.
The skill
Structure first, pinpoints always, and read the summary against the judgment — never instead of it.
The planted fact — lost in the middle, live
DemonstrationThe room predicts
The room watches the case-deciding line go into the middle of a long paste, and predicts: will the model's summary surface it?
What is going on
Attention over a long window is not uniform: mid-context material is found less reliably, and the effect is probabilistic — which is itself worth seeing, so the demonstration is run twice.
Shown live
Asked cold, the summary misses or mangles the planted line (if it doesn't, the pre-run capture shows the miss — and the room learns the effect is a probability, not a law). Re-supplying the passage above the question fixes it instantly.
The skill
Never assume pasted means read. Re-supply the operative passage right before the question that turns on it.
Free versus paid · you watch this one
The two-hundred-page judgment — where the window is real
The task, on both: Build an issues-and-holdings table, with paragraph pinpoints, for a very long multi-opinion constitution-bench judgment — the kind with concurrences that decide everything.
The presenter runs the same prompt on a free-tier account and then on a stronger model or a higher reasoning-effort tier, side by side. Students watch rather than replicate — nobody needs a paid plan to take this course.
Free tier
The free tier cannot take the full text: the window is ~27–32K tokens and the upload path is capped. Chunked, staged prompting — the Session 3 technique — recovers a serviceable table in four passes, at the cost of effort and some cross-opinion synthesis at the seams.
Stronger model / higher effort
The presenter's paid large-window model ingests the entire judgment in one pass and produces the table with pinpoints, including the cross-opinion tally the chunked run had to assemble by hand.
The window is a real ceiling you can pay to lift — and a technique ceiling you can engineer past for free, most of the way. Diagnose before you spend: if the task fits the window and the output is poor, it was never the model. And hold the thought until Session 4: there is a free tool built exactly for this problem.
The legal thread
Everything you put in the window is a disclosure to a third party, and the free tiers train on conversations by default — so the rule that makes this whole session professionally safe is: judgments, statutes and published material go in; clients never do. The opt-out switches, the workspace exceptions and the statutory frame arrive in Session 4; the privilege doctrine in Session 7. For now: if you would not read it aloud in a crowded café, it does not go in the window.
Hands-on · 30 minutes · on your own laptop
The Judgment in the Window
Do to a judgment what the build-along just did on screen: supply the operative text, build the brief in stages with paragraph pinpoints, then read the model's summary back against the judgment and log every divergence — the flattened concurrence, the dropped condition, the overstated holding. The deliverable is a brief you could actually take to a moot practice, plus the divergence log that proves you checked it.
1 · Watch — the instructor demonstrates
On
Gemini or ChatGPT (free tier, temporary chat); the judgment as a text-layer excerpt from the session set
The exact prompt
The text between the markers is an extract from a judgment of the Supreme Court of India. Using only this text: (1) list the issues the court framed; (2) for each issue, the holding, with the paragraph number(s); (3) anything said obiter, with paragraph numbers; (4) only after all that, a ten-line summary. Quote the single operative sentence of the judgment word for word, with its paragraph number. If something is not addressed in this extract, say 'not in the supplied text' rather than answering from memory. === JUDGMENT BEGINS === [pasted extract] === JUDGMENT ENDS ===
Point at
Point at the pinpoints as they appear — then open two of them in the actual PDF, in front of the room. One will check out; read it aloud. Wherever one is off — a paragraph number pointing at the wrong passage, a holding stated wider than the paragraph supports — that divergence is the lesson, not an embarrassment.
Roughly what comes back
A clean staged brief with mostly-accurate pinpoints and, typically, one soft spot: a blurred distinction between holding and obiter, or a summary line that outruns its paragraph. On a judgment with a concurrence, expect the summary to flatten it into 'the court held'.
If it misbehaves
The pre-captured staged run on the same extract (fallback folder, S3). If the free tier truncates the paste, use the shorter alternate extract marked 'S3-short' — the set includes one sized for the smallest window.
2 · Your turn — a variant, not a copy
Your turn, on a different judgment: pick one of the five in the session set (each from a different practice area, each sized for a free window) — or your own moot or seminar judgment if you have its text layer. Run the staged sequence, then open the operative paragraphs and read them against the model's brief. Mark every divergence and classify it: flattened dissent or concurrence, lost qualification, dropped distinguishing fact, overstated holding, wrong pinpoint.
Free tier
Four to six prompts on one pasted extract — inside every free window because the set's extracts are sized for it (that constraint is itself the lesson). If your tier rate-limits, complete the read-back on paper against the transcript so far: the checking is the assessed skill, and it needs no tokens.
3 · The reveal
Phones tally two numbers: how many of your two checked pinpoints survived, and which divergence class you caught. The class distribution goes on screen — flattening is never zero — and one student's best catch is read aloud with the paragraph on screen beside the summary that lost it.
Deliverable
A staged judgment brief with pinpoints, plus a divergence log with each divergence classified — the second entry in the Prompt & Context Portfolio, and the artefact you will rebuild inside a workspace in Session 4.
Run of show · 30 minutes
- 0–10 min — Watch: the staged brief built live, two pinpoints opened against the PDF, one divergence found and classified.
- 10–13 min — Your turn: choose your judgment, open a temporary chat, paste the operative extract with the markers.
- 13–21 min — Run the staged sequence: issues → holdings with pinpoints → obiter → summary.
- 21–27 min — Read back: open the operative paragraphs, check two pinpoints, log and classify every divergence.
- 27–30 min — Reveal: pinpoint survival and divergence classes tallied on screen; the best catch of the day shown against its paragraph.
For the instructor · before the session
- Prepare the five-judgment extract set: text-layer PDFs, operative extracts pre-cut to fit the smallest free window, one 'S3-short' alternate, direct Indian Kanoon URLs for each.
- Pre-run the planted-fact demonstration at least twice that morning and capture a run where the model misses — the effect is probabilistic and the capture is the fallback.
- Stage Compare B: the full constitution-bench PDF loaded in the paid large-window account; the four chunked prompts for the free tier written out and tested.
- Confirm temporary-chat mode and memory-off on the demo accounts.
- Launch the Session 3 poll deck; the recall-versus-reading bet opens the hour.
- Print or post the divergence-class list (flattened dissent / lost qualification / dropped fact / overstated holding / wrong pinpoint) — the lab logs against it.
Key sources & cases
Liu et al., 'Lost in the Middle: How Language Models Use Long Contexts' (TACL 2024), with 2025 follow-ups (NoLiMa, ICML 2025; Chroma, 'Context Rot', 2025)
The evidence that position and length matter: performance highest at the edges of the context, degrading in the middle; degradation with length on all 18 models Chroma tested; distractors compound. The 2026 frontier generation is claimed by vendors to hold up across the window — treat that as a vendor claim; the ordering advice still presupposes the effect. Verified to abstracts and the Chroma report 2026-08-27.
Laban et al., 'LLMs Get Lost in Multi-Turn Conversation' (2025)
The 39% average multi-turn drop and 'models that take a wrong turn do not recover' — the empirical backbone of the fresh-window habit. Verified to the abstract 2026-08-27.
Google DeepMind, FACTS Grounding benchmark (2024–)
With the document supplied and grounding instructed, the best model at launch still produced ungrounded material in roughly one answer in six. The number behind 'grounding changes the failure mode; it does not abolish the read-back'. Re-pull the current leaderboard figure each cohort and cite its date.
Vendor long-context guidance: Anthropic, OpenAI and Google documentation (read 2026-08-27)
Concordant current guidance: long material first, question last (Anthropic claiming 'up to 30%' improvement — a vendor-internal figure, quote as such); delimit sections; quote before analysing; Google warns accuracy drops with multiple 'needles'. Links in the codebook; pages moved hosts in 2026 — re-link each cohort.
Free-tier context windows, as checked 2026-08-27
ChatGPT Free ≈27K tokens; Gemini free tier 32K (vendor pages, one read via text proxy); Claude Free unstated by the vendor; paid tiers up to 1M. These numbers are the most volatile facts in the course — RE-CHECK EACH COHORT and re-state the checked date on the slide.
Anvar P.V. v. P.K. Basheer (2014) 10 SCC 473; Arjun Panditrao Khotkar v. Kailash Kushanrao Gorantyal (2020) 7 SCC 1
The session's worked judgments for the recall-versus-reading opener and the staged brief — chosen because Session 4 builds its workspace on the same electronic-evidence line, so the doctrine compounds across sessions. Both verified 2026-08-26 as recorded in the codebook.
Readings
- Liu et al., 'Lost in the Middle' (TACL 2024) — abstract and Figure 1
- Chroma, 'Context Rot' (2025) — the summary page; skim the charts
- Laban et al., 'LLMs Get Lost in Multi-Turn Conversation' (2025) — abstract
- Arjun Panditrao Khotkar v. Kailash Kushanrao Gorantyal (2020) 7 SCC 1 — the operative paragraphs, before the session; you will watch a model read them
- Your worst prompt from Session 2's homework — rewritten as a document-first, question-last brief
Sixteen hours, one professional discipline.
Using AI well is not a knack — it is craft, competence and verification, practised until they are habits you could defend in court.