Grounding the Machine — Retrieval, Tools and the Authority Check
The difference between a model that composes an authority and one that fetches it — and the four-step check that turns either into something you can rely on.
The hook
“A generic chatbot and a grounded research engine can answer the same question with the same confidence, but only one of them can show you where the answer came from. In practice, the distance between those two is the distance between competent lawyering and a cost order.”
What you'll be able to do
- Explain what retrieval-grounding actually does, and why it changes the failure mode from invention to misreading without eliminating the need to check.
- Apply the strongest single move in legal AI: supply the document yourself and ask questions about it.
- Run the four-step authority check as a drilled habit — identify, verify it exists, confirm it says that, check it is still good law.
- Choose a tool on grounding, provenance and verifiability rather than on fluency, and say where AI genuinely accelerates legal work.
On the syllabus
- Generative versus grounded: composing an authority versus fetching one
- What retrieval actually does, in plain terms — and what it fixes and does not fix
- The strongest move in legal AI: supply the document and ask questions about it
- The four-step authority check — exists, says that, still good law, and the step everyone skips
- Choosing tools by provenance; where AI genuinely accelerates legal work and where it introduces risk
In short
Session 6 is the craft session that makes the rest of the course operational. It draws the line between generative and grounded tools properly: a chatbot composes an authority that looks right, while a retrieval-grounded system answers from a document it actually fetched and can point at. It explains retrieval in plain terms, and is honest about what grounding fixes and what it does not — a grounded tool stops inventing and starts misreading, which is a better problem but still a problem. From there the hour installs the four-step authority check as a drill, and the hands-on half hour runs that drill on five authorities until it is automatic.
Why it matters for using AI well
You leave with one workflow you can run on autopilot and one instinct you cannot switch off: identify, verify it exists, confirm it says that, check it is still good law — and when you have the document, hand it over rather than asking the model to remember it.
What they leave with
The skill
Run the four-step authority check every time, and prefer supplying the document over asking the model to recall it.
The insight
The dangerous authority is not the one that does not exist — it is the real one that says something else, or the real one that has been overruled.
The moment they remember
The four-step check, run live against the room's own prediction. Students vote on how many of three confident AI-supplied authorities will survive; the presenter then checks all three on Indian Kanoon on the projector. Typically one does not exist, one says something materially different, and one is fine — and the room's prediction was optimistic. Then comes the reversal that makes the session: the same question through a grounded tool, with the provenance sitting right there in the answer. The realisation is not that AI cannot do law. It is that they had been holding the wrong tool and skipping the last step.
In this session
- 01
Generative versus grounded, properly stated: a generative chatbot predicts a plausible-looking answer with no connection to the databases it appears to be citing; a retrieval-grounded system searches real documents, puts what it finds in front of the model, and answers from those. Identical confidence on the surface, opposite reliability underneath.
- 02
What retrieval actually is, without the jargon: find the relevant documents, place them in the model's working context, and answer from them with a pointer back to the passage. The failure mode shifts from invention to misreading — the tool now cites something real, and may still characterise it wrongly — which is why grounding raises the floor without removing the duty. This is measured, not asserted: the Stanford RegLab evaluation found leading grounded legal research tools still hallucinating on roughly one in six to one in three queries — far better than a generalist model, and nowhere near the “hallucination-free” the marketing claimed. The Supreme Court's own White Paper cites the same study at p.56, and adds the rule that follows from it: “Users shall not employ one generative AI tool to verify or authenticate the content generated by another generative AI tool.”
- 03
The strongest single move available to a student or a junior: stop asking the model what the law is, and start giving it the judgment, the statute or the contract and asking questions about that. Everything Session 3 taught about the model's strengths says this is where it is reliable — it is reshaping text it was given.
- 04
The four-step authority check, as a habit rather than an aspiration: (1) ask the model to identify the authorities; (2) verify each exists on a legal database — Indian Kanoon, SCC Online, Manupatra; (3) confirm it says what the model claims, by opening it and finding the paragraph; (4) check it is still good law with a citator — SCC Online's Note Up, Manupatra's citation analysis, and KeyCite or Shepard's abroad. No step is optional, and the fourth is the one people skip, because a case that exists and says the right thing feels finished.
- 05
The three ways an authority fails, in ascending order of danger: it does not exist; it exists and says something different; it exists, says the right thing, and has been overruled. Only the first is obvious, and only the first is what most people check for.
- 06
Choosing a tool: judge it on grounding, provenance and verifiability rather than fluency. Ask what corpus it searches, whether it links to the primary source, whether it is citator-backed, and what it does when it does not know. A tool that never says “not found” is telling you something about itself.
- 07
Where AI genuinely accelerates: first-pass issue-spotting, drafting scaffolds you then fill from verified sources, document question-answering over material you supply, cross-jurisdictional comparison, and navigation of long documents. Where it introduces risk: any unsupplied fact, any citation, any figure, anything recent, and anything confidential entered into a public tool.
The four-step mirror
Run it on the class. Then on the machine.
An experiment on the room, the same effect explained in the model, a live demonstration on a real chatbot, and a named takeaway skill.
Generic chatbot vs. grounded tool — the provenance gap
On the class
The room commits to which it would trust more for an Indian case-law question — a fluent general chatbot answer, or a grounded research engine's answer — and the split is captured before anything is revealed.
In the model
A generative chatbot predicts a plausible-looking answer with no connection to the databases it appears to cite; a grounded, citator-backed engine retrieves from real sources. Same confidence on the surface, opposite reliability underneath.
Live chatbot
The same research question runs live through a generic chatbot and through a grounded tool, and the room compares the provenance of what each returns — can every citation be traced to a real, checkable source?
The skill
Match the tool to the task. For legal research, choose a grounded, citator-backed engine and verify provenance to source — fluency is not authority.
The four-step check, live
On the class
The room is shown a confident AI research answer with three authorities and votes on how many of the three will survive the four-step check.
In the model
The model's authorities are predictions of what a supporting case would look like. Some will exist and say something else; some will exist and have been overruled; some will not exist at all.
Live chatbot
The presenter runs all four steps on each authority in front of the room — database, paragraph, citator — and the tally goes up against the room's prediction.
The skill
Identify, verify it exists, confirm it says that, check it is still good law. Every time, and especially the fourth.
Hand it the document
On the class
The room predicts which will be more accurate: asking a model what a judgment held, or pasting the judgment in and asking the same question.
In the model
Asked to recall, the model predicts; given the text, it reshapes what it was handed — which Session 3 already established is where it is strong. The move converts a recall task into a reading task.
Live chatbot
The presenter asks about a judgment cold, then supplies the text and asks again, and the room compares accuracy and the availability of paragraph-level support.
The skill
When you have the source, supply it. Turn every recall question you can into a reading question.
Hands-on · on your own laptop
The Authority Check Drill
Drill the four-step check until it is automatic. Elicit five authorities from a generic chatbot on a proposition in your own area, run all four steps on each, and log the verdicts — then re-run the same question by supplying a real judgment yourself, and compare what changes.
Run of show · 30 minutes
- 0–8 min — Ask a generic chatbot for five authorities supporting a proposition in your area. Capture the output verbatim, citations and all.
- 8–22 min — Run all four steps on each authority: does it exist on a grounded database; does it say what was claimed (find the paragraph); is it still good law (citator). Log a verdict per authority.
- 22–28 min — Re-run the same question, but supply a real judgment yourself and ask questions about that text. Compare the reliability and the availability of paragraph-level support.
- 28–30 min — The room's aggregate survival rate goes on screen, broken down by the three failure classes: does not exist, says something else, no longer good law.
Deliverable
A five-row authority-check table — each authority with its four-step verdict and the trail — plus a short note comparing the recall-based answer against the answer given from a supplied document.
Key sources & cases
Indian research stack: SCC Online, Manupatra, Indian Kanoon, CaseMine
The grounded tools used to open the authority itself and to check every case in a traced line.
Citators: SCC Online Note Up, Manupatra citation analysis; KeyCite and Shepard's (comparative)
The lawyer's existing verification analogue — now applied to everything a model hands you. Step four of the authority check.
Kevin D. Ashley, Artificial Intelligence and Legal Analytics (2017)
What machines can and cannot do with legal texts — the academic ground under the generative/grounded distinction.
Supreme Court of India, White Paper on Artificial Intelligence and Judiciary (Centre for Research and Planning, Nov 2025)
Verified 2026-08-26 against the official PDF. Note the exact title — “Artificial Intelligence and Judiciary”, with no “the”. States that “Judges must remain the ultimate decision-makers, AI may assist, but it cannot substitute human judgement” (p.10) and that “AI may assist the judges, but cannot replace them” (p.65). Guideline 14 requires all AI-derived information to be independently verified before reliance; Guideline 15 forbids using one generative tool to verify another; Guideline 12 directs that private, confidential or legally privileged information not be input into ANY AI tool. On deployment it *suggests* courts prioritise secure in-house tools over “open-source or publicly accessible” ones — it does NOT restrict cloud AI, and never uses that framing.
Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023)
What skipping the check costs. Re-read here as a workflow failure rather than a cautionary anecdote.
Magesh, Surani, Dahl, Suzgun, Manning & Ho, “Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools,” Journal of Empirical Legal Studies (2025)
Stanford RegLab/HAI. The first preregistered empirical evaluation of grounded legal-research tools: Lexis+ AI hallucinated on roughly 17% of queries and Westlaw AI-Assisted Research on roughly 33%, against about 43% for GPT-4 — against vendor claims of “100% hallucination-free linked legal citations”. Preprint May 2024, peer-reviewed in JELS 2025. Verified 2026-08-26. Teach the finding *and* the controversy: vendors disputed the methodology and Stanford augmented the study — which is itself a lesson in checking a striking number.
Readings
- Supreme Court of India, White Paper on Artificial Intelligence and Judiciary (Centre for Research and Planning, Nov 2025)
- Kevin D. Ashley, Artificial Intelligence and Legal Analytics (2017)
- Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023) — re-read as a workflow failure
- Documentation for SCC Online Note Up and Manupatra citation analysis (citator practice)
- Magesh et al., “Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools,” J. Empirical Legal Studies (2025)
Next session
Session 07 / 08
AI and the Law I — Judgments, Doctrine and Comparative Analysis
Apply it to judgmentsSixteen hours, one professional discipline.
Using AI well is not a knack — it is competence, candour and verification, practised until they are habits you could defend in court.