Responsible Use >
Legal Research
Core Concepts
Generative AI research tools exist on a spectrum of reliability
Legal AI systems like Westlaw CoCounsel and Lexis Protégé are trained on curated legal data sets and are designed specifically for legal research. General-purpose AI systems like ChatGPT and Claude (whether free or paid tiers) draw from the internet and proprietary data and are not reliable for legal research.
Generative AI legal research tools can augment – not replace – traditional research
Legal AI systems can help identify relevant authorities, summarize case holdings, and flag potential issues, such as invalid citations or quotations. However, even these tools remain susceptible to errors and require independent verification. The types of errors a legal AI system might make include confusing a dissent with a concurrence, confusing trial and appellate court holdings, or confusing a quote from a party’s brief with a quote from a court’s holding. A 2025 study conducted by Stanford University researchers found that although RAG-based proprietary legal AI tools hallucinate less often than general-purpose AI systems, hallucinations and other errors still appear 17 - 33% of the time. Given the pace at which AI technology evolves, the models used in the Stanford study may have improved since these results were reported. However, no GenAI system is currently hallucination-proof.
General-purpose AI systems should not be used for legal research
General-purpose AI systems, whether free or paid tiers, (e.g., ChatGPT, Claude, or Gemini) are not trained on curated legal datasets and are prone to hallucination. They may fabricate citations or legal authorities that do not exist, misstate holdings, or invent quotations. They are not reliable tools for legal research.
Westlaw and Lexis offer citation and quotation verification tools
Tools and prompting options integrated into the Westlaw and Lexis legal AI suites can flag potential hallucinations as well as citation and quotation errors in the parties’ filings and in draft decisions and orders. Westlaw and Lexis also have non-GenAI tools that can be used to check for hallucinations. No tool, however, can identify citations to hallucinated authority with perfect accuracy.
Generative AI may be problematic when used for novel or nuanced legal questions
GenAI (whether general-purpose or legal) cannot reason through unsettled law or navigate genuinely novel or nuanced legal issues the way a human can. The use of GenAI for such questions risks delegating core judicial functions as well as importing errors that are difficult to detect because there is no established correct answer to verify against.
Verification is non-negotiable
Every citation must be independently verified using authoritative sources. Every holding and quotation must be checked against the original. This is true even for legal AI systems like Westlaw CoCounsel and Lexis Protégé.
Short Videos
-
GenAI Systems and Data Protection: How to Configure Privacy Settings
This video explains the implications of the type of GenAI system used for data protection and includes a demonstration of how to configure privacy settings.
-
Using Westlaw to Check for Hallucinations
This video offers a quick tutorial on how to use Westlaw drafting assistant to check for hallucinated citations and quotations.
-
Using Lexis to Check for Hallucinations
This video offers a quick tutorial on how to use Lexis Document Analysis to check for hallucinated citations and quotations.
Practical Guides
-
Talking to your Law Clerks about GenAI
Checklist of topics to address when discussing GenAI use with law clerks.
Frequently Asked Questions
This is not advisable. General-purpose AI systems are trained primarily on the internet rather than on curated legal data sets. As a result there is a greater tendency for those systems to produce less reliable information and hallucinate.
GenAI does not engage in legal reasoning the same way a judge or law clerk might. GenAI responds to a query by analyzing statistical probabilities. When training data does not have a clear-cut answer, the AI system may make one up. With novel legal questions, there is no way to verify the reliability of a system’s response. In addition, small differences in the wording of prompts can significantly affect output.
While dictionary content is typically incorporated into the training data of GenAI systems, these systems are not dictionaries. Dictionary content is static and does not depend on the phrasing of the prompt or the training data. GenAI responses are fluid, change over time, are shaped by the prompt, and can reflect cultural and other biases.
Curated Resources
-
Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
Journal of Empirical Legal Studies (2025)
-
AI Hallucination Cases
Damiencharlotin.com (updated on an ongoing basis)