Skip to main content

Responsible Use >

Summarization

Core Concepts

Generative AI can be helpful for summarization tasks, but it is imperfect

GenAI can be effective at distilling lengthy documents into concise summaries, creating chronologies from disparate information, and organizing information into structured formats. However, LLMs focus on word patterns and frequency. They may miss language that prioritizes or qualifies concepts, such as negation (e.g., the words ‘no,’ ‘not’ and ‘without’). This is an inherent part of LLM architecture, not a flaw per se. But this architecture can undermine the accuracy of a GenAI summary. As with all GenAI output, the summary must be reviewed for accuracy, neutrality, and completeness.

Accuracy of Generative Ai may degrade as input volume increases ('context rot')

An LLM’s context window is the maximum amount of information a GenAI system can process or analyze during a single interaction. It represents all the information the AI system can see in a given conversation; the system cannot process any information outside of its context window. Information entered into a GenAI system is broken down into tokens, sequences of textual characters that make up part or all of a word. The context window is measured in tokens and its size varies system to system, depending on the AI system’s architecture. Earlier AI models have context windows of about two to three thousand tokens (about three to six pages or fifteen hundred to three thousand words), while newer models range up to one to two million tokens (fifteen hundred to five thousand pages or seven hundred fifty thousand to one million words).

The larger the context window, the more information an LLM can process. As engagement with a GenAI system proceeds (through a series of prompts or conversations), its context window fills. GenAI systems tend to favor tokens at the beginning and those at the end of the input. Notably, the size of a system’s context window is not always obvious and the system may not provide this information to the user.

When using GenAI to summarize documents, if the quantity of information the system is asked to summarize exceeds its context window, the system will not retrieve information outside the context window and may struggle to retrieve information near the middle of the context window, leading to diminished accuracy. This phenomenon is called 'context rot.' Even LLMs with large context windows often prioritize information at the beginning and at the end of the input, potentially omitting important information in the middle. This is referred to as 'lost in the middle.'

Appropriate use cases for Generative AI summarization and organization

GenAI can summarize lengthy briefs, exhibits, or transcripts, create case timelines, organize facts, and structure complex regulatory materials. It also has been used to decipher handwriting in filings by self-represented litigants (with moderate but not 100% accuracy). Although GenAI tends to hallucinate less when working with a circumscribed set of materials, it still can make errors. In a recent case, counsel used GenAI to summarize documents and deposition transcripts for a declaration provided to the court; the declaration included hallucinated facts. GenAI summaries require human review and verification. 

The neutral prompting principle

How a prompt is phrased has an impact on GenAI output. Prompts given to GenAI tools should be objective and neutral, without presupposing or signaling a particular outcome. There is a difference between asking GenAI to ‘objectively summarize the parties’ arguments’ and asking it to ‘explain which party has the better argument and why.’

Summarization and organization versus analysis

While it is not improper to use GenAI for some types of analysis, it is vital to be clear about what you are asking it to do. There is an important distinction between asking GenAI to summarize or organize the parties’ arguments (appropriate) and asking GenAI whether these arguments are correct (not appropriate). This issue can manifest in subtle ways. For example, there is a difference between asking GenAI to identify ‘all allegations’ versus ‘key allegations.’ The former is more clearly summarization while the latter involves analysis and requires the GenAI to make choices about what is legally significant. One cannot rely on a GenAI tool to clearly differentiate among the tasks of summarization, organization, and analysis. GenAI-generated summaries and chronologies should always be verified and reviewed for accuracy. 

Short Videos

Practical Guides

  • Talking to your Law Clerks about GenAI

    Checklist of topics to address when discussing GenAI use with law clerks.

Frequently Asked Questions

Curated Resources