Loading...
All articles
AI Governance 5 min read

AI Technology for Lawyers: Tokens, Context, and Inference

A document broken into tokens and processed through a model context window

Note: This article is part of AI Technology for Lawyers, a series explaining the technical foundations of AI for legal professionals. Start with the series introduction.

Tokens: The Currency of AI

Tokens are the units that models actually process. They are chunks of text, not necessarily whole words. In English, a common rule of thumb is roughly four characters per token, but this varies by model and content. A common word like "the" is typically a single token; a technical term like "indemnification" might be three or four tokens. Understanding tokens helps you understand two things: costs and limits.

Costs. AI services typically charge per token, counting both input (what you send) and output (what you receive). A simple query might use a few hundred tokens; analyzing a long document might use tens of thousands. When vendors quote pricing, ask whether they count input tokens, output tokens, or both.

Limits. Context windows - how much text the model can consider at once - are measured in tokens. This matters more than most users realize.

Context Windows and Their Limits

The context window is the maximum amount of text a model can "see" at once, encompassing both your input and the model's output. Context window sizes vary significantly across models and versions. They vary from roughly 8,000 tokens in earlier systems to over one million tokens in some current models. Treat context window size as a versioned specification that should be confirmed during procurement and tracked over time, as providers frequently update these limits.

These numbers sound large, but there are two problems lawyers should understand:

  1. Truncation. When input exceeds the context window, content is dropped, often without any notification to the user. If you ask an AI to analyze a 500-page contract and the model's context window can fit only 100 pages, the system is not analyzing the full document. It may analyze the first 100 pages and ignore the rest. It may sample from different sections. The behavior varies by implementation, but the result is the same: the AI is not considering everything you provided. For legal document review, this creates serious risk.
  2. Attention degradation. Context window size does not equal attention quality. Research has documented a phenomenon called "lost in the middle": models often perform worse on information buried in the middle of long inputs compared to information at the beginning or end. Just because content fits within the context window does not mean the model will properly consider it.

The practical implication: when using AI for document analysis, verify that the model can actually process what you are asking it to process, and consider whether critical information might be lost to truncation or poor attention.

A note for practitioners: Silent truncation creates a foreseeable failure mode. When an AI system fails to consider material provisions because it silently dropped half the document, this is not an exotic edge case - it is a predictable consequence of known technical limitations. Treat it as a competence and supervision issue. Controls include document chunking strategies, section-by-section prompting, and requiring the system to quote or cite the specific provisions it relied on.

Deterministic vs. Probabilistic

Users are used to software being deterministic. What that means is that every time you use the software a certain way, it will give you the same answer. Your calculator app doesn’t usually say that 2+2=4, but occasionally say that 2+2=5. Except for in the case of some very rare bugs, software does the same thing, every time. Traditional AI software generally shares this quality.

In contrast, the newest AI software is probabilistic. That means that it varies what it returns, even when you provide the same inputs (i.e., the same prompt). It may give very similar answers, but they are not guaranteed to be the same. That is both helpful and unhelpful. We value consistency in computers. On the other hand, it is this probabilistic nature that gives current AI models a lot of their power.

From a controls perspective, this is what makes AI governance separate from other types of governance. No matter how “good” your AI is, you have to plan on the fact that a certain percentage of the time it will make the “wrong” decision or return the “wrong” output.

It is an area of current research on how to make these new AI systems more deterministic while maintaining their power. However, for right now you should assume that AI models will fail a certain percentage of the time.

Temperature and Determinism

Temperature is a parameter controlling how random the model's outputs are. At temperature 0, outputs are more deterministic. At higher temperatures (commonly ranging up to 1.0 or 2.0), outputs become more variable and creative, but potentially less reliable. In theory a temperature setting of zero (together with some other parameters) should make the output fully deterministic. In practice, details associated with how computers process numbers can cause changes in the outputs.

This matters for two reasons:

Reproducibility. If you need consistent outputs, like for a validated legal workflow, for audit purposes, or for testing, then document the temperature settings. Results may vary run to run even with identical inputs unless temperature is controlled.

Understanding the nature of AI outputs. Most generative AI is probabilistic at its core. Outputs are generated by sampling from learned probability distributions, not by executing logical rules or retrieving stored facts. Even when configured for more deterministic operation, the model is selecting the statistically most likely output based on its training patterns. It is not computing a provably correct answer.

This explains hallucination: models can be confidently wrong because they are producing statistically plausible text, not verified facts. It also explains why traditional software warranties do not map cleanly to AI systems. A database either returns the correct record or it does not. An AI system produces outputs that exist on a spectrum of plausibility, and plausibility is not the same as accuracy.

This is why using a model as a legal search tool is fraught. A database returns only records of cases that actually exist. A model will return a case that is highly likely to support the legal principle you want to support. Sometimes this case may exist - cases that really exist have higher probabilities than imaginary cases - but not necessarily.

Continue here to the next article in the series: Model vs. System

Van Lindberg

CEO, Model Monster