Note: This article is part of AI Technology for Lawyers, a series explaining the technical foundations of AI for legal professionals. Start with the series introduction.
The Transformer Revolution
In 2017, Google researchers published a paper titled "Attention Is All You Need" that introduced the transformer architecture. This was the inflection point that made modern AI possible.
The key innovation is the attention mechanism. When a transformer processes text, it determines which parts of the input are relevant to producing each part of the output. If a document mentions "the company" in paragraph one and uses "it" in paragraph fifteen, the attention mechanism can connect those references statistically - something prior architectures struggled with.
Every major chat-based AI system you have encountered is likely built on a transformer architecture. When you see "transformer-based" in vendor materials, that signals a system with LLM-like capabilities.
An Analogy for Training
One persistent misunderstanding is that AI models work like sophisticated databases, storing everything they see and retrieving pieces to assemble into outputs. This is not how they work.
Consider two analogies that illuminate the training process: the art inspector and the law student.
The Art Inspector. Imagine an art inspector hired to examine every painting in the Louvre. The inspector has no background in art and no preconceived ideas about what makes a Picasso a Picasso. Lacking guidance, the inspector measures everything about each painting: brushstrokes per square inch, paint thickness, color relationships, the corner where artists sign their names. Before each measurement, the inspector tries to predict the answer based on what has been observed so far. Early predictions are wrong. But after studying thousands of paintings, the inspector becomes remarkably accurate. He does this not by memorizing paintings, but by learning patterns that distinguish each artist’s work.
The art inspector analogy shows how models start with a clean slate. Before training, models are filled with random data. All associations emerge from observation. But what is recorded is not the raw measurements - it is the changing probabilities associated with different inputs. One limitation of the analogy: the inspector's "measurements" are human-interpretable features like brushstrokes and color. A neural network's learned representations are not generally interpretable in the same way. While researchers have found some places where one particular measurement corresponds to something we would find meaningful, in most cases the measurements are abstract numerical patterns that do not map directly to concepts humans would recognize.
The Law Student. When students begin law school, they are told their job is not just to learn the law but to learn how to "think like a lawyer." The case method involves studying judicial decisions to understand legal principles, then applying those principles to new situations. Through the Socratic method, students struggle to answer questions, receive feedback, and refine their understanding. A successful student can take a hypothetical situation and generate an appropriate analysis. She does this not by memorizing holdings, but by learning how courts have weighted various facts and principles.
The law student analogy shows how models derive principles through inference. Just as the student is never explicitly taught which principles to apply, models are not instructed what their inputs "mean." The meaning that emerges is a complex probability function with billions of parameters.
The law student analogy also demonstrates why direct memorization is antithetical to training goals. Models that simply memorize training examples perform poorly on new inputs. Training processes use regularization techniques, such as dropout, weight decay, and data augmentation, specifically to discourage memorization and encourage the model to learn generalizable patterns instead.
The Training Process
Training is fundamentally a guessing game played at enormous scale. For most modern LLMs, training uses next-token prediction: given a sequence of text, predict the next token. The process follows five steps, repeated billions of times:
- Convert training data to tokens. Text and code are broken into tokens. Tokens are what the units models actually process. Tokens are chunks of text, not necessarily whole words. In multimodal models, non-text inputs like images are also encoded into token-like units. In English, a common rule of thumb is roughly four characters per token, though this varies by model, tokenizer, and content. Technical material and non-English languages often require more tokens per concept.
- Present a partial sequence. "It was a dark and stormy": what comes next?
- Predict the next token. The model guesses "night."
- Measure the error. A loss function measures how wrong the prediction was.
- Adjust weights to reduce error. The model's weights - the millions or billions of learned numerical parameters - are adjusted slightly to make that prediction more likely next time.
After processing millions of examples, the model has learned statistical patterns about language, concepts, and reasoning. This is critical for legal analysis: models learn patterns, not facts. A model does not "know" that Paris is the capital of France; it has learned that in text about France and capitals, "Paris" frequently appears. This explains both their capabilities and their tendency to produce confident-sounding errors.
Three Distinct Phases
Contracts and governance frameworks often distinguish three phases of model use, each with different legal implications:
| Phase | What Happens | Who Does It | Legal Issues |
|---|---|---|---|
| Pre-training | Creates base model from broad data | Model provider | Training data provenance, copyright exposure |
| Fine-tuning | Customizes for specific tasks | Provider or customer | Customer data rights, artifact ownership |
| Inference / Generation | Uses model to process new inputs and create outputs | Customer/users | Transient use, logging, retention |
Pre-training is the initial phase where the foundation model is created using massive datasets, usually including content scraped from the internet, but also including text found in articles, books, government documents, and other sources. Pre-training data provenance is the primary driver of copyright exposure for foundation models.
Fine-tuning takes a pre-trained model and trains it further on specialized data. A law firm might fine-tune on its own work product to create a model familiar with its writing style. Fine-tuning raises sharp IP and confidentiality issues: who owns the resulting weights? Can the customer take the fine-tuned model elsewhere?
Inference is the operational phase. When you interact with a deployed model, you are using the model to process new inputs and infer new relationships. During inference, data flows through the model transiently to produce a response but does not change the weights. This is why many agreements distinguish inference from training: "you may use our data for inference but not for training" means data can flow through the model but cannot be used to modify it.
Generation is a special case of inference, where the outputs are iteratively created using inference in a loop. This is what allows “chat” or “reasoning” AI systems to function. Specifically, text generation is inference performed repeatedly: each step predicts the next token conditioned on the tokens generated so far.
Many enterprises choose RAG (discussed in Article 5) over fine-tuning precisely to avoid "training on sensitive data" and to simplify deletion, retention, and IP ownership issues. This architectural choice often has as much to do with legal risk management as with technical requirements.
A note for contract drafters: Most "no training on our data" clauses address weight updates, specifically whether customer content can modify the model. They do not automatically address logging, retention, human review of prompts and outputs, or use of conversation data for safety monitoring. Each of these requires separate treatment.
Continue here to the next article in the series: Tokens, Context, and Inference

