Note: This article is part of AI Technology for Lawyers, a series explaining the technical foundations of AI for legal professionals. Start with the series introduction.
Making Models Behave
A language model trained solely to predict the next token might produce harmful content, give dangerous advice, or refuse to be helpful. The model learns from internet text, which includes everything from encyclopedias to manifestos. How do models learn to follow instructions, be helpful, and decline harmful requests?
The answer involves multiple training stages. After pre-training, models typically undergo instruction tuning, that is supervised fine-tuning on examples of instruction-following behavior. The primary alignment technique is RLHF, or Reinforcement Learning from Human Feedback.
RLHF works as follows:
- The model generates multiple responses to prompts.
- Human raters evaluate which responses are better, meaning more helpful, more accurate, less harmful.
- A separate "reward model" learns to predict human preferences.
- The main model is then trained to maximize its reward model scores.
RLHF explains why modern AI systems feel more useful than earlier versions. GPT-5 is not just a better language model than GPT-3; it has been extensively trained to follow instructions, be helpful, and decline harmful requests. The base model might know how to write malware; the RLHF-tuned model refuses.
Constitutional AI, developed by Anthropic, uses a set of principles, which it refers to as a "constitution," to guide model behavior. The model critiques its own outputs against these principles and revises accordingly. It is like giving the model an ethical framework it can consult.
An emerging method is Reinforcement Learning with AI Feedback, or RLAIF. Constitutional AI can be thought of as one type of RLAIF. RLAIF has the benefit of being more scalable, but it is unclear whether the models providing feedback - the reward signal for learning - fully express human preferences.
These techniques explain what "safety training" means in practice. When a vendor says their model has safety training, they are typically referring to RLHF, Constitutional AI, or similar alignment work. It also explains why jailbreaks exist: adversarial prompts can sometimes override the trained preferences, causing the model to ignore its safety training.
An important caveat: Alignment training makes models generally safer and more helpful, but it does not guarantee compliance with domain-specific requirements. A model trained to be "helpful, harmless, and honest" may still produce outputs that violate your organization's policies, regulatory obligations, or professional standards. Alignment is a foundation, not a substitute for use-case-specific evaluation and controls.
Evaluations: How Do We Know If It Works?
Model quality is not self-evident. Evaluation requires systematic testing, and there are several types:
Benchmarks are standardized tests for measuring model performance on specific tasks, enabling comparison across models. MMLU tests knowledge across academic subjects. HumanEval tests coding ability. These enable comparison: "Model A scores 87% on MMLU, Model B scores 82%."
But benchmarks have significant limitations. They measure performance on specific test conditions that may not reflect your use case. A model excelling at academic knowledge might fail on practical legal reasoning. There is also risk of "teaching to the test" - models may be optimized for benchmarks without corresponding real-world improvement.
Safety evaluations specifically test for harmful behaviors. Will the model help with CBRN (chemical, biological, radiological, nuclear) threats? Generate child sexual abuse material? Produce discriminatory outputs? These evaluations inform safety claims and regulatory compliance.
Red teaming involves dedicated teams attempting to find vulnerabilities and failure modes, like jailbreaks, prompt injections, confidentiality breaches. Red teaming is referenced in the EU AI Act and was a requirement under President Biden's AI Executive Order (though that order was subsequently rescinded).
For procurement and governance, ask: What evaluations were performed? By whom? On what tasks? How recent? Evaluation results from six months ago may not reflect current model versions.
Reasoning Models
The newest development in AI capabilities is reasoning models: systems specifically trained for complex, multi-step reasoning.
Traditional LLMs generate responses token by token, predicting the next token. Reasoning models take time to "think" before responding. They break problems into steps, consider alternatives, check their work, and revise their reasoning. OpenAI's GPT-4 and GPT-5, the most recent versions of Anthropic’s Claude, Google’s Gemini, and DeepSeek's R1 are prominent examples.
For legal work, reasoning models show promise for complex analysis tasks: contract review requiring synthesis across multiple provisions, legal research requiring evaluation of competing authorities, risk assessment requiring consideration of multiple factors.
But they come with tradeoffs. Reasoning models use significantly more computation per query because they are literally doing more work internally. This affects both cost and latency. A query that costs fractions of a cent on a standard model might cost several cents on a reasoning model. Response times stretch from seconds to minutes.
There is also a transparency consideration. Some reasoning models expose a chain-of-thought trace showing intermediate steps; others keep reasoning internal. When visible, these traces can help evaluate whether the analysis is sound, but they are not guaranteed to reflect the model's actual computational process. The displayed "reasoning" may be a post-hoc rationalization rather than a faithful representation of how the model reached its conclusion. When reasoning is entirely hidden, you are trusting a black box on more complex questions.
The Explainability Problem
When an AI system produces an incorrect or harmful output, such as a discriminatory hiring recommendation, a flawed legal conclusion, a credit denial, a natural question follows: Why did it do that?
The honest answer, in most cases, is that we cannot say. We can observe the mathematical operations: billions of weight values were multiplied and added, producing a probability distribution from which the output was sampled. But we cannot right now extract a human-interpretable rationale. The model does not "reason" in the sense of applying discrete logical rules that can be audited and explained. It produces outputs based on statistical patterns learned from its training data. These patterns exist as abstract numerical relationships, not as propositions we can inspect.
Think back to the example of the art inspector. The inspector can say "It's a Picasso" with 99% accuracy but cannot testify in court why, other than "it looks like the other ones."
This creates challenges for at least three areas of law:
- Administrative law. Many regulatory schemes require reasoned decision-making. If an agency uses AI to make or inform decisions, can it satisfy requirements to explain the basis for its actions? The "black box" nature of neural networks sits uneasily with doctrines requiring articulated rationales.
- Discrimination claims. If an AI system produces disparate outcomes, proving discriminatory intent or identifying the source of bias is difficult when no one, including the system's creators, can fully explain why the model produces specific outputs. Statistical auditing of outcomes may be the only practical approach.
- Liability defense. When an AI system causes harm, defendants may struggle to explain what went wrong or demonstrate that they exercised reasonable care in deployment. "The model made a mistake we cannot explain" is not a satisfying defense, but it may sometimes be the truthful one.
Some techniques offer partial transparency. Attention visualization shows which parts of the input the model weighted heavily. Chain-of-thought prompting produces intermediate reasoning steps. But these are approximations, not true explanations. The displayed "reasoning" may not reflect the actual computational process that produced the output. The fundamental opacity of neural networks remains a legal and governance challenge without a clear technical solution.
System Updates and Drift
AI systems are not static - they change over time. Usually new models are superior, but they also sometimes change behavior in unexpected ways. This creates challenges for validated workflows and regulatory compliance.
Model updates are intentional changes to the model itself: new weights, different safety settings, improved capabilities. Providers update models regularly, often without detailed notification.
System drift refers to behavior changes across the broader AI system. Model drift is one source, but enterprise AI systems can also drift due to:
- Changes to system prompts or developer instructions
- Updates to retrieval indexes (new documents added, old ones removed)
- Modifications to connector permissions or data sources
- Changes in guardrail configurations
- Updates to orchestration logic or tool definitions
For regulated industries or critical workflows, this is a significant concern. A system validated in January might behave differently in July, even if the underlying model has not changed. If you are relying on consistent behavior, drift can cause compliance failures without anyone noticing until something goes wrong.
The primary responses to this issue are technical: have a suite of tests and observe how answers change over time, or characterize outputs and their distribution over time. But many times you may not have direct access to a model provided by a vendor. In that case, contractual responses include:
Version pinning. Lock to a specific model version so behavior does not change without your consent.
Notice requirements. Require the vendor to inform you of material changes before they take effect.
Regression testing obligations. Require the vendor to test updates against your use cases before deployment.
Deprecation policies. Specify how much notice you get before a model version is retired.
System-level change control. Extend version control and change notification beyond just the model to cover prompts, retrieval configurations, and tool permissions.
The EU AI Act requires post-market monitoring - ongoing observation of how high-risk systems perform after deployment. This is not just about bugs; it is about detecting when behavior drifts in ways that affect compliance.
This is the last article in the series. Continue here for resources and further reading .

