Note: This article is part of AI Technology for Lawyers, a series explaining the technical foundations of AI for legal professionals. Start with the series introduction.
Why Retrieval Matters
Models only know what was in their training data, and that data has a cutoff date. A model's knowledge cutoff depends on the specific model and version; treat it as a versioned specification to confirm during procurement. Models cannot know about events, documents, or information that came after their training cutoff. They also do not know the contents of your company's internal documents, contracts, or policies unless you provide them.
One amusing thing is that sometimes when you are working with an AI to review a document that has factual assertions past the cutoff date, the AI will warn that you have speculative or fictional data in your document. You should take that as an indicator that the model is only “checking” the document against its own probabilistic knowledge, and is not grounding its results on some sort of search.
Retrieval-Augmented Generation (frequently referred to as RAG) addresses this limitation by connecting the model to external knowledge bases. RAG is the most common architecture for enterprise AI applications, and understanding it is essential for advising on AI deployments.
How RAG Works
RAG operates in five steps:
- User asks a question. "What does our policy say about remote work eligibility?"
- The system converts the question to an embedding. An embedding is a list of numbers, known as a vector, that represents the semantic meaning of the text. Similar meanings produce similar numbers. This allows the system to find relevant content based on meaning, not just keyword matching.
- The system searches a knowledge base for similar embeddings. The knowledge base contains embeddings of your documents, pre-computed and stored in a vector database. The system finds documents whose embeddings are close to the query embedding. Usually this returns documents that are semantically related to the question.
- Retrieved documents become context for the model. The relevant document chunks are inserted into the prompt, giving the model information it would not otherwise have.
- The model generates a response grounded in those documents. Instead of relying solely on training data patterns, the model can draw on specific, current, authoritative sources.
RAG is why vendors say "talk to your documents" or "search your knowledge base with AI." In a typical RAG setup, the model is not trained on your documents. The system retrieves and supplies relevant excerpts as context.
A Note on Embeddings
Although the model does not learn from your documents in a RAG configuration, the vector database does store embeddings derived from them. Embeddings are mathematical representations of semantic meaning - compressed, transformed versions of the source content. This raises a question: do embeddings constitute "derived data" with independent legal significance?
The answer may matter for data governance. Embeddings are not designed to be reversible, but they can leak information about source content and may be partially invertible in some settings. Research has shown that semantic relationships encoded in embeddings can reveal sensitive information about document contents. When negotiating data handling terms for RAG systems, consider whether embeddings are covered by deletion rights, retention limits, and confidentiality obligations, not just the source documents themselves.
Benefits of RAG
RAG offers three principal benefits:
Reduced hallucination. By grounding outputs in actual sources, RAG can reduce (though not eliminate) fabricated information.
Current information. RAG systems can include recent documents that postdate the model's training cutoff.
Citations. Well-designed RAG systems can provide source attribution, showing which documents informed the response.
Risks of RAG
RAG also introduces risks that require governance attention:
Access control. Which documents can the system retrieve? Are the same access controls that govern human access to documents also governing AI access? If an employee asks about salary ranges and the system retrieves confidential compensation data the employee should not see, you have a confidentiality incident. RAG systems must respect document-level permissions.
Retrieval logging. Is there an audit trail of what documents the system accessed for each query? If something goes wrong, can you reconstruct what the AI considered? Retrieval logs can become discoverable records, and they can also be critical evidence in incident reconstruction.
Indirect prompt injection. If malicious content exists in documents the system might retrieve (like a poisoned webpage, a manipulated email, a document crafted by an adversary) that content could influence model behavior. This is particularly concerning when RAG systems index content from sources that are not fully trusted. We discuss this attack vector further in Article 6.
False confidence. RAG outputs may cite sources, but citation does not guarantee accuracy. The model might misread the source, summarize it incorrectly, or combine information from multiple sources in misleading ways. Citation reduces but does not eliminate the need for verification.
IP risk. Using RAG with data sources to which you do not have full rights likely implicates various IP risks. In particular, RAG outputs are designed to frequently include parts of source documents in the outputs. This can create trade secret or copyright risks.
A note for practitioners: RAG transforms abstract "AI risk" into concrete access-control, logging, and data-governance risk. The central question becomes: what can the retriever see? If the answer is "everything in the enterprise," the AI system inherits every access-control problem in the organization. Pay special attention in litigation contexts; if the system can display output from privileged documents to non-privileged users, that can waive privilege.
Continue here to the next article in the series: Agentic AI

