Building an agent that can do something useful is becoming surprisingly easy. Give a model access to a few tools, connect it to enterprise data, and define a goal. The agent can quickly begin breaking a request into steps, gathering information, making decisions, calling external systems, and continuing based on what it observes.
The model gets most of the attention because it is the visible intelligence in an agentic system. But an agent cannot reason about evidence it never receives, cannot safely act on information it was not authorized to see, and cannot make reliable decisions from information that is stale, incomplete, or incorrectly ranked.
Consider a request like this:
"Audit the Q3 compliance risks for Project Titan based on internal audit reports and active team discussions, highlight immediate remediation steps, and draft a risk mitigation ticket."
The agent needs to understand what Project Titan refers to. It can locate the relevant audit reports and identify recent team discussions. It might need to determine which information the user is permitted to access. Often there is a need to reconcile information that may have changed since those documents were created. It then needs to assemble enough evidence to identify meaningful risks, determine what requires attention, and draft an actionable ticket.
Crucially, the agent may need to search more than once. The first search might identify the project. Subsequent steps might find the relevant audit reports or investigate a particular finding. Another search might be necessary after the agent discovers something unexpected or calls an external system, possibly followed by a final verification search to ensure the state hasn't changed before taking action. Search architecture for AI agents determines which evidence reaches the model at each step of a workflow.
Retrieval is no longer simply a relevance problem once an agent can take action. There is now a question of whether the agent is operating from the right information, with the right permissions, at the right time, and at a cost the system can sustain. Without good search architecture there is at best some hindrance to the process and at worst a catastrophic failure waiting to happen.
Authorization: Retrieving Evidence Within the User’s Permissions
The first question in an enterprise search system is often not what the user is looking for but who the user is. An enterprise agent operates on behalf of someone which means the audit report relevant to the request may not be accessible to this particular user. A discussion about Project Titan may exist but parts of it might be restricted to a specific team. A document may contain exactly the information the agent needs while being classified in a way that prevents the user from viewing it.
These constraints cannot simply be ignored until after retrieval. While it is possible to retrieve broadly and remove unauthorized documents post-query, doing so fundamentally alters the context available to the agent after the original ranking has already taken place. If five highly relevant documents are retrieved and four are removed by an authorization check, the remaining context is vastly different from what produced the ranking. In some environments, the system may have to over-fetch candidates dramatically to ensure enough authorized material remains after filtering. This has the potential to add unpredictable latency, load, and inaccuracy to the process which must somehow be accounted for.
Search systems have been solving versions of this problem for years using document-level security, tenant isolation, metadata constraints, and query execution strategies that incorporate authorization directly into retrieval rather than treating it as a post-processing cleanup. Since the context of an agent's request is critical to its success this carries consequences. If the search layer exposes something the user should not see, this unauthorized data actively influences the agent's downstream reasoning and actions. The architectural challenge for user security is maintaining adequate authorized recall and performance while strictly enforcing the security boundary across growing numbers of users, tenants, and security policies.
Data Freshness and Agent Reliability
Data freshness dictates how accurately retrieved information reflects the live environment an agent operates in. Freshness is easy to overlook in a prototype, where a static document collection indexed once a day works perfectly well for a demonstration. But an agent executing tasks in a live enterprise environment must reason about information that changes constantly throughout the day. A compliance assessment might depend on a chat message sent ten minutes ago. A support agent might need to know an incident was resolved five minutes earlier. A financial workflow might require the live transaction state rather than a record from a nightly batch run.
Search architecture must often manage several ingestion patterns simultaneously where historical archives can be processed through batch pipelines, transactional updates can arrive via change-data-capture (CDC) streams, and live communications can flow continuously from platforms like Teams or Slack. Increasing freshness, however, is never free. Frequent near-real-time (NRT) indexing directly impacts throughput, segment creation and merging, memory overhead, disk I/O, and query latency. At scale, unmanaged indexing pressure can trigger severe performance bottlenecks, deadlocks, or catastrophic cluster failures.
The fundamental architectural question is how fresh the information needs to be for the specific decision the agent is about to make. For a historical audit archive, a few hours or even a day may make no practical difference. For an active compliance investigation, a five-minute-old conversation might completely alter the conclusion. Because an agent uses retrieved context to execute actions, operating on yesterday's state means the agent isn't just returning stale data. Agents could be executing actions on an obsolete world state. Freshness is therefore a direct component of action reliability.
Agents Need Both Meaning and Exactness
The real fun begins in the actual semantic retrieval. Enterprise information is full of things that need to be understood semantically alongside things that must be matched exactly. The request may refer to "compliance risks," "remediation steps," or "unresolved controls." These are conceptual ideas where semantic vector retrieval excels. But enterprise data also contains project names, contract numbers, error codes, regulatory citations, customer identifiers, product SKUs, and internal acronyms. These are specific tokens whose exact presence determines whether a result is useful.
Lexical, keyword search with decades of lessons learned can do a great deal of heavy lifting and very lightweight cost. BM25 scoring and inverted indexes provide precise matching for terms, acronyms, and domain-specific vocabulary that semantic models frequently underweight. Analysis chains are powerful tools that provide custom tokenizers, character filters, edge-ngrams, stemming, and synonyms. These simple and mechanical techniques handle hyphenated identifiers, variant codes, and specialized terminology that cannot be solved simply by increasing vector dimensions.
An agent investigating Project Titan needs to search for a concept like a "control failure" while guaranteeing that the results explicitly match the correct project, organization, date range, and tenant. A production system must apply these structured metadata filters and security constraints before the most expensive stages of retrieval take place. Rather than hunting for a single magic retrieval technology, a mature architecture combines lexical and dense retrieval, applies structured filters during candidate generation, and uses reranking to supply the agent with precise, actionable evidence.
Evidence Assembly: From Documents to Useful Context
Let's imagine that the search system identifies a 60-page audit report as highly relevant, and passes all 60 pages to the LLM. This technically gives the agent access to the information, but it also introduces severe latency, ballooning token costs, and the risk that critical evidence gets buried in noise. The retrieval layer has the opportunity to transform raw source material into curated evidence.
Parent/child chunking addresses this directly by creating small passages (maybe 200–400 tokens) that are assembled for precise keyword and semantic matching while maintaining a reference to their parent document section. When a passage matches, the system retrieves the surrounding parent context to supply coherent, self-contained evidence without passing the entire source document downstream. An inexpensive first stage might fetch a broad candidate pool, after which a cross-encoder reranker evaluates those candidates against the agent's specific sub-goal to pass only the highest-confidence evidence to the model. Multi-stage, agentic workflows excel in this process, and the search architecture paves the way.
The ability to do this efficiently depends on foundational decisions made long before a reranker is introduced. We take some care to consider how documents are modeled, which fields are indexed, how content is chunked, how metadata is represented, and how the index is partitioned. At small scale the data can be somewhat sloppy, a system can overfetch candidates, apply expensive filters, and pass bloated context because the underlying costs are trivial. Production removes that luxury. Index topology, schema design, and retrieval strategy directly dictate whether the system can meet its latency and cost requirements when thousands of agents search concurrently.
Search architecture for AI agents determines which evidence reaches the model at each step of a workflow.
Every Agent Step Changes the Retrieval Problem
In a traditional search interface, a user submits a query and receives a result. In an agentic workflow, the agent workflow may search, interpret the result, take an action, observe the outcome, and possibly repeat. In the Project Titan scenario, the first search likely identifies the project and compliance material, finding an unresolved control issue prompts a second search for details. That search may uncover a recent team discussion suggesting the issue was addressed. This might trigger a third search to verify current status. Only then does the agent decide whether a remediation ticket is necessary.
Because every search is part of a multiturn chain, a retrieval failure may cascade polluted and inaccurate data down the chain. If an early search misses a critical document, the agent forms an incomplete understanding that dictates what it searches for next. Subsequent queries on a flawed premise created by the initial retrieval failure will make a failed outcome much more likely. On the other hand, precise retrieval at an early step enables the agent to ask better follow-up questions, investigate the right systems, and verify assumptions before taking action.
For an agentic system, retrieval is simultaneously an evidence quality problem and a workload scaling problem. Search architecture for agents thrive when the retrieval layer is designed at the workflow level rather than the single query level.
Evidence Provenance and Search Evaluation
A very important part of agentic workflows is understanding and explaining where evidence came from. By binding each retrieved passage to its underlying document ID, section anchor, version number, and ingestion timestamp, the search layer preserves an immutable chain of evidence. If an auditor asks why a specific remediation ticket was opened, the system can trace the recommendation back to the exact version of the source material that informed it, which is a critical requirement when documents evolve over time.
This is data lineage, or provenance, that answers where the evidence came from. Search evaluation answers a distinct, complementary question: did retrieval provide the evidence the workflow actually needed? When an agent produces a flawed recommendation, engineering teams often default to tweaking prompts or swapping LLM models. But if the source document was never retrieved, if the critical passage ranked below the context cutoff, or if security filters removed valid data, no amount of prompt engineering can compensate for missing context. Worse, a prompt won't reliably ensure access data restrictions are deterministically followed.
Search evaluation must be treated as a distinct engineering discipline. Standard metrics like Recall@K, NDCG, and MRR measure retrieval accuracy, while security benchmarks test for ACL leakage and freshness metrics track ingestion latency. Ultimately, these measurements connect directly to workflow success on whether the task completed, how many retrieval loops were required, and how often missing evidence forced human intervention.
Production Changes the Economics of Retrieval
A retrieval decision that looks insignificant in isolation can become consequential when it sits inside an agentic workflow. An agent may search several times during a single task, and those searches may span multiple data sources and retrieval stages. Candidate generation, filtering, reranking, context construction, and indexing are no longer costs associated with an occasional user query. They become part of a repeated execution path.
A few hundred milliseconds added to each retrieval operation can materially increase the latency of a multistep workflow. A modest increase in candidate count can drive up reranking costs across millions of executions. A larger context can increase model inference costs even when the underlying search operation is inexpensive.
The challenge is that these are not independent trade-offs. Reducing retrieval cost too aggressively can lower recall and leave the agent without evidence it needs. Retrieving more broadly may improve recall, but increase latency and the amount of context sent to the model. Pushing for fresher data can improve the quality of decisions that depend on current state, while increasing indexing work and potentially reducing query throughput. Applying security constraints after retrieval can preserve a simple query path, but may remove so many candidates that the agent is left without sufficient authorized evidence.
At production scale, these decisions have to be considered together. The goal is not to maximize search precision, minimize infrastructure cost, or achieve the freshest possible index in isolation. The goal is to provide the agent with sufficient evidence to complete its task reliably, while keeping retrieval within the workflow's latency, security, reliability, and cost constraints.
From Evidence to Action
Return to the analyst’s original request: auditing Q3 compliance risks for Project Titan, identifying remediation steps, and drafting a mitigation ticket. The value of that ticket depends on whether the agent found enough reliable evidence to identify what still requires attention.
The agent must evaluate that evidence, prioritize findings, and decide what to do next. Search architecture shapes the information available to support those decisions.
In an agentic workflow, retrieval matters at every decision point. Production search architecture must provide evidence that is authorized, current, relevant, and traceable, within the workflow’s latency and cost requirements. That makes search architecture central to how reliably an agent can move from gathering information to taking action.