AWS has documented how to use agentic retrieval with LangChain and Amazon Bedrock Managed Knowledge Bases, giving RAG developers a planning-based path for multi-part questions and trace data for inspecting its decisions. The approach can improve evidence coverage for comparative or multi-hop queries, but it requires different integration and permissions choices than ordinary retrieval and is not AWS’s recommended default for simple questions.
Two APIs, different retrieval behavior
Amazon Bedrock Managed Knowledge Bases handles chunking, embedding, storage and retrieval after an application configures a data source. AWS’s walkthrough contrasts its Retrieve API, which runs one hybrid search and returns scored chunks, with AgenticRetrieveStream, which plans retrieval, issues sub-queries, assesses the evidence and can search again.
In LangChain, AmazonKnowledgeBasesRetriever wraps the standard API and can be used directly in a chain. For managed knowledge bases, AWS says the configuration should use managedSearchConfiguration; vectorSearchConfiguration is the older path for knowledge bases with a customer-run vector store. Standard results are LangChain Document objects with relevance in metadata["score"] and original document metadata under metadata["source_metadata"].
Agentic retrieval is instead exposed by langchain-aws as the standalone agentic_retrieve function. With response generation enabled, the service returns a grounded answer and citations as well as retrieved chunks. The function is limited to Amazon Bedrock Managed Knowledge Bases; it is not a flag that can be switched on for AmazonKnowledgeBasesRetriever.
Why AWS separates complex questions
AWS illustrates the limitation of a single search with a query comparing checkout and inventory services across on-call escalation, backup and restore targets, and deployment rollback procedure. That creates six sub-intents. In the walkthrough’s corpus, retrieving five chunks—10% of the corpus—covered four of six intents and missed checkout on-call information and inventory restore information. Increasing the result count to 10, or 19% of the corpus, covered all six, although two sub-intents appeared twice and one chunk contributed no evidence.
That example is a walkthrough result, not a general performance benchmark. Its point is that increasing a standard retrieval result count can restore coverage but may introduce redundant context, whereas a planning loop can split the question into smaller retrieval tasks and decide whether further evidence is needed.
AWS also reports an evaluation on the public MuSiQue multi-hop benchmark. The company says agentic retrieval improved recall over single-shot retrieval, with the largest gains on the hardest questions; the reported improvement on single-hop questions was under five points. AWS’s practical recommendation follows that trade-off: use standard Retrieve for short, well-scoped requests, and consider agentic retrieval for multi-part, comparative or exploratory queries.
Tracing the planner and its limits
The convenience helper returns final chunks but does not expose the plan. To observe it, AWS directs developers to call Boto3’s agentic_retrieve_stream on the bedrock-agent-runtime client. Boto3 1.43.32 or later is required because earlier releases did not include agentic_retrieve_stream; the walkthrough also specifies Python 3.12 or later, langchain-aws 1.6.3 or later, and LangChain 1.0 or later.
Trace events identify SpeculativeRetrieval, Planning, Retrieval and, when a passage needs more context, FullDocumentExpansion. They carry IN_PROGRESS, SUCCEEDED or FAILED statuses. The final result event is separate from those steps and contains chunks deduplicated across iterations; a chunk retrieved by several sub-queries can still appear repeatedly in the traces. AWS advises logging the nested sub-query text at attributes.actions[].retrieve.inputQuery.text, not merely the step and status.
AWS recommends leaving maxAgentIteration at its default of 5. The setting accepts values from 2 through 10, but AWS says settings of 2 or 3 run one cycle without sub-query decomposition, effectively producing single-shot behavior at agentic-retrieval cost. Decomposition begins at 4, and the planner may finish before the configured ceiling when it considers its evidence sufficient.
There are output differences to account for. Standard Retrieve returns a typed relevance score for each chunk; agentic results instead provide content, metadata and source-retriever information without an equivalent typed score. Applications that rank or filter context by score therefore need a different approach after switching APIs.
Integration, security and guardrails
To use agentic retrieval inside a LangChain Expression Language chain, AWS wraps the function in a RunnableLambda that extracts text from its results. Developers can let the service generate the answer, or turn service generation off and supply the retrieved context to their own prompt and model. AWS notes that passing raw Document objects into a prompt renders their representation and adds metadata noise, so standard-retrieval documents should be formatted first.
The walkthrough separates the knowledge base’s service role from the AWS STS caller identity used to query it. The service role needs access to the S3 document source and embedding model. The caller needs distinct retrieval and model permissions. In particular, AWS says bedrock:GetDocumentContent is needed if FullDocumentExpansion requests a whole document; a policy limited to bedrock:Retrieve can fail partway through a query.
Both paths support Amazon Bedrock Guardrails, but their interfaces differ. Agentic retrieval uses policyConfiguration.bedrockGuardrailConfiguration and supports only BLOCK mode, while the LangChain retriever uses guardrail_config. AWS identifies reliance on MASK mode as a reason to remain on the standard API.
Cost and routing implications
AWS says standard Retrieve is cheaper and lower latency because it is one call, works with self-managed knowledge bases and leaves answer generation under application control. Agentic retrieval makes several model invocations, has higher latency and costs more per call. It can also register up to five knowledge bases in one request and route sub-queries using a natural-language description, a capability AWS says Retrieve does not provide.
The recommended production pattern is query-shape routing: send routine, single-intent traffic to the lower-cost path and reserve the planner for requests likely to benefit from decomposition. Teams should also budget for document storage and ingestion, retrieval calls and foundation-model inference. Ingestion is asynchronous, and AWS advises polling for a terminal state rather than waiting a fixed interval. Resources and stored documents should be removed after an experiment because they can continue to incur storage charges.
Source: AWS Machine Learning Blog
Definition. Agentic retrieval plans sub-queries, evaluates retrieved evidence and can retrieve again to improve coverage of complex questions.
| Standard Retrieve | Agentic retrieval |
|---|---|
| One hybrid search returning scored chunks | Plans sub-queries, assesses evidence and can retrieve again |
| Lower cost and latency | Higher cost and latency from several model invocations |
| Works with self-managed knowledge bases | Limited to Amazon Bedrock Managed Knowledge Bases |
| Returns a typed relevance score per chunk | Does not provide an equivalent typed score |
| Suitable for routine, well-scoped requests | Best reserved for multi-part, comparative or exploratory queries |
Key takeaways
- Standard Retrieve performs one hybrid search and is the lower-cost, lower-latency path for short, well-scoped requests.
- Agentic retrieval can decompose comparative, multi-part or exploratory questions into sub-queries and assess whether more evidence is needed.
- AWS says agentic retrieval costs more and has higher latency because it makes several model invocations.
- Agentic retrieval is limited to Amazon Bedrock Managed Knowledge Bases, while standard retrieval also works with self-managed knowledge bases.
- Use Boto3's agentic_retrieve_stream to inspect planner traces, including nested sub-query text.
- A caller may need bedrock:GetDocumentContent when FullDocumentExpansion requests an entire document.
FAQ
When does AWS recommend agentic retrieval?
AWS recommends considering it for multi-part, comparative or exploratory queries that are likely to benefit from sub-query decomposition.
When should standard Retrieve be used?
AWS recommends standard Retrieve for short, well-scoped and routine single-intent requests because it is cheaper and lower latency.
Can agentic retrieval be enabled on AmazonKnowledgeBasesRetriever?
No. AWS exposes it as the standalone langchain-aws agentic_retrieve function rather than a switch on AmazonKnowledgeBasesRetriever.
How can developers inspect the agentic retrieval plan?
AWS directs developers to Boto3's agentic_retrieve_stream on the bedrock-agent-runtime client, which emits trace events.
Why can a retrieval-only permission policy fail with agentic retrieval?
If FullDocumentExpansion is needed, AWS says the caller requires bedrock:GetDocumentContent; a policy limited to bedrock:Retrieve can fail during a query.