Postman has documented the architecture behind Agent Mode, its AI-native interface for testing, documentation, discovery and implementation, as it serves a community of 40 million developers on Amazon Bedrock. The implementation guide gives agent builders a concrete design pattern: limit the tools and context a model receives for each task, while treating inference routing, data controls and caching as production concerns.

Postman says the central challenge was adapting a product that has evolved over 11 years for an agent that reasons over data rather than navigates tabs, sidebars and other interface state. Agent Mode works directly against the application, but users must approve actions that modify application state. Enterprise administrators can also enable Amazon Bedrock Guardrails to redact personally identifiable information before it reaches the underlying large language model.

Tool catalogs need to be scoped

Postman initially favored small, atomic tools for actions such as opening a request, changing a field or retrieving metadata. That approach helped with control, but lengthy workflows became slow because each result had to return to the model before it selected the next action.

In Postman’s testing, tool-selection errors rose once the visible catalog exceeded about 40 tools. The reported failures included nonexistent tool calls, invalid arguments despite valid schemas, and contextually wrong selections. Larger or newer models reduced the problem but did not remove it, according to the company.

The resulting design retrieves tools by need and isolates execution threads. In the example described by Postman, a root agent queries a vector database of tool embeddings, reduces a catalog of more than 170 tools to roughly 15 relevant ones, and passes those to a context-isolated sub-agent. The company is also decoupling agent actions from user-interface state: Agent Mode can send a request in the background without an open tab, although approval is still required for state-changing work.

Structured data can replace one-off read tools

For its API Catalog, Postman consolidated narrow views into a query tool that uses ClickHouse table schemas. The agent can generate queries with joins and filters across structured data such as service uptime, test results and endpoint response times, rather than relying on a dedicated read tool for each possible question.

The trade-off is to invest in a well-modeled data layer instead of continually expanding the tool catalog. Postman presents schema-aware read access as a way to support a broader range of analysis while keeping the model’s tool context smaller.

Context, rather than capability, caused more failures

Postman says incomplete context caused more failures than missing tools. The agent needs an accurate view of the user’s location in the product, active entities and established state; without it, a correct tool can still be ineffective.

Sending the application’s existing rendering and data-transfer objects to the model was not sufficient, the company found. Postman instead built dedicated handlers that distill each entity into information useful for reasoning. Its design combines broad, automatically gathered background context that is minimized for the prompt with deeper, user-selected context processed by entity-specific handlers.

That approach still faces a context-budget constraint. Request descriptions, OpenAPI specifications and request payloads can contain open-ended user-generated material that consumes the context window. Postman says it is exploring a filesystem-backed approach intended to reduce the need for each handler to implement its own truncation and expansion logic.

Retrieval and Bedrock controls support the runtime

Agent Mode combines client-side action tools, generic agent instructions and a retrieval-augmented knowledge base. Postman seeded the knowledge base with concise feature-specific Learning Center articles, then selects articles at runtime using the query and available context. Selecting a mock server, for example, can automatically add the related article instead of placing the product’s full knowledge base in a static prompt.

For inference, Postman uses Amazon Bedrock to access supported Anthropic Claude models and route work by its latency, quality and cost requirements. It says moving among supported models is primarily a configuration change, and that newer and larger models reduced tool hallucinations in its testing.

Postman also uses Bedrock cross-Region inference for bursty developer traffic. Geographic profiles route work among supported Regions within a defined geography, while global profiles may use supported destination Regions worldwide when a workload does not require a geographic processing boundary. The selected profile, IAM and service-control policies, and quotas must permit every Region that Bedrock could choose.

A geographic profile does not mean inference runs in Postman’s own AWS environment; it constrains processing to the eligible Bedrock Regions in that profile. AWS states that Bedrock does not use prompts and completions to train AWS models or distribute them to third parties. Postman says it has configured data_retention_mode to none for supported Agent Mode models, though retention availability and behavior depend on the model and must be checked for each production deployment.

Two cache tiers separate stable and variable prompts

Postman uses Bedrock prompt caching because an agent repeatedly sends system instructions, agent behavior, core tools, selected knowledge and conversation context. Its near-immutable core uses a one-hour checkpoint, while more variable context uses a five-minute checkpoint that refreshes on a cache hit. Bedrock requires the longer-lived checkpoint to appear before the shorter-lived one.

The company characterizes production inference as a routing-and-caching problem as well as a model-selection decision. Cache benefits and supported time-to-live settings remain model-dependent; Postman recommends examining cache-read and cache-write token fields and measuring time to first token for the workload itself.

This is a vendor-published implementation account, and the reported tool-selection threshold and model observations are Postman’s own testing results. Its production implementation is proprietary and not available as a public sample repository. Source: AWS Machine Learning Blog

Definition. Postman Agent Mode is an AI-native interface that uses scoped tools, distilled context and Amazon Bedrock runtime controls to work across testing, documentation, discovery and implementation.

Prompt content tierPostman cache checkpoint
Near-immutable coreOne-hour checkpoint
More variable contextFive-minute checkpoint that refreshes on a cache hit

Key takeaways

  • Postman reduced a catalog of more than 170 tools to roughly 15 relevant tools for a described sub-agent workflow.
  • The company says incomplete context caused more failures than missing tools.
  • Schema-aware queries over structured ClickHouse data can replace many narrow read tools.
  • State-changing actions still require user approval.
  • Bedrock model routing is selected around latency, quality and cost requirements.
  • Postman separates near-immutable and variable prompt content into one-hour and five-minute cache checkpoints.

FAQ

Why does Postman scope tool catalogs for Agent Mode?

Postman says tool-selection errors rose when the visible catalog exceeded about 40 tools. Its design retrieves tools by need and gives a reduced set to a context-isolated sub-agent.

How does Postman handle context for Agent Mode?

It combines minimized automatically gathered background context with deeper user-selected context handled by entity-specific processors, rather than sending existing rendering and data-transfer objects directly to the model.

How does Postman use Amazon Bedrock prompt caching?

Its near-immutable prompt core uses a one-hour checkpoint, while more variable context uses a five-minute checkpoint that refreshes on a cache hit; the longer-lived checkpoint must appear first.

Can Agent Mode change application state without approval?

No. Postman says users must approve actions that modify application state, including when actions are decoupled from an open user-interface tab.

Sources