Postman describes Agent Mode as an AI feature integrated into its API platform, with Amazon Bedrock supplying foundation-model inference. Its account outlines a production design that selects tools for each task, prepares different kinds of context, routes inference by workload and applies model-dependent retention and caching settings. The 40 million developers figure refers to Postman’s developer community, not to Agent Mode users.
What Agent Mode does
Postman is the API platform; Agent Mode is a feature within it that assists with API testing, documentation, discovery and implementation. In the described setup, Amazon Bedrock provides model inference, while Agent Mode coordinates Postman tools and workspace context.
The user-facing feature can help generate API tests from plain-language prompts, investigate failing requests and draft CI/CD workflow files. Suggested changes to tests can be reviewed and approved by the user.
Postman announced on March 31, 2026, that Claude on Amazon Bedrock had become Agent Mode’s default model provider. Its technical account describes selecting among supported Claude models by workload, so that announcement does not mean every request uses one model.
How Postman scopes tools to a task
Rather than show an agent every available tool at once, Postman describes a retrieval flow that searches a catalog of more than 170 tools and narrows it to about 15 relevant to a task. A context-isolated sub-agent then receives the selected tools.
Postman also reports that tool-selection errors increased in its testing when the visible set exceeded roughly 40 tools. That observation helps explain the design choice: the system searches for a smaller, task-specific set instead of exposing the full catalog to the agent at once.
How context and structured data reach the agent
Postman describes two context paths. Background context is gathered automatically from the workspace and reduced before it enters the prompt. Focused context is chosen by the user—for example, through selections or mentions—and passed through handlers tailored to the relevant type of Postman object.
That distinction matters when a workspace contains large request descriptions, OpenAPI specifications or payloads: Postman says these can crowd the model’s context window, making it necessary to reduce or truncate context. Broad background context and user-selected detail are therefore handled differently rather than treated as one undifferentiated prompt.
For structured information in API Catalog, Postman describes schema-aware queries against ClickHouse instead of maintaining a separate narrow tool for every question. A sample query uses a seven-day window, calculates p95 latency—the 95th-percentile value—and filters for results below 100 ms. Those are example query conditions, not measured service results.
Inference routing and user approval
Postman says it selects an Amazon Bedrock inference profile for each workload. Geographic profiles can route among supported Regions within a defined geography; global profiles can route among supported Regions worldwide.
Agent Mode can also send requests in the background without an open tab. Postman says actions that change application state still require user approval, so background processing does not remove that approval step.
Model-dependent retention and prompt caching
Postman says it sets the Bedrock data_retention_mode to none for supported Agent Mode models. Availability and behavior depend on the selected model, so the setting is not a universal rule for every model.
The described prompt-caching setup uses a one-hour checkpoint for stable prompt content and a five-minute checkpoint for more variable context. The longer-lived checkpoint must come before the shorter-lived one, and cache support varies by model.
Postman also says it uses Amazon Bedrock Guardrails to redact personally identifiable information before it reaches the large language model. Enterprise administrators can enable this in Agent Mode’s guardrail settings.
What the technical account establishes
Postman’s description offers a view of how one API-platform feature combines task-scoped tools, workspace context and managed model inference. It also identifies the practical choices that shape the design: limiting the tools presented for a task, trimming context, selecting inference routes and applying controls that vary by model.
Postman says its production implementation is proprietary and is not available as a public sample repository.