Platform integration

Agent infrastructure

AWS API deep-dive: what Amazon API Gateway does, and where it stops for agents

AWS API deep-dive: what Amazon API Gateway actually does across REST, HTTP, and WebSocket APIs, where the gateway model holds up for agents, and where it stops.

8 minute read
Decorative imagery showcasing Pontil's brand

The question this article answers: what is Amazon API Gateway actually doing when you put it in front of your services, and does that shape still fit when the caller is an AI agent rather than a mobile app or a partner backend?

Our view in one sentence: Amazon API Gateway is a strong request-plane product for the traffic it was designed for — humans, apps, and machine-to-machine partners hitting HTTP or WebSocket endpoints — and it does none of the work the tools layer of the agent stack has to do.

The next five sections cover: what the AWS API surface actually is once you stop reading the marketing page; how the three endpoint types (REST, HTTP, WebSocket) differ in ways that matter for agent workloads; where the gateway model holds up under agent traffic; where it stops short; and what the shape of an answer looks like if you're a SaaS team building agents on top of AWS.

What Amazon API Gateway actually is

Amazon API Gateway is a managed service that sits between callers and your backend. It terminates HTTP or WebSocket connections, authenticates the request, applies throttling and quotas, transforms payloads if you want, and forwards to a target — usually a Lambda function, an HTTP endpoint, or another AWS service via direct integration. On the way back it can transform the response, log the call to CloudWatch, and emit metrics.

That description covers the request plane. It does not cover contract design, tool definition, capability discovery, or anything about how a caller decides which endpoint to hit. Those are the caller's problem. That's the right split when the caller is a human developer writing code — they'll read your docs, generate an SDK, and wire it in. It stops being the right split when the caller is a model deciding at runtime which of your endpoints matches the user's intent.

Under the hood there are three flavours worth naming: REST APIs (the original, richest, most expensive), HTTP APIs (leaner, faster, cheaper, announced in 2019 and generally available in 2020), and WebSocket APIs (for stateful bidirectional channels). Same console, same billing surface, materially different products. Most teams pick one for the wrong reason — either the first tutorial they read or the cheaper of the two — and inherit constraints they don't discover until later.

REST APIs, HTTP APIs, and WebSocket APIs are not interchangeable

The naming is unhelpful. HTTP APIs also do REST. REST APIs also speak HTTP. The distinction is about feature depth and cost model, not protocol.
‍

REST API
HTTP API
WebSocket API

Launched

2015

2020 (announced 2019)

2018

Per-request cost

Higher

~70% cheaper than REST

Per-message + connection-minute

Request/response transforms

Yes

Limited

N/A

API keys and usage plans

Yes

No

No

Private endpoints (VPC)

Yes

Yes

No

WAF integration

Yes

No (via CloudFront)

No

Custom authorizers

Lambda, IAM, Cognito

Lambda, JWT*, IAM

Lambda, IAM

Best fit

Complex request shaping, monetised APIs

Straight-through auth + proxy

Streaming, chat, live updates


*HTTP APIs support Cognito user pools via the generic JWT authorizer, pointed at the Cognito issuer URL — there's just no dedicated Cognito authorizer type the way REST APIs have.

For an agent workload, the choice usually collapses to HTTP API unless you specifically need request transformation, API keys tied to usage plans, or WAF. Agents don't need most of what REST API's extra surface gives you, and they will pay for it at every tool call.

WebSocket APIs are a separate conversation. They're relevant if you're streaming model output through your own infrastructure or holding a persistent channel to a browser-based agent UI. They're not what the tool call itself flows over.

What the gateway does well when agents are calling

Credit where it's due. A lot of what AWS built into API Gateway is exactly what you want in front of agent traffic.

Auth termination. Lambda authorizers, JWT validators, and IAM auth all work the same whether the caller is a browser or a model. If you've already wired Cognito or your own OAuth provider into API Gateway for your existing customers, agent traffic slots into the same auth path. The identity boundary — which is the hard part of zero trust for AI agents — can be enforced at the same edge.

Throttling and quotas. Agents fan out. A single user request can turn into a dozen tool calls. REST API's per-key usage plans and both flavours' route-level throttling give you a place to defend the backend without writing rate-limit code in every Lambda. The mechanics still work; what changes is the shape of the limits, which we cover below.

Observability. CloudWatch metrics, access logs, and X-Ray tracing give you the request-side view of what agents are doing. You'll still need agent-side tracing to correlate tool calls to model turns, but the gateway side is solid.

Regional and private endpoints. For SaaS companies whose customers demand VPC-private access, both REST and HTTP APIs support private endpoints via VPC endpoints. That's a real answer to "the agent can't touch the public internet" requirements — and those requirements are showing up in enterprise deals now.

None of this is unique to AWS — Azure API Management, Kong, and Google Cloud API Gateway offer overlapping capabilities. But if you're already deep in AWS, the gateway is a sensible request-plane choice for agent-facing endpoints. It's not the wrong product. It's just not the whole answer.

Where the gateway model stops short for agents

Here's where a fair deep-dive has to make the harder point. API Gateway solves the request plane. Agent projects stall on things the request plane doesn't touch.

It doesn't decide what the tools should be. A gateway fronts endpoints you already have. If your product does 100 things through its UI and your API exposes 15 of them, the gateway happily serves those 15. It won't generate the other 85. This is the shape of the problem behind most stalled agent projects — the API layer exposes a small fraction of what the product can actually do — and no gateway product on any cloud addresses it.

It doesn't give the caller a way to discover capabilities. Agents need machine-readable descriptions of what a tool does, what arguments it takes, and what it returns. API Gateway can serve an OpenAPI spec at a URL. It doesn't help you make that spec agent-usable — which is a design problem, not a hosting problem. Descriptions matter, argument names matter, response shapes matter. Tool schema design is the real work, and it happens upstream of the gateway.

Per-user rate limits are awkward. API Gateway throttling is naturally per-API-key or per-route. Agent traffic wants per-authenticated-user limits — so a single misbehaving agent burning through one user's quota doesn't take down the tenant. You can get there with Lambda authorizers writing to a custom counter, but you're building infrastructure the gateway doesn't ship.

Identity flows through as a token, not as an actor. The gateway will happily forward a JWT to your Lambda. Whether that Lambda then executes as the authenticated user or as a shared service role is your problem. For AI agents, the answer has to be the former — agent identity vs user identity matters at every tool call — and the gateway isn't the layer that solves it.

Idempotency is your job. Agents retry. Models retry. Orchestrators retry. If the same tool call arrives twice, the gateway won't dedupe. API idempotency for AI agents is a design pattern you implement in the backend, not a checkbox in the gateway console.

Contract drift is invisible. API Gateway doesn't tell you when your handler's real response shape has diverged from the OpenAPI you published. That's a full-stack problem — OpenAPI spec drift breaks agents before it breaks humans — and the gateway isn't looking for it.

None of this is a criticism of API Gateway. It's a statement about what layer it operates at. Gateways front APIs; they don't create the tools agents need.

Cost, at agent traffic

A quick honest word on cost. API Gateway pricing looks reasonable at web-app volumes and gets uncomfortable at agent volumes. HTTP APIs at roughly $1 per million requests are fine. REST APIs at roughly $3.50 per million requests start to sting when a single user session generates fifty tool calls and each tool call generates two or three internal calls behind it.

The deeper problem isn't the per-request rate — it's that agent traffic is bursty and hard to forecast. A demo that costs pennies can become a $12,000 monthly line item once ten enterprise customers turn on the agent for their whole team. This is a broader point about API gateway pricing in the agent era: the per-request model was designed for a world where request volume was roughly proportional to human users. Agents break that assumption.

Budget with a multiplier. Whatever you estimate agent traffic will be based on user counts, multiply by five for tool fan-out and again by two for retries. If that number is uncomfortable, the answer is usually to reduce tool call count through better tool granularity, not to switch gateways.

How Pontil fits

We operate one layer up from the gateway. Amazon API Gateway is a good place to terminate agent traffic — auth, throttling, observability, private endpoints all work the way you'd want them to. What it doesn't do is generate the tools your agents need in the first place, keep those tools honest as your product changes, or execute them under the authenticated user's identity.

Pontil is the tools layer of the agent stack. We scan the systems you already own — your existing services, whatever shape their APIs are in — and generate tools your agents can actually invoke. Tools stay current as your product changes. Runtime executes as the real user, not a shared service account. If you're building agents on AWS, you'll keep the gateway. We sit behind it, doing the work it was never designed to do. If the gap between what your product can do and what your API exposes is the thing stalling your agent project, that's the layer we solve for.

What's the right question to ask about AWS and agents?

Not "is API Gateway good enough for agents?" — it's a strong request-plane product and will remain so. The better question is: what does your team need to build on top of the gateway to make your product actually reachable by an agent, and how much of that is work you want to own?

For a single service with a well-designed API, the answer might be "almost nothing" — publish the OpenAPI spec, wire up an authorizer, and let the model figure it out. For a SaaS company with a portfolio of products, patchy API coverage, and a two-year rewrite backlog, the answer is materially more work than the gateway page suggests. The gateway is the easy part.

The honest way to plan an agent project on AWS is to size the tools-layer work separately from the request-plane work, budget them separately, and not let the fact that API Gateway is well-understood convince you that the rest of the problem is well-understood too.

Join our weekly newsletter

Stay up to date on the ever changing agentic landscape.

POSTS

Related content

Platform integration

API strategy

Azure API Gateway: what Azure API Management actually does, and where it stops for agents

8 minute read

API strategy

Platform integration

API gateway vs iPaaS: which one fits the problem you actually have

7 minute read

Agent infrastructure

API strategy

AI gateway vs API gateway: what each one actually does, and when you need both

7 minute read