API strategy
Agent infrastructure
Token bucket vs sliding window rate limiting compared, with working Redis implementations, distributed rate limiter pitfalls, and how to pick per endpoint.

By the end of this guide you'll know which rate limiting algorithm fits your API, how to implement it in a distributed setup, and how to expose limits agents can actually read and respect. We'll cover token bucket, sliding window log, and sliding window counter — the three algorithms that show up in every serious production discussion.
Prerequisites: a working API, a shared data store (Redis or equivalent), and a rough sense of your traffic shape. Time required: about an hour to work through, longer to ship end to end.
This guide assumes you're throttling API calls at the edge or in middleware. If you're deep into API rate limiting best practices for SaaS in the agent era, you already know why the algorithm choice matters more when the caller is an agent bursting through your endpoints.
Before you pick, know what you're picking between. The three algorithms behave differently under burst, differently under sustained load, and differently under the failure modes that hit production.
Token bucket: a bucket holds N tokens. Each request removes one. Tokens refill at a fixed rate. When the bucket is empty, requests are rejected or queued. Bursts up to the bucket size are allowed; sustained throughput is capped at the refill rate.
Sliding window log: keep a timestamped log of every request in the last window (say, 60 seconds). Count entries. If the count exceeds the limit, reject. Exact but memory-heavy.
Sliding window counter: divide time into fixed buckets (say, per-second). Keep counts per bucket. On each request, sum the counts in the last N buckets and weight the partial bucket at the boundary. Approximate, cheap, and close enough for most cases.
If you take one thing from this table: token bucket permits bursts, sliding window log forbids them, sliding window counter sits in the middle.
One question decides most of this: do you want to allow bursts?
If your API is called by agents that batch tool calls, or by clients that run cron jobs on the minute, you'll see traffic that's spiky by design. Token bucket handles that gracefully — the bucket absorbs the spike, and sustained load is still capped by the refill rate. Reject a legitimate burst and you'll be answering support tickets by lunch.
If you're enforcing a hard quota — 1,000 requests per hour, no negotiation — sliding window log gives you the exact answer. Compliance tiers, paid quotas, and anything a customer's finance team will audit belong here.
If you want a good default for a general-purpose public API, sliding window counter is the right pick. It approximates the sliding log at a fraction of the memory cost, and the approximation error is invisible to normal callers.
Agent traffic is worth a special note. Agents are the most bursty legitimate callers most APIs will see — an agent chains dependent tool calls and can fire ten requests in a second, then nothing for a minute. Token bucket suits them. Fixed-window counters (the fourth algorithm nobody should still be using) do not.
Here's a working token bucket using Redis and a Lua script for atomicity. The script updates the bucket and returns whether the request is allowed in a single round trip.
-- token_bucket.lua
-- KEYS[1] = bucket key
-- ARGV[1] = capacity, ARGV[2] = refill_rate (tokens/sec)
-- ARGV[3] = now (unix ms), ARGV[4] = requested tokens
local capacity = tonumber(ARGV[1])
local refill_rate = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
local requested = tonumber(ARGV[4])
local bucket = redis.call('HMGET', KEYS[1], 'tokens', 'last')
local tokens = tonumber(bucket[1]) or capacity
local last = tonumber(bucket[2]) or now
local elapsed = math.max(0, now - last) / 1000
tokens = math.min(capacity, tokens + elapsed * refill_rate)
local allowed = 0
if tokens >= requested then
tokens = tokens - requested
allowed = 1
end
redis.call('HMSET', KEYS[1], 'tokens', tokens, 'last', now)
redis.call('EXPIRE', KEYS[1], math.ceil(capacity / refill_rate) * 2)
return {allowed, tokens}
Call it from your API middleware, keyed by whatever identifies the caller — API key, user ID, or (for agents) the delegated user token subject.
Expected result: the script returns {1, remaining} when the request is allowed and {0, remaining} when it isn't. remaining is what you'll put in the response headers in Step 6.
Sliding window counter is almost as cheap as token bucket and closer to what most product managers describe when they say "100 requests per minute." Here's the working version.
-- sliding_window.lua
-- KEYS[1] = key prefix
-- ARGV[1] = limit, ARGV[2] = window_seconds, ARGV[3] = now (unix seconds)
local limit = tonumber(ARGV[1])
local window = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
local current_bucket = math.floor(now / window)
local previous_bucket = current_bucket - 1
local current_key = KEYS[1] .. ':' .. current_bucket
local previous_key = KEYS[1] .. ':' .. previous_bucket
local current = tonumber(redis.call('GET', current_key)) or 0
local previous = tonumber(redis.call('GET', previous_key)) or 0
local elapsed_in_current = now - (current_bucket * window)
local weight = 1 - (elapsed_in_current / window)
local estimated = current + (previous * weight)
if estimated >= limit then
return {0, math.floor(limit - estimated)}
end
redis.call('INCR', current_key)
redis.call('EXPIRE', current_key, window * 2)
return {1, math.floor(limit - estimated - 1)}
The key trick: the previous bucket is weighted by how far into the current window we are. This smooths the boundary that catches out fixed-window counters, where a client can send 2 × limit requests in two seconds by straddling the boundary.
A distributed rate limiter is only as consistent as its coordination layer. Three things go wrong when teams first ship one.
Node-local counters that don't share state. If you rate-limit per-node with in-memory counters, ten nodes with a "100/minute" limit will happily serve 1,000 requests per minute in aggregate. You need a shared store — Redis, DynamoDB, or a purpose-built system like Envoy's global rate limit service.
Race conditions on read-modify-write. Read the counter, check the limit, write the increment — done in three round trips, this races. Every scheme in this guide uses a Lua script or an atomic operation for a reason. Don't split it.
Clock skew. Sliding window counters use timestamps. If your API nodes disagree on the time by more than a few hundred milliseconds, you'll see off-by-one errors at bucket boundaries. Use the Redis server's time (redis.call('TIME')) rather than the app node's clock for the now argument.
For context on why coordinated failures matter more when the caller is an agent, see agent tool granularity. A single agent turn can generate ten calls, and inconsistent limiting across nodes means the agent sees non-deterministic 429s.
Agents and well-behaved clients back off correctly only if you tell them what the limits are. Return them on every response, not just on 429s.
RateLimit-Limit: 100
RateLimit-Remaining: 47
RateLimit-Reset: 23
This three-header format is the widely deployed de facto convention — inherited from early drafts of the IETF RateLimit header fields work and shipped by GitHub, Twitter, and many SDKs. Note that the current IETF draft has moved to a single structured field, e.g. RateLimit: limit=100, remaining=47, reset=23 alongside RateLimit-Policy: 100;w=60. If you're building fresh, consider emitting both so old and new clients both understand you. On 429 responses, add Retry-After with the number of seconds until the caller should retry.
For agents specifically, structured error responses matter more than headers alone. A 429 with a machine-readable body — see the pattern in API error responses for AI agents — lets the calling agent decide whether to back off, switch tools, or surface the failure to the user.
Unit tests catch the arithmetic. Load tests catch the algorithm choice.
Run three scenarios against your implementation:
2 × limit requests in one second, then nothing for the rest of the window. Token bucket accepts up to capacity, rejects the rest. Sliding window counter rejects roughly half.limit requests in the last second of one window and limit more in the first second of the next. Fixed-window counters fail this. Sliding window counter smooths it. Sliding window log rejects the second batch entirely.Record which requests were rejected and what the response headers said. If a client following the Retry-After header would still get a 429 on retry, your limiter is lying — fix the header calculation before shipping.
Rate limiting at the wrong key. Limiting by IP address breaks behind NAT and corporate proxies. Limiting by API key breaks when agents share a service account. Limit by the identity that actually maps to the quota — usually the authenticated user, or the delegated subject on the agent's token. See AI agent authorization for how to get the identity right first.
Fixed windows in disguise. "Requests per calendar minute" resets at :00 and lets a caller send 2 × limit in two seconds across the boundary. If your window resets on a clock tick rather than sliding, you have a fixed window, not a sliding one. Fix it or accept the burst.
Silent Redis failures. If your Redis call times out, do you fail open or fail closed? Fail open and you serve unlimited traffic during an incident. Fail closed and a Redis blip takes your API down. Pick deliberately — most teams fail open with a circuit breaker and an alert.
No visibility into who's hitting the limit. Log every 429 with the caller identity, the endpoint, and the algorithm's state at rejection time. Without this, you can't tell abuse from a legitimate customer who needs a higher tier.
Assuming one algorithm fits every endpoint. Search endpoints tolerate bursts; write endpoints don't. It's fine to run token bucket on read-heavy routes and sliding window log on the routes that mutate state. Pick per endpoint if the traffic shape justifies it.
Stay up to date on the ever changing agentic landscape.