429, Retry After, and Runbooks: API Rate Limiting for Engineers

API rate limiting, or request throttling, caps how many calls a client can make over a fixed window. The recommended posture is to enforce per-principal limits at the gateway, return HTTP 429 with a Retry-After header, and require clients to implement backoff. Done right, it protects system stability, keeps usage fair across callers, and controls infrastructure cost.
TL;DR:
- Implement layered rate limits at the edge, gateway, and service level to prevent different failure modes and internal bottlenecks.
- Use token bucket algorithms at the gateway for burst tolerance, while leaky bucket or sliding window methods better protect internal resources and prevent boundary loopholes.
- Enforce per-identity, per-endpoint, and per-tier limits with clear response headers like 429 and Retry-After to improve user experience and facilitate self-throttling.
- Apply adaptive limits based on real-time latency and queue signals, and simulate load scenarios to prevent over-reliance on static thresholds.
- Distinguish between short-term velocity limits and long-term quotas to ensure proper system protection without confusing billing or usage policies.
Table of Contents
- Why rate limiting matters for stability, security, and cost
- Common rate-limiting algorithms and when to use each
- Where to enforce limits and multi-layer patterns
- Design considerations and practical rules
- Client-side retry strategies and safe backoff
- Testing, observability, and tuning rate limits
- Implementation examples and short snippets
- Operational perspective and the Nullbit runbook checklist
- Distinguishing between rate limiting and quota management
- Handling burst traffic and spikes effectively
- Rate limiting in distributed systems and microservices architectures
- Security considerations related to rate limiting
- Impact of rate limiting on user experience and design strategies to mitigate negative effects
- Common mistakes engineering teams make with rate limiting
- How Nullbit helps you put this into practice
- Sources
- FAQ
Why rate limiting matters for stability, security, and cost
A single misbehaving client, a retry storm, or a traffic spike can cascade through a system in seconds. When one service slows down, callers retry, that retry volume fans out to downstream dependencies, and the whole chain saturates. Rate limiting breaks that chain before it starts.
Unbounded traffic also translates directly into cost: compute, database connections, and third-party API calls all scale with request volume, and a scraping bot or credential-stuffing script can burn through a budget meant for real users. OWASP’s guidance on unrestricted resource consumption maps specific endpoint risks, login, signup, search, and public read endpoints each need different velocity limits because attackers target them differently.
- Retry storms turn a slow dependency into a full outage through fan-out amplification.
- Unbounded traffic drives infrastructure and third-party API costs up without warning.
- Login and signup endpoints face credential stuffing and need tighter per-identity limits than public reads.
- Limits should be tied to p99 latency targets, not set arbitrarily.
Common rate-limiting algorithms and when to use each
Each algorithm trades accuracy for state and complexity differently, and the right choice depends on where you enforce it. Microsoft’s throttling pattern guidance frames the decision around burst tolerance versus implementation cost.
- Token bucket: tokens refill at a steady rate and callers can spend a burst up to the bucket size, making it a good fit for APIs that expect occasional spikes from normal usage.
- Leaky bucket: requests queue and drain at a constant rate, smoothing bursts into steady output, useful when downstream systems need predictable load.
- Fixed window: counts requests in a fixed time slice, simple to implement but vulnerable to boundary bursts where a client sends double its quota across a window edge.
- Sliding window: tracks requests across a rolling interval for more accurate limits at the boundary, at the cost of more state to maintain.
Gateways handling many tenants often lean on token bucket for its burst tolerance, back-end services protecting a database favor leaky bucket for its steady drain, and high-burst public endpoints frequently need sliding window to close the fixed-window loophole.
Where to enforce limits and multi-layer patterns
A single rate limiter at one layer leaves gaps. Layered enforcement catches different failure modes at each boundary.
- Edge, CDN, or WAF: filters volumetric traffic and obvious bot floods before they reach application infrastructure.
- API gateway: enforces per-key quotas, applies standardized response headers, and is the natural place for per-principal and per-endpoint policy.
- Service level: throttles at fan-out points and around resource-bound components like databases or search indexes, where a gateway limit alone would not protect an internal bottleneck.
- Outbound limits: protect downstream and third-party dependencies from your own service’s retry behavior.
Each layer answers a different question: is this traffic volumetric abuse, is this caller over its contract, or is this internal component about to saturate. Skipping any one layer means one class of failure has no defense.
Design considerations and practical rules
The details of what you return and how you structure policy determine whether rate limiting protects your system or just annoys your users. Microsoft’s Azure Architecture Center throttling pattern draws a clear line: return HTTP 429 when a caller has exceeded its own contract, and reserve 503 for service-level constraints where the system itself is overloaded regardless of who is asking.
- Return 429 for per-principal or per-key limit breaches, and 503 when the service itself cannot handle any more load.
- Include a
Retry-Afterheader on every throttled response, as RFC 6585 specifies for the 429 status code. - Where possible, expose standardized
RateLimitandRateLimit-Policyheaders so clients can self-throttle before hitting the wall. - Build policy around per-principal, per-endpoint, and per-tier limits rather than one global bucket that punishes every caller for one bad actor.
- Make rejection cheap: reject before expensive work starts, and instrument every rejection path so you can see who is hitting limits and why.
- Consider adaptive limits that tighten or loosen based on live latency and queue-depth signals rather than a fixed number set once and forgotten.
The IETF rate-limit-headers draft defines a standard format for communicating remaining quota, such as RateLimit: "default";r=0;t=30, giving client libraries a consistent way to back off without parsing custom headers.
Client-side retry strategies and safe backoff
A client that retries a 429 immediately just makes the problem worse. Google Cloud’s retry guidance recommends truncated exponential backoff with jitter as the default pattern for handling transient errors.
- Wait an initial delay, then double it on each subsequent failure up to a configured maximum delay.
- Add random jitter to each wait so multiple clients retrying the same failure do not synchronize into another spike.
- Only retry on transient errors (429, 408, and 5xx responses) and only for operations that are safe to repeat.
- Cap the number of attempts and set a total deadline so a failing call gives up instead of retrying indefinitely.
- Log every retry attempt and its outcome so retry volume itself becomes a visible signal.
Pro Tip: Never retry a non-idempotent write without a dedicated idempotency key, or a retried request can double-charge, double-create, or double-send.
Testing, observability, and tuning rate limits
You cannot tune a limit you cannot see. Track requests per second or minute per principal, p50, p95, and p99 latency, and the rate of 429 and Retry-After responses your system issues.
- Load-test the rejection path itself, not just the happy path, to confirm throttled responses stay fast under pressure.
- Simulate retry storms deliberately to see how your backoff and jitter settings behave under real synchronized load.
- Chaos-test downstream failures to verify that limits actually prevent fan-out rather than just delaying it.
- Run an observe, decide, act loop with a named owner responsible for adjusting limits when saturation signals fire, an approach Microsoft’s throttling pattern frames as a control loop rather than a static gate.
- Tune by comparing measured p99 latency against your SLO and adjusting window size or per-tier weight from there, not from guesswork.
Implementation examples and short snippets
A gateway policy, a sample response, and a bit of pseudocode cover most of what a team needs to get started. Azure API Management, for instance, supports rate-limit and rate-limit-by-key policies that set a call count and renewal period and can expose remaining calls through response headers.
A throttled response typically looks like a 429 status with a Retry-After header set to a number of seconds, alongside RateLimit and RateLimit-Policy headers naming the active limit. Token bucket logic itself is simple: on each request, check if a token is available, consume it if so, and refill tokens at a fixed rate on a timer. The real decision is where that bucket’s state lives.
| Backend | Consistency | State cost | Typical latency impact |
|---|---|---|---|
| Local in-memory counter | Per-instance only | Low | Minimal |
| Redis shared counter | Consistent across instances | Moderate | Small added round trip |
- Local counters are fast but let each instance drift from the others under load balancing.
- A shared store like Redis keeps limits accurate cluster-wide at the cost of a network hop per check.
Operational perspective and the Nullbit runbook checklist
Throttling incidents move fast, and the teams that recover quickest have a rehearsed loop: observe dashboards tracking top callers and saturation metrics with tools like an AI Crawlability Audit, decide using a clear prioritization matrix for which traffic to shed first, and act through targeted blocking, temporary limit changes, or traffic shaping. Runbooks should be rehearsed in advance, DRIs trained on the tooling, and logs built to mask personally identifiable information from the start. Our anomaly detection case study reflects this same observe-decide-act structure applied to infrastructure monitoring.
Distinguishing between rate limiting and quota management
Rate limiting and quota management solve related but distinct problems, and conflating them leads to policies that fit neither purpose well. Rate limiting governs velocity: how many requests a caller can make within a short, rolling window, measured in requests per second or minute. Its job is to protect system stability in real time, smoothing bursts and rejecting traffic that would otherwise overwhelm a component right now.
Quota management governs volume over a much longer period, typically a day, month, or billing cycle, and it is usually tied to a commercial agreement or a fair-use policy rather than to system protection. A customer on a plan that includes 10,000 calls a month is bound by a quota, while the same customer might also be limited to 50 requests per second by a rate limit that has nothing to do with their plan tier.
The two mechanisms often live in different parts of the stack. Rate limits are commonly enforced at the gateway with fast in-memory or Redis-backed counters because they need to reject in milliseconds. Quotas are frequently tracked in a billing or account system, checked less frequently, and enforced with a softer response, sometimes a warning or a feature downgrade rather than an outright rejection.

Confusing the two produces two common mistakes. Teams sometimes implement only a quota and assume it protects against a burst, when a client can exhaust an entire monthly quota in a single minute and still crash the system before the quota check even triggers a block. The reverse mistake is enforcing only a rate limit and having no visibility into which customers are approaching their commercial usage caps, leading to billing disputes or unexpected plan upgrades that surprise the customer. Both controls belong in a complete API strategy, addressing different time scales and different owners.
Handling burst traffic and spikes effectively
Legitimate bursts, a marketing launch, a batch job kicking off, or a mobile app reconnecting after a network drop, look identical to abusive traffic at the network level. The difference is in how you absorb them rather than how you detect them.
Token bucket algorithms exist specifically for this case: by allowing a caller to spend accumulated tokens in a short burst while still enforcing a steady average rate, they let normal spiky usage through without opening the door to sustained abuse. Setting the bucket size larger than a single request but smaller than a client’s maximum plausible batch gives you burst tolerance without unlimited exposure.

Queueing is another lever. Rather than rejecting every request above a threshold outright, a system can queue excess requests briefly and drain them at a controlled rate, which works well for asynchronous or batch-friendly workloads but poorly for latency-sensitive user-facing calls where a queued response arrives too late to matter.
Autoscaling complements rate limiting rather than replacing it. A limit protects the system during the seconds or minutes before autoscaling can add capacity, and once new instances are online, temporarily relaxed limits can let traffic through at the new, higher capacity. The two need to be designed together: a rate limit that never adjusts as capacity grows just becomes an artificial ceiling.
Finally, communicate proactively where you can. An API that expects a predictable spike (a product launch, a scheduled batch window) can pre-provision higher per-tenant limits for known high-volume clients ahead of time rather than discovering the mismatch through a wave of 429 responses during the event itself.
Rate limiting in distributed systems and microservices architectures
A single-instance rate limiter is straightforward: keep a counter in memory and check it on every request. Distributed systems break that model immediately, because a request can land on any of dozens of instances behind a load balancer, and a counter local to one instance has no idea what the others have already allowed.
The common fix is a shared state store, most often Redis, that every instance checks and updates atomically. This keeps limits accurate across the fleet but introduces a network round trip on every request and makes the store itself a dependency that needs its own availability guarantees; if the shared counter goes down, you need a defined fallback, whether that is failing open, failing closed, or falling back to a degraded local limit.

Microservices architectures add another layer of complexity because a single user-facing request often fans out into several internal calls. A rate limit applied only at the edge gateway will not protect an internal service that gets hit disproportionately hard by one particular downstream call pattern. This is why layered enforcement, discussed earlier for edge, gateway, and service tiers, matters even more in a microservices context: each service needs its own protection against its own bottlenecks, not just a shared perimeter check.
Service meshes have made this more manageable by pushing rate limiting into a sidecar or mesh-level policy that applies consistently across services without each team reimplementing the logic. Whether enforcement happens in a mesh, a shared library, or a central gateway, the underlying requirement is the same: limits need consistent, low-latency access to shared state, and every service in the call chain should assume its neighbors might reject requests too.
Security considerations related to rate limiting
Rate limiting is one of the most practical defenses against denial-of-service attempts, but its security value goes well beyond stopping raw volumetric floods. OWASP’s API Security guidance recommends layering limits by IP, by authenticated identity, and by endpoint, because each layer defeats a different attack pattern.
Per-IP limits catch unsophisticated scripted attacks but do little against distributed attacks that rotate through thousands of IP addresses, sometimes via botnets or residential proxy networks. Per-identity limits catch a compromised or malicious account regardless of which IP it connects from, which matters for credential stuffing attacks where an attacker cycles through leaked username and password pairs against a login endpoint. Per-endpoint limits recognize that a login form, a password reset flow, and a public search endpoint each carry different risk profiles and need different velocity thresholds.
Rate limiting also reduces the blast radius of automated scraping and enumeration attacks, where an attacker walks through sequential IDs or common usernames looking for valid accounts. A tight per-identity or per-session limit on these endpoints slows enumeration to a pace that is no longer economically useful for an attacker.
It is worth being explicit that rate limiting is a mitigation, not a complete defense. It slows and constrains abuse rather than eliminating it, and it works best alongside authentication, input validation, and anomaly detection rather than as a standalone control. A sudden spike in 429 responses from a particular identity or IP range is itself a useful signal worth feeding into a broader security monitoring pipeline, since it often precedes a more targeted attack attempt.
Impact of rate limiting on user experience and design strategies to mitigate negative effects
A rate limit that a legitimate user never notices is well designed, and one that blocks a normal workflow is a bug wearing a security feature’s clothing. The gap between the two usually comes down to communication and threshold design rather than the underlying algorithm.
Exposing remaining quota through standard headers, as described in the IETF rate-limit-headers draft, lets well-built client applications warn users or slow their own request rate before a hard rejection happens, turning a jarring error into a smooth degradation. A mobile app that can show “approaching your request limit” because it read a RateLimit header gives its user a far better experience than one that just shows a generic failure after a 429.
Error messaging matters as much as the header. A 429 response with a clear Retry-After value and a human-readable message lets a well-built client tell its user exactly when to try again, instead of leaving them to guess or hammer the retry button. Vague or missing error bodies push users toward exactly the retry behavior that made the rate limit necessary in the first place.
Tiering limits by account type also shapes experience directly. A free tier with a visibly generous but finite limit, paired with a clear upgrade path when a user approaches it, reads as a product decision rather than a punishment. The same numeric limit presented as an opaque wall, with no explanation and no path forward, reads as a bug report waiting to happen.
Finally, consider the cost of false positives. An overly aggressive per-IP limit can catch legitimate users behind a shared corporate NAT or a university network, throttling many innocent people because of a few heavy users sharing the same visible IP. Layering identity-based limits alongside IP-based ones, as covered earlier, protects against this specific failure mode.
Common mistakes engineering teams make with rate limiting
The most common failure is the single global bucket: one limit applied to all traffic regardless of caller or endpoint, which means one noisy tenant degrades service for everyone else sharing that bucket. A close second is hiding downstream 429 responses behind a generic 500, which strips the client of any signal about what actually happened and guarantees a blind retry.
Rate limiting is not always the right tool. When the real problem is insufficient capacity rather than unfair usage, a bulkhead pattern or autoscaling addresses the root cause better than a tighter limit ever will. Throttling belongs in capacity planning and incident runbooks from the start, not bolted on after the first outage.
— Matija
How Nullbit helps you put this into practice
Getting rate limiting right across a gateway, a fleet of services, and a handful of downstream dependencies is a system engineering problem, not a single configuration file. Nullbit’s System Engineering and Cloud & Infrastructure services cover exactly this kind of architecture review and gateway policy implementation, and our AI Automation work has included anomaly detection integration similar to the observe-decide-act loop described above, drawn from projects like our infrastructure anomaly detection case study.

Whether you need a focused audit of your existing limits or a full implementation across gateway and service layers, our Agile Engagement and Fixed-Price Project options cover both. Get in touch to scope the work against your own traffic patterns and SLOs.
Sources
- Throttling pattern - Microsoft Azure Architecture Center
- API4:2023 Unrestricted Resource Consumption - OWASP API Security
- Retry strategy | Vertex AI / Google Cloud Documentation
- IETF rate-limit-headers draft
FAQ
How do I fix an API rate limit reached error?
Slow down and retry using truncated exponential backoff with jitter rather than retrying immediately, and check the response for a Retry-After header to know exactly how long to wait. If the limit reached is a monthly or account quota rather than a short-term rate limit, you will need to wait for the reset period or request a higher tier.
What is a good rate limit for an API?
There is no universal number since it depends on your infrastructure capacity, endpoint sensitivity, and client tier, which is why OWASP recommends setting limits per endpoint based on that endpoint’s specific risk and resource cost. Set limits by measuring your p99 latency against your service level objective and adjusting from there rather than picking a round number upfront.
How do I rate limit an API to 10 requests per minute?
Configure a token bucket or fixed-window counter with a limit of 10 units per minute at your gateway or rate-limiting middleware, refilling or resetting the counter on that interval. Return a 429 status with a Retry-After header once the tenth request in the window is exceeded, and consider a sliding window instead of fixed window if boundary bursts (two batches of 10 sent just before and after a window edge) are a concern.





