Ctrl + K
API27 min read

Rate Limiting Explained

Understand how API rate limiting works, why it matters, which algorithms are commonly used, and how to design and handle rate limits in modern applications.

Published: 2026-10-05

Rate limiting is a mechanism that controls how many requests a client can make to an API or other service during a given period. Instead of allowing a client to send an unlimited number of requests, the server defines a limit such as 100 requests per minute and rejects or delays requests that exceed it.

Rate limiting is important for both security and reliability. It can protect an API from accidental request floods, abusive clients, brute-force attempts, excessive resource consumption, and traffic spikes. It also helps distribute shared resources more fairly when many users or applications access the same service.

A good rate-limiting strategy is more than simply returning HTTP 429 when too many requests arrive. The server must decide what is being limited, how limits are counted, which algorithm is appropriate, how limits are communicated to clients, and what should happen when a limit is reached.

What Is Rate Limiting?

Rate limiting restricts the frequency at which a client can perform an operation. The operation is commonly an HTTP request, but a limit can also apply to a specific endpoint, user action, API key, IP address, account, or resource.

For example, an API might allow a client to make 60 requests per minute. If the client sends 40 requests during the first 30 seconds and another 20 during the next 10 seconds, it has reached the configured limit. Additional requests may receive a 429 Too Many Requests response until the client is allowed to continue.

The important distinction is that rate limiting usually controls request frequency rather than the total number of requests a client can ever make. A limit of 100 requests per minute resets or replenishes according to the algorithm being used.

LimitMeaning
10 requests per secondUp to 10 requests can be processed within the configured rate interval.
100 requests per minuteThe client is limited to a defined request rate over approximately one minute.
10,000 requests per dayThe client has a larger quota intended for daily usage control.
5 login attempts per minuteA specific operation receives a stricter limit than general API traffic.

Why Do APIs Use Rate Limiting?

Without rate limiting, a single client can potentially consume a disproportionate amount of server capacity. This can happen because of a programming mistake, an aggressive crawler, a retry loop, malicious traffic, or simply an application that generates more requests than the service was designed to handle.

Protecting Server Resources

Every API request consumes resources. Depending on the endpoint, processing a request can involve CPU time, memory, database queries, network bandwidth, external API calls, file operations, or expensive computations.

Rate limiting provides a control mechanism that prevents individual clients from continuously consuming an unreasonable share of those resources.

Preventing Abuse

Rate limits can make automated abuse more difficult. They are commonly used around authentication endpoints, password reset requests, verification codes, search endpoints, expensive operations, and public APIs.

Rate limiting does not replace authentication, authorization, bot detection, or other security controls. Instead, it is one layer in a broader defense strategy.

Handling Traffic Spikes

Even legitimate users can generate sudden traffic spikes. For example, a frontend bug could accidentally send the same request repeatedly. A rate limiter can reduce the impact of such behavior while the underlying problem is investigated.

Providing Fair Resource Usage

When an API serves many clients, rate limits can prevent one client from consuming most of the available capacity. Limits can be configured per user, API key, organization, IP address, or another identity depending on the application's architecture.

HTTP 429 Too Many Requests

The standard HTTP response used when a client has exceeded a rate limit is 429 Too Many Requests. The status code tells the client that the server has received too many requests within a given period and that the client should reduce its request rate.

HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 30

{
  "error": "rate_limit_exceeded",
  "message": "Too many requests. Please try again later."
}

The Retry-After header can tell the client how long it should wait before retrying. Depending on the HTTP specification and server implementation, Retry-After can communicate a delay in seconds or a specific HTTP date.

💡 A 429 response should contain enough information for a well-behaved client to back off instead of immediately sending the same request again.

Rate Limit Headers

APIs often expose rate-limit information through HTTP response headers. This allows clients to understand the configured limit, how much capacity remains, and when they may be able to retry.

HTTP/1.1 200 OK
RateLimit-Limit: 100
RateLimit-Remaining: 42
RateLimit-Reset: 30

The exact headers used by an API depend on its design and the rate-limiting infrastructure. A client should follow the API documentation rather than assuming that every service exposes the same header names or semantics.

When a request is rejected, the response can also expose the remaining limit or reset information where appropriate. Retry-After is especially useful because it gives clients an explicit retry delay.

What Exactly Should Be Rate Limited?

Before choosing an algorithm, an API needs to decide what constitutes a client. There is no universally correct identifier because different applications have different security and usage requirements.

KeyTypical use
IP addressPublic endpoints where no authenticated identity is available.
User IDAuthenticated applications where limits should follow individual accounts.
API keyDeveloper and server-to-server APIs.
OrganizationMulti-tenant systems where an entire organization shares a quota.
EndpointExpensive operations that need stricter limits than ordinary requests.
CombinationSystems that need several layers of protection at the same time.

IP-based limiting is easy to implement but can be inaccurate. Multiple legitimate users may share an IP address through NAT, a corporate network, a mobile carrier, or a proxy. Conversely, an attacker can sometimes distribute traffic across multiple IP addresses.

Authenticated APIs can often use user IDs, API keys, or organization IDs as the primary rate-limit key. Many production systems combine identity-based and network-based limits instead of relying on a single identifier.

Fixed Window Rate Limiting

The fixed window algorithm divides time into fixed intervals. For example, a server might allow 100 requests during each one-minute window. A counter is associated with the current window, and each request increments that counter.

Limit: 100 requests
Window: 60 seconds

12:00:00 - 12:00:59 -> up to 100 requests
12:01:00 - 12:01:59 -> counter starts again

Fixed windows are simple and inexpensive. A counter can be stored in memory or in a distributed store such as Redis, with the key expiring at the end of the window.

The main weakness is the boundary problem. A client could send many requests near the end of one window and another large batch immediately after the next window begins. Although the client technically respects both individual windows, the server can still receive a large burst within a short period.

Advantages of Fixed Windows

  • Simple to understand and implement.
  • Low storage and computational overhead.
  • Easy to explain to API consumers.
  • Works well for many basic quota systems.
  • Can be implemented efficiently with expiring counters.

Disadvantages of Fixed Windows

  • Allows bursts around window boundaries.
  • Does not represent request distribution within the window.
  • Can produce uneven traffic patterns.
  • May be too coarse for sensitive endpoints.

Sliding Window Rate Limiting

A sliding window evaluates requests within a continuously moving time interval rather than resetting the counter at a fixed boundary. For example, with a limit of 100 requests per minute, the server can consider the requests that occurred during the previous 60 seconds at the moment each new request arrives.

This reduces the boundary problem found in fixed windows because there is no artificial reset that suddenly makes a large number of requests available.

A straightforward sliding-window implementation can store timestamps for individual requests. For high-volume systems, storing every timestamp may become expensive, so optimized approaches often use buckets or approximate counters.

Advantages of Sliding Windows

  • Provides smoother request-rate control.
  • Reduces fixed-window boundary bursts.
  • Can represent recent traffic more accurately.
  • Useful when request distribution matters.

Disadvantages of Sliding Windows

  • More complex than a simple fixed counter.
  • Exact implementations may require more storage.
  • Distributed implementations need careful synchronization.

Token Bucket Algorithm

The token bucket algorithm is one of the most widely used approaches for rate limiting. Imagine a bucket that can contain a maximum number of tokens. Tokens are added to the bucket at a fixed rate. Each request consumes one or more tokens.

If enough tokens are available, the request is allowed. If the bucket is empty, the request is rejected or delayed. Because tokens can accumulate up to the bucket capacity, the algorithm naturally allows controlled bursts while still enforcing an average rate.

Bucket capacity: 100 tokens
Refill rate: 10 tokens/second
Request cost: 1 token

A request can be processed while at least 1 token is available.
Unused capacity accumulates up to the bucket's maximum.

Token bucket is useful when an API should allow short bursts but should not permit sustained unlimited traffic. For example, a client might normally make a few requests per second but occasionally need to send several requests immediately.

Token Bucket Parameters

ParameterPurpose
CapacityMaximum number of tokens that can accumulate.
Refill rateNumber of tokens added over time.
Request costNumber of tokens consumed by an operation.
Initial tokensNumber of tokens available when the bucket starts.

A useful property of token buckets is that different operations can have different costs. A cheap endpoint might consume one token, while an expensive report-generation operation could consume ten.

Leaky Bucket Algorithm

The leaky bucket algorithm models traffic as a queue that is processed at a controlled rate. Incoming requests enter the queue, and the system processes them at a configured rate. If the queue becomes full, additional requests can be rejected.

This makes the leaky bucket particularly useful when the goal is to smooth traffic rather than simply count requests. Instead of allowing a burst to reach the backend immediately, requests can be released at a more consistent rate.

A token bucket generally emphasizes allowing controlled bursts while maintaining an average rate. A leaky bucket is often associated with smoothing or queueing traffic. Real-world implementations can vary, so the exact behavior depends on the infrastructure and configuration.

Token Bucket vs Leaky Bucket

CharacteristicToken bucketLeaky bucket
Primary ideaRequests consume accumulated tokens.Requests enter a queue processed at a controlled rate.
BurstsNaturally supports controlled bursts.Typically smooths bursts through queueing.
Traffic patternControls average rate with burst capacity.Produces a more consistent processing rate.
Typical useAPI request limiting and service quotas.Traffic shaping and controlled processing.

Concurrency Limits vs Rate Limits

A rate limit and a concurrency limit solve different problems. A rate limit controls how frequently requests can arrive over time. A concurrency limit controls how many operations can be running at the same time.

For example, an API might allow 100 requests per second but permit only 20 expensive requests to execute concurrently. This prevents a large number of long-running operations from exhausting server resources even when the request rate itself is within the configured limit.

The two mechanisms can be combined. Rate limiting controls traffic entering the system, while concurrency controls how much work is allowed to execute simultaneously.

Rate Limiting by Endpoint

Not every endpoint consumes the same amount of resources, so applying one global limit to every operation is often too simplistic.

Endpoint typePossible strategy
Health checkHigh limit because processing is cheap.
Read-only APIModerate limit based on expected usage.
SearchStricter limit because queries can be expensive.
File generationLow limit or concurrency control because operations are resource intensive.
LoginVery strict limit to reduce automated abuse.
Password resetStrict limit combined with account and IP controls.

Endpoint-specific limits are especially useful when a small number of expensive operations could otherwise consume the same resources as many inexpensive requests.

Global and Per-Client Limits

A production API may use multiple rate-limit layers. A per-user limit can prevent one account from dominating resources, while a global limit can protect the entire service from an overall traffic spike.

For example, an API could enforce a limit per API key, a stricter limit for expensive endpoints, and a global infrastructure limit. The exact combination depends on the service's architecture and capacity.

💡 Layered limits are often more practical than trying to create one perfect global number. Different limits can protect different resources and abuse scenarios.

Rate Limiting and Authentication

Authentication changes how a server can identify a client, but it does not automatically solve rate limiting. An authenticated API can apply limits to user IDs, API keys, organizations, or service accounts.

Unauthenticated endpoints often need additional IP-based or device-related controls because there may be no reliable account identifier. Login and registration endpoints deserve particular attention because attackers can generate large numbers of attempts without successfully authenticating.

A useful security design is to apply different limits before and after authentication. For example, a login endpoint can have a limit per IP address as well as a limit associated with the attempted account.

Rate Limiting and Retries

Clients commonly retry failed HTTP requests. This becomes dangerous when the retry mechanism ignores rate limits. A client that immediately retries every 429 response can create an even larger request storm.

async function requestWithBackoff(
  request: () => Promise<Response>,
  maxRetries = 3
): Promise<Response> {
  for (let attempt = 0; attempt <= maxRetries; attempt++) {
    const response = await request();

    if (response.status !== 429) {
      return response;
    }

    if (attempt === maxRetries) {
      return response;
    }

    const delay = 1000 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delay));
  }

  throw new Error("Request failed");
}

Exponential backoff increases the delay between retries. A production client can additionally respect Retry-After when the server provides it.

Retry behavior should also use jitter in systems with many clients. Without jitter, many clients that receive the same response at approximately the same time may retry simultaneously, creating another synchronized traffic spike.

How to Handle a 429 Response

A client receiving 429 should normally avoid immediately sending the same request again. It should inspect Retry-After if available, wait for an appropriate period, and then retry according to the API's documented behavior.

If Retry-After is not provided, the client can use an exponential backoff strategy with a reasonable maximum delay. Repeated 429 responses should not result in an unbounded retry loop.

⚠️ Do not treat HTTP 429 as a generic server error that should be retried immediately. Doing so can make the rate-limit condition worse and increase load on the API.

Distributed Rate Limiting

Rate limiting becomes more complicated when an application runs on multiple server instances. If every instance keeps its own in-memory counter, a client can potentially exceed the intended global limit simply by having requests distributed across different instances.

For example, if an API has five application servers and each server independently allows 100 requests per minute, a client could potentially receive a much higher effective limit than 100 requests per minute if traffic is distributed between instances.

A shared data store can provide a common source of rate-limit state. Redis is frequently used because counters, expiration, atomic operations, and fast access are important for many rate-limiting designs.

Distributed rate limiting must account for atomicity and race conditions. Two requests arriving at nearly the same time should not both observe the same remaining capacity and then incorrectly exceed the limit.

Rate Limiting with Reverse Proxies and API Gateways

Rate limiting does not always need to be implemented directly inside application code. Reverse proxies, API gateways, load balancers, CDNs, and managed infrastructure can enforce limits before requests reach the application.

Enforcing some limits at the edge can reduce the amount of unwanted traffic that reaches application servers. Application-level limits can then provide more context-aware controls based on authenticated users, endpoint cost, business rules, or database state.

A common architecture uses infrastructure-level protection for broad traffic control and application-level protection for business-specific limits. The appropriate split depends on the platform and operational requirements.

Choosing the Right Rate-Limit Algorithm

There is no single algorithm that is best for every API. The choice depends on whether the system needs strict quotas, controlled bursts, smooth traffic, low implementation complexity, or accurate recent-history tracking.

AlgorithmGood fitMain trade-off
Fixed windowSimple APIs and straightforward quotas.Boundary bursts.
Sliding windowMore accurate recent request-rate control.Higher complexity or storage requirements.
Token bucketAPIs that should allow controlled bursts.Requires careful capacity and refill configuration.
Leaky bucketTraffic smoothing and queue-based processing.Queueing can increase latency or require additional capacity.

For many general-purpose APIs, token bucket or a sliding-window approach provides a useful balance between protection and usability. Fixed windows remain attractive when simplicity and low overhead are more important than precise traffic smoothing.

How to Choose a Rate Limit

A rate limit should be based on actual system capacity and expected client behavior rather than an arbitrary number. Start by understanding how expensive the endpoint is, how frequently legitimate clients need to call it, and how much traffic the underlying infrastructure can handle.

For example, a simple metadata endpoint might tolerate a much higher request rate than an endpoint that performs a complex database query or invokes an external AI service.

It is also useful to distinguish normal usage from burst capacity. A client may legitimately need to send several requests quickly, even if its long-term average request rate is relatively low. Token bucket parameters can model this distinction explicitly.

Rate Limits for Public APIs

Public APIs should document their rate limits clearly. Developers need to know the expected request rate, what identity the limit applies to, what response indicates that the limit has been exceeded, and how they should retry.

If different plans or API keys have different limits, the documentation should make those differences clear. Clients should not have to discover the limit through repeated 429 responses.

Clear limits also make client-side optimization easier. Developers can batch requests, cache responses, reduce unnecessary polling, and schedule background work more efficiently when they understand the API's constraints.

Rate Limiting and Caching

Caching can reduce the number of requests that need to reach an application and therefore reduce pressure on rate limits. If a client repeatedly requests the same data and the response can safely be cached, serving cached data can be more efficient than executing the same backend operation repeatedly.

However, caching and rate limiting solve different problems. A cache reduces repeated work, while a rate limiter controls request frequency. A cached endpoint can still receive excessive traffic, so caching should not be treated as a replacement for rate limiting.

Rate Limiting and Polling

Polling is a common source of unnecessary API requests. A frontend might repeatedly request the same resource every few seconds even when its state has not changed.

Where appropriate, clients can reduce polling frequency, use conditional requests, cache responses, or use server-sent events and WebSockets when real-time communication is actually required. Rate limits can protect the server, but good client design reduces the need to hit those limits in the first place.

Rate Limiting Expensive Operations

Some operations are expensive enough that a simple request count is not a good representation of resource usage. Generating a report, processing an uploaded file, running a complex search, or invoking a costly external service may consume significantly more resources than a basic GET request.

In these cases, weighted rate limiting can assign different costs to different operations. A lightweight operation could consume one unit while an expensive operation consumes ten or more.

Another option is a separate quota for the expensive operation. This can make the system easier to reason about when different resources need independent protection.

Rate Limiting for Authentication Endpoints

Authentication-related endpoints often require stricter controls than ordinary API requests. Login, password reset, email verification, and one-time-code endpoints can be targeted by automated systems that generate large numbers of requests.

A single IP limit may not be sufficient because attackers can distribute requests across many addresses. Conversely, an account-based limit alone can make it possible to affect legitimate users sharing an account or identity.

Authentication protection often combines several signals and controls, such as IP-based limits, account-based limits, progressive delays, anomaly detection, and additional verification. The exact design depends on the threat model.

Common Rate-Limiting Mistakes

Using Only IP Addresses

IP addresses are useful but are not always reliable identities. Multiple users can share an address, while a single attacker can use multiple addresses. For authenticated APIs, an account, API key, or organization identifier may provide a better primary key.

Applying One Limit Everywhere

A single global limit ignores the fact that different endpoints have different costs and security requirements. Authentication and expensive operations often need much stricter limits than simple read endpoints.

Ignoring Distributed Deployments

Independent counters on multiple application servers can result in a much higher effective limit than intended. Distributed systems need shared or coordinated rate-limit state when the limit is supposed to apply across instances.

Not Telling Clients When to Retry

Returning 429 without useful retry information makes it harder for clients to behave correctly. Retry-After and documented rate-limit information can significantly improve client behavior.

Retrying Immediately After 429

Immediate retries can create a feedback loop where the client continuously hits the limit. Exponential backoff, jitter, and Retry-After handling are better approaches.

Choosing Arbitrary Limits

A limit that is too strict can break legitimate applications, while a limit that is too generous may provide little protection. Limits should be based on measured resource usage, expected traffic, and acceptable capacity.

Testing Rate Limiting

Rate limiting should be tested deliberately rather than only during production traffic. A test environment can generate requests at different rates and verify that requests are allowed and rejected according to the configured algorithm.

Useful test cases include requests exactly at the limit, requests just above the limit, bursts, requests around window boundaries, concurrent requests, multiple clients, and requests after the reset or refill period.

Distributed systems should also be tested with multiple application instances. This helps reveal race conditions and inconsistencies that may not appear when the rate limiter is running on a single server.

Tools such as an HTTP Request Builder and REST API Mock Generator can help create repeatable requests and test how an API responds when request limits are reached.

Observability and Rate Limits

A rate limiter should be observable. Metrics can show how often requests are rejected, which endpoints generate the most throttling, which clients consume the most capacity, and whether legitimate users are frequently hitting the limit.

Useful metrics include the number of 429 responses, requests per client, requests per endpoint, rate-limit utilization, rejection rates, and the distribution of traffic over time.

Logging should be designed carefully. Rate-limit logs can contain identifiers such as IP addresses, user IDs, or API keys. Sensitive credentials should never be written to logs in plain text.

A Practical Rate-Limiting Strategy

A practical implementation can start with a small number of clearly defined rules instead of attempting to solve every possible abuse scenario at once.

  • Identify the resources and endpoints that need protection.
  • Decide which client identity should be used for each limit.
  • Measure normal request patterns and resource consumption.
  • Choose an algorithm appropriate for the traffic pattern.
  • Define separate limits for expensive or security-sensitive operations.
  • Return HTTP 429 when a request exceeds the configured limit.
  • Provide Retry-After or equivalent rate-limit information where appropriate.
  • Implement client retries with backoff and jitter.
  • Use shared state when limits must work across multiple servers.
  • Monitor rejected requests and adjust limits based on real usage.

Example API Response

A useful rate-limit error should be machine-readable and should provide a stable error identifier. The exact response format depends on the API's conventions.

{
  "error": {
    "code": "rate_limit_exceeded",
    "message": "Too many requests.",
    "retryAfter": 30
  }
}

The response should not require clients to parse human-oriented text to determine what happened. A stable error code makes it easier for client applications to handle rate-limit responses programmatically.

Rate Limiting Checklist

  • Define what resource or operation is being protected.
  • Choose an appropriate client identity.
  • Determine whether the limit is global, per client, per endpoint, or layered.
  • Choose between fixed window, sliding window, token bucket, or leaky bucket.
  • Decide whether controlled bursts should be allowed.
  • Consider concurrency limits for long-running operations.
  • Return HTTP 429 when the rate is exceeded.
  • Provide retry information when possible.
  • Make clients respect Retry-After and use backoff.
  • Use shared state for distributed rate limiting when required.
  • Monitor 429 responses and adjust limits based on real traffic.
  • Document limits clearly for API consumers.

Frequently Asked Questions

What is API rate limiting?

API rate limiting controls how frequently a client can send requests to a service. A server defines a limit, such as 100 requests per minute, and rejects or delays requests that exceed the configured rate.

What HTTP status code is used for rate limiting?

HTTP 429 Too Many Requests is the standard status code used when a client has exceeded a rate limit. The response can include Retry-After or other rate-limit information to help the client determine when to retry.

What is the difference between rate limiting and throttling?

The terms are often used interchangeably, but throttling can more broadly refer to controlling or slowing traffic, while rate limiting specifically defines a permitted request rate. The exact terminology varies between systems and documentation.

Which rate-limiting algorithm should I use?

It depends on the API's requirements. Fixed windows are simple, sliding windows provide more continuous request tracking, token buckets allow controlled bursts, and leaky buckets are useful for smoothing traffic. The correct choice depends on traffic patterns and resource constraints.

Should rate limits be based on IP addresses?

IP addresses can be useful, especially for unauthenticated endpoints, but they are not always reliable client identities. Multiple users can share an IP address and attackers can distribute traffic across several addresses. Authenticated APIs can often use user IDs, API keys, or organization identifiers as additional rate-limit keys.

How should a client handle HTTP 429?

The client should avoid immediately retrying the request. It should respect Retry-After when provided and otherwise use an appropriate exponential backoff strategy, preferably with jitter to avoid synchronized retries.

Can rate limiting be implemented across multiple servers?

Yes. Distributed applications commonly use shared state or infrastructure-level rate limiting so that all application instances enforce a consistent limit. A local in-memory counter on each server is insufficient when the limit is intended to apply globally.

Helpful API Tools

When working with rate-limited APIs, an HTTP Request Builder can help create repeatable requests for testing headers, status codes, and request parameters. An HTTP Header Viewer is useful for inspecting response headers such as Retry-After and rate-limit metadata, while an HTTP Status Codes Lookup helps verify the meaning of 429 responses and related HTTP statuses. For controlled API testing, a REST API Mock Generator can simulate endpoints and different response scenarios, and an HTTP Response Formatter can make raw API responses easier to inspect.

Conclusion

Rate limiting is a fundamental part of reliable API design. It protects server resources, reduces the impact of abusive or accidental traffic, helps distribute capacity between clients, and provides a predictable boundary for API consumption.

The implementation should be based on the actual behavior of the service. Fixed windows are simple, sliding windows provide more continuous tracking, token buckets support controlled bursts, and leaky buckets can smooth traffic. In larger systems, rate limits can be combined across clients, endpoints, authentication states, and infrastructure layers.

A robust rate-limiting design also includes clear 429 responses, useful retry information, appropriate client backoff, distributed coordination where necessary, and monitoring of real-world traffic. The goal is not simply to reject requests, but to keep the API predictable and available while allowing legitimate clients to use it effectively.

Found an issue?

Found an error, outdated information, or something missing from this article? Let me know through the Contact page.

Your feedback helps improve our articles and keep them accurate and useful.