What Is API Rate Limiting?
Rate limiting is a technique that controls how many requests an API client can make within a defined time window. When a client exceeds the limit, the server rejects additional requests until the window resets.
Rate limiting serves multiple purposes:
- Abuse prevention. Without limits, a single client could exhaust server resources.
- Cost control. Caps the cost exposure from runaway clients or scripts.
- Fair usage. Prevents one heavy consumer from degrading the experience for others.
- DoS protection. A basic layer of defense against denial-of-service attacks.
Rate Limiting Strategies
Fixed Window
Count requests in fixed time intervals. The counter resets at the start of each window. Example: 1,000 requests per hour.
Problem: A client can burst requests at the boundary of two windows, effectively doubling the allowed rate briefly.
Sliding Window
Tracks requests over a rolling period. The count at any moment is the sum of requests in the last N seconds. This prevents the burst problem of fixed windows.
Token Bucket
Tokens accumulate in a bucket at a defined rate, up to a maximum capacity. Each request consumes one token. This allows short bursts up to the bucket capacity while enforcing a long-term average rate.
Leaky Bucket
Requests are added to a queue and processed at a fixed rate. If the bucket overflows, new requests are rejected. This enforces a strict constant request rate.
HTTP Headers for Rate Limit Communication
| Header | Meaning |
|---|---|
X-RateLimit-Limit | Maximum requests allowed per window |
X-RateLimit-Remaining | Requests remaining in the current window |
X-RateLimit-Reset | Unix timestamp when the window resets |
Retry-After | Seconds until the client can retry |
When a client exceeds the limit, return HTTP 429 (Too Many Requests) with a Retry-After header.
Implementation in Node.js
Using the rate-limiter-flexible library with Redis:
typescriptimport { RateLimiterRedis } from "rate-limiter-flexible"; import redis from "./redis"; const limiter = new RateLimiterRedis({ storeClient: redis, keyPrefix: "api_rate_limit", points: 100, // requests duration: 60, // per 60 seconds }); export async function rateLimitMiddleware(req, res, next) { try { const key = req.headers["x-api-key"] || req.ip; const result = await limiter.consume(key); res.set("X-RateLimit-Remaining", result.remainingPoints); next(); } catch (rejRes) { res .status(429) .set("Retry-After", Math.ceil(rejRes.msBeforeNext / 1000)) .json({ error: "Too many requests. Please slow down." }); } }
Conclusion
API rate limiting is a necessary component of any production API. Token bucket and sliding window algorithms are the most practical choices. Use Redis for distributed rate limiting and communicate limit state to clients via standard response headers.