Backend Development7 min readOct 1, 2025

Understanding API Rate Limiting: Why It Matters and How to Implement It

Rate limiting protects APIs from abuse, controls costs, and ensures fair usage across consumers. This guide explains rate limiting strategies, implementation patterns, and how to communicate limits to API clients.

UL

Ullass Engineering Team

Ullass — Software Development & Digital Products

What Is API Rate Limiting?

Rate limiting is a technique that controls how many requests an API client can make within a defined time window. When a client exceeds the limit, the server rejects additional requests until the window resets.

Rate limiting serves multiple purposes:

  • Abuse prevention. Without limits, a single client could exhaust server resources.
  • Cost control. Caps the cost exposure from runaway clients or scripts.
  • Fair usage. Prevents one heavy consumer from degrading the experience for others.
  • DoS protection. A basic layer of defense against denial-of-service attacks.

Rate Limiting Strategies

Fixed Window

Count requests in fixed time intervals. The counter resets at the start of each window. Example: 1,000 requests per hour.

Problem: A client can burst requests at the boundary of two windows, effectively doubling the allowed rate briefly.

Sliding Window

Tracks requests over a rolling period. The count at any moment is the sum of requests in the last N seconds. This prevents the burst problem of fixed windows.

Token Bucket

Tokens accumulate in a bucket at a defined rate, up to a maximum capacity. Each request consumes one token. This allows short bursts up to the bucket capacity while enforcing a long-term average rate.

Leaky Bucket

Requests are added to a queue and processed at a fixed rate. If the bucket overflows, new requests are rejected. This enforces a strict constant request rate.

HTTP Headers for Rate Limit Communication

HeaderMeaning
X-RateLimit-LimitMaximum requests allowed per window
X-RateLimit-RemainingRequests remaining in the current window
X-RateLimit-ResetUnix timestamp when the window resets
Retry-AfterSeconds until the client can retry

When a client exceeds the limit, return HTTP 429 (Too Many Requests) with a Retry-After header.

Implementation in Node.js

Using the rate-limiter-flexible library with Redis:

typescript
import { RateLimiterRedis } from "rate-limiter-flexible"; import redis from "./redis"; const limiter = new RateLimiterRedis({ storeClient: redis, keyPrefix: "api_rate_limit", points: 100, // requests duration: 60, // per 60 seconds }); export async function rateLimitMiddleware(req, res, next) { try { const key = req.headers["x-api-key"] || req.ip; const result = await limiter.consume(key); res.set("X-RateLimit-Remaining", result.remainingPoints); next(); } catch (rejRes) { res .status(429) .set("Retry-After", Math.ceil(rejRes.msBeforeNext / 1000)) .json({ error: "Too many requests. Please slow down." }); } }

Conclusion

API rate limiting is a necessary component of any production API. Token bucket and sliding window algorithms are the most practical choices. Use Redis for distributed rate limiting and communicate limit state to clients via standard response headers.

Related Articles

Ullass — Software Development Company

We build web apps, SaaS platforms, and digital products

Ullass designs, engineers, and scales software products for ambitious businesses. Also try our free online tools at tools.ullass.com. Questions? hello@ullass.com