Guides

What Is an API Rate Limit?

Ingrid Solbergยทยท7 min read

A rate limit is not the API provider being stingy. It is the one thing standing between a working API and one that falls over the first time a buggy script hits it in a loop โ€” and sooner or later that script is yours. Understanding the limit is less about appeasing someone else's rules and more about making your own integration reliable.

What a rate limit actually is

An API rate limit is a cap on how many requests a single caller may make within a window of time. "100 requests per minute" is a rate limit. So is "5,000 per hour" or "10 per second." Every request you send is counted against your allowance for the current window; cross the line and the API stops serving you until the window resets.

The crucial detail is who the limit applies to. The API has to identify you to count your requests, and it does that by your API key, your access token, or โ€” if you send neither โ€” your IP address. That identity is the bucket your requests are tallied into, which is why two different API keys get two independent allowances. If you're fuzzy on how those credentials identify a caller, what an API key is and what a bearer token is both cover the mechanics.

A turnstile at a venue

Picture the entrance to a busy venue with a turnstile that only lets a set number of people through per hour. It isn't there to keep you out โ€” it's there so the room never fills past what's safe, and so no single tour group can shove through and block the door for everyone behind them. Each ticket gets counted as it passes; once the hour's quota is reached the turnstile locks, and when the next hour begins it opens again. An API rate limiter is that turnstile, your API key is the ticket it scans, and the time window is the hour that resets the count.

How the counting actually works

Under the hood, providers use a handful of standard models, and the main thing separating them is how they treat a sudden burst:

  • Fixed window โ€” count requests per clock interval (per calendar minute, say) and reset to zero at the boundary. Simple, but it allows a double-rate burst right across the reset line.
  • Sliding window โ€” track the trailing 60 seconds continuously, which smooths out that boundary spike.
  • Token bucket โ€” you hold a bucket of tokens that refills at a steady rate; each request spends one, and a full bucket lets you burst briefly before settling to the refill rate.
  • Leaky bucket โ€” requests queue and drain at a constant rate, enforcing a smooth, steady flow.

You rarely have to care which one a given API uses, but it explains why some APIs tolerate a short burst and others reject the eleventh request in a second flat.

How you'll know you hit one: 429

When you go over, a well-behaved API answers with HTTP 429 Too Many Requests. That's distinct from a 401 (your credentials are bad) or a 403 (you're not allowed) โ€” a 429 means you're authenticated just fine, you've simply knocked too often. Alongside it you'll usually get headers worth reading:

HTTP/1.1 429 Too Many Requests
Retry-After: 30
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1696430400

Retry-After is the API telling you, in seconds, exactly how long to wait. X-RateLimit-Remaining is your live fuel gauge, and X-RateLimit-Reset (often a Unix timestamp) is when the tank refills. Honoring these is the difference between a resilient client and one that gets temporarily banned for hammering a locked turnstile.

How to stay under the limit

  • Read the headers, don't guess. Watch X-RateLimit-Remaining and slow down as it approaches zero instead of discovering the limit by slamming into it.
  • Back off with jitter. On a 429, wait for Retry-After, and if you're running many clients, add a small random delay so they don't all retry in the same instant and cause a second stampede.
  • Cache and batch. The cheapest request is the one you never send. Cache responses you'll reuse, and combine operations into a single call wherever the API offers a batch endpoint.
  • Authenticate. Anonymous, IP-based tiers are almost always far stricter than the per-key quota you get once you send credentials. Generating a proper key โ€” you can create a strong one with the API key generator โ€” usually multiplies your allowance.
  • Spread scheduled work. If a cron job fires a thousand calls at the top of every hour, stagger them. A steady trickle stays under a limit that an instant flood would blow straight through.

Tokens expire, too

One failure that masquerades as a rate-limit problem: a flood of retries firing because the access token behind them has quietly expired, each retry bouncing and triggering another. If your requests start failing in bursts, check whether the token is still valid before you blame the limit โ€” you can inspect a token's claims with the JWT decoder, read its parts with the bearer token parser, and check its expiry with the JWT expiration checker. When a JWT expires covers the lifetime side in more depth.

The short version

A rate limit is a ticket-counting turnstile that keeps an API stable, fair, and affordable โ€” including for you. Treat the 429 and its headers as instructions rather than errors: wait the Retry-After, watch your remaining budget, cache what you can, authenticate for a bigger allowance, and spread your traffic out. Do that and you'll almost never see a 429 in the first place.

Try the tools

Frequently Asked Questions

What is an API rate limit in simple terms?

It's a cap the API provider puts on how many requests you're allowed to make in a given window of time โ€” say 60 requests per minute. Every request you send counts against that allowance; once you hit the cap, further requests are rejected until the window resets. The limit is tracked per caller, usually keyed to your API key, access token, or IP address, so your usage doesn't affect anyone else's allowance.

What does HTTP 429 Too Many Requests mean?

429 is the status code an API returns when you've exceeded its rate limit. It means your request was understood and you're authenticated fine โ€” you've simply sent too many requests too quickly. A well-designed API pairs the 429 with a Retry-After header telling you how many seconds to wait, so the correct response is to pause for that long and then retry, not to immediately hammer the endpoint again.

How do I avoid hitting a rate limit?

Read the rate-limit headers the API sends (X-RateLimit-Remaining and X-RateLimit-Reset) and slow down before you hit zero; cache responses you'll need again instead of re-fetching; batch multiple operations into one call where the API allows it; spread scheduled jobs out rather than firing them all at once; and always authenticate, since anonymous or IP-based tiers are usually far stricter than the per-key quota you get with credentials.

What's the difference between rate limiting and throttling?

They overlap and the terms are often used interchangeably. A rate limit is the rule โ€” the fixed ceiling, like 1,000 requests per hour. Throttling is a server's act of slowing or rejecting requests to enforce that rule, sometimes gradually (adding delay as you approach the limit) rather than cutting you off hard at the boundary. In practice, hitting the rate limit is what triggers throttling.

Why do APIs have rate limits at all?

Three reasons. First, stability: one client stuck in a retry loop could otherwise overwhelm the service and take it down for everyone. Second, fairness: shared capacity means one heavy user shouldn't be able to starve all the others. Third, cost and abuse control: every request consumes compute and bandwidth, and limits make denial-of-service attacks and scraping far more expensive to attempt.

IS

Ingrid Solberg writes for CodeUtilityKit, where the team builds free, privacy-first developer tools that run entirely in your browser. Every guide is written and reviewed by developers who use these tools daily.