I set a per-IP rate limit on an API. It wasn't per-IP. Every visitor on the internet shared one counter, and I found that reading the code rather than from a bug report, which was luck. The first real symptom would have been the site refusing everybody at once while the single attacker I'd built it for carried on unaffected.
The mistake isn't exotic. It's the default outcome any time the thing enforcing a limit can't see who it's limiting.
The address on the socket isn't the visitor
Almost nothing reaches that API from a browser directly. Someone loads a page, the web tier renders it, and the web tier calls the API. So when the API asks the operating system who's on the other end of the connection, the honest answer is "the web server". On a managed platform it's usually a hop further out than that, and the peer your process sees is a router belonging to the host.
Every visitor in the world therefore turns up wearing the same address. Key a per-IP counter on it and you've built one bucket with everybody's name on it. The limit isn't too loose or too tight. It's aimed at the wrong thing: it throttles the crowd and never touches the individual, which is the precise inverse of the job.
The obvious fix hands over a perfect evasion
The standard answer is x-forwarded-for, the header a proxy adds to record who it is forwarding for. Read that instead of the socket and the numbers start looking sensible immediately.
It's also just a header, and anyone can send one. Put a different fake address on every request and your allowance never runs out, so the limit now stops exactly the people who weren't trying to get round it.
That's worse than having no limit, and not by a small margin. No limit is a gap you know about. A limit that reads as protection and stops nobody is a gap you've stopped looking at.
So which header can you believe?
Only one that an outsider can't set. Some edges add a client-IP header themselves and refuse any request that arrives already carrying one, Cloudflare's being the well documented case. Where that holds, the value reaching your process is the edge's word rather than the sender's.
Where it holds. Don't assume it, because it's two commands. Point them at the host that actually reads the header, which is the origin your application runs on, not the CDN in front of your website:
# 1. Does the edge reject a forged header, or pass it through?
curl -sI https://your-api.example/ -H 'CF-Connecting-IP: 1.2.3.4'
# 2. Is there any route to the origin that skips the edge?
curl -sI https://your-origin-host.example/ | grep -i serverThe first should be refused before it ever reaches your application. The second should show the same edge answering, because a design that assumes every request is inspected is only as good as the absence of a back door around it.
A 200 on the first one isn't a failure, by the way. It's an answer: whatever you just tested does not strip that header, so nothing behind it may believe the header. Run it against a CDN, or against anything not sitting behind that particular edge, and 200 is the only result you can get. It tells you about the host you aimed at and nothing else.
Re-run both whenever anything about the hosting changes, and write down what you saw. A measurement nobody recorded gets repeated as an assumption.
The request with no visitor behind it
Here's the part I hadn't considered at all. Some of the traffic hitting the API isn't a person and never was.
Server-rendered pages read their content by calling the API directly, not through the browser-facing route. There's no end user's address to forward, because there's no end user. It's a server drawing a page.
Lump that in with real traffic and your own rendering competes with your visitors for the same allowance. The busier the site gets the more of the limit it spends on itself, and when it finally trips, the failure looks exactly like an attack you aren't having. Your own web tier is not the threat the limit exists for, and it has to be told apart from one.
How you separate them is its own decision and it's the piece I'd think hardest about, because whatever proves "this request is my own server" immediately becomes a thing that must not leak.
What the whole design rests on
One assumption, and it belongs in a comment next to the code rather than in somebody's head: that the edge is not optional. Every request arrives through it, so the header it sets is something the sender doesn't control.
Move somewhere that isn't true and this becomes worse than the bug it replaced. Today an attacker can at most evade their own limit. With a trusted but forgeable header they can write someone else's address into it and spend that person's allowance instead. Denial of service aimed at one named victim, using your own protection to do it.
That's the kind of assumption that stays true right up until an infrastructure decision, taken for a completely unrelated reason, quietly stops it being true. Nobody will think to tell you. The commands above are cheap; the note saying why you ran them is what makes somebody run them again.
One API and one bug proves nothing general about rate limiting, and I'd be careful reading it as a pattern. What it does say is that "per IP" is a claim about what your process can actually see, and that's worth checking before you write the number into a config file.
Counting is the part that isn't optional
Five years in procurement before I wrote software for a living. Nobody there signs for a delivery because the paperwork says what arrived. You count what's on the floor and you match it against the order, and the signature means you did that.
I wrote "per IP address" into a config file without once checking which address the process could see. That's the whole bug. I skipped the counting.