Skip to content
Brian Sithu
Go back

Rate limits weren't built for an agent army

Cover image for Rate limits weren't built for an agent army

An unattended agent fleet hammering a map API shows what request counts miss. I would check identity and budgets instead.

Brian Sithu

Written by

Brian Sithu

View profile

On this page

05 sections

On October 5, independent researchers posted preliminary findings about a fleet of AI agents running on Tencent infrastructure and hitting Alibaba’s Amap map service with repetitive queries. I am going on the TechCrunch writeup here, since the original research is not linked in the coverage. The reported shape: unattended agents asking for directions to different entrances of public places. A park, a zoo, a hospital. Over and over.

TechCrunch frames it as rule-sidestepping rather than anything worse. No breach, no exfiltration, just a fleet quietly ignoring how the API wants to be used. That framing undersells it. The interesting part is not what this fleet took. It is that nothing in a normal API defense noticed a hundred callers doing the same chore forever.

The researchers reached for a careful phrase. Agent fleet, not swarm. Many parallel agents on the same kind of task, with no sign they coordinate with each other. That distinction is the whole post.

Every API I have run assumes the caller is one of two things. A human at a keyboard or a batch job on a schedule. Humans are bursty and slow. They get bored, they sleep, they stop. Batch jobs are the opposite and just as legible. One credential, a predictable hour, a volume someone signed off on, and a person to email when it misbehaves.

A fleet is neither. It does not sleep or get bored. It retries forever. It splits across keys and addresses so no single one looks loud. It does what the report describes: a hundred quiet actors doing the work of one loud one, with nobody to email.

Map data is the perfect first target, and my guess is that is no accident. Fresh local facts like which entrance to use are exactly the ground truth that training data lacks and agents get asked about constantly. Each query is cheap. The ten thousandth is not.

A rate limit answers the wrong question

A rate limit asks one thing. Is this key too loud right now. The fleet never asks that question. Each key stays polite. The abuse lives in the aggregate, and the aggregate has no key. A single polite agent that stays under the limit but never stops is worse than a burst in one specific way. A burst ends.

The usual answers fail in order. Per-IP blocks punish whoever shares the infrastructure. There is no browser, so there is no CAPTCHA. The User-Agent string is self-declared. Revoke one key and the fleet mints ten more, because keys are cheap and the task queue is long.

Check identity, not loudness

What I would build instead starts with identity at the product boundary. Every inbound call carries a signed token that says three things. Which agent this is, who runs it, and on whose behalf it acts. The edge verifies the signature without calling home. No token, no full service. My guess is most casual abuse dies right here, because signing means the operator can be found.

Spend budgets, not request counts

Then stop counting requests and start counting spend. Each verified principal gets a budget per day. Compute, egress, downstream calls, all drawn from one allowance. When the budget runs thin the agent degrades instead of dying. Cached answers, coarser results, a slower lane. Abuse is only cheap when someone else pays. A budget makes the operator feel every query, including the ten thousandth directions lookup.

My proposal: three layers between any agent and the API. Identity verifies the caller, the budget ledger deducts the cost, and the upstream serves or degrades

Trace every request to something billable

The last layer is the one that survives key rotation. Every request maps to an account that can be charged, throttled, or cut off. The point is not punishment. It is giving the operator a handle that holds when keys do not. Kill one key and the fleet mints ten. Cut the budget and all ten starve together.

What I would build instead

Rate limits assume someone is home on the other end of the key. Someone who notices the 429, feels embarrassed, backs off. Agents removed the someone. There is only a task queue and infinite patience.

So I would stop asking how loud a key is and start asking who is behind it, what they have spent, and which account answers for it. Identity at the edge, budgets behind that, billing behind that. Three layers, each one cheap to check on the hot path. The fleet in the report did nothing exotic. It just showed up in numbers, politely, forever. Defenses built for humans at keyboards will keep missing that shape until the unit of enforcement becomes the operator instead of the request.



Previous Post
Three hundred tokens a second
Next Post
The postmortem is already written