System DesignOctober 7, 2026~9 min read

Running a URL Shortener on Cloudflare for $0 (Part 3): Writes and Abuse on a Free Plan

TL;DR: A create request is validated without ever fetching the destination. The Worker takes the next ID from a block of 100 leased from a D1 counter, and turns it into a 7-character code with a keyed Feistel permutation. It writes the link and its optional idempotency record in one D1 batch. Abuse control has two layers: a zone WAF rule that blocks floods before the Worker runs, and a rate-limit binding inside the Worker. Neither one fully protects the daily write quota, and this part explains why.

The create path

What happens inside POST /api/v1/links?

The body has three fields: long_url, an optional alias, and an optional expires_in in seconds. Validation is pure code with no I/O:

Input Rule
Body At most 16 KiB, checked against Content-Length and again against the bytes read
long_url At most 2,048 characters, http or https, a non-empty host, no user or password, no control characters
alias 3 to 32 characters from A-Z a-z 0-9 _ -. api and app are reserved
expires_in 60 seconds to 5 years. The default is 5 years
Idempotency-Key Optional, 1 to 128 printable ASCII characters

The URL is parsed but never fetched, resolved or previewed. A write path with no outbound requests cannot be used to reach an internal address. It also keeps the Worker far from the free plan’s limit of 50 outbound requests per invocation.

A valid request then gets a code, unless it brought its own alias, and goes to D1.

Challenge 1: codes without a coordinator

How do many isolates hand out unique codes with one counter?

There are four common ways to make short codes:

Option Problem here
Random 7 characters Every create risks a collision, and the risk grows as the table fills
A counter update per create One more D1 write for every create, all on one hot row
Snowflake-style IDs They need a stable worker ID, and a Worker isolate has none
A leased block of IDs, then a keyed permutation (chosen) One counter write per 100 creates. The permutation hides the order

The lease is one statement:

UPDATE counters SET counter = counter + 100 WHERE region = 'global' RETURNING counter

The isolate keeps the block [counter - 100, counter) in module memory and hands out IDs from it. When the block runs out, it leases another. If an isolate is evicted, the rest of its block is skipped. That leaves gaps but never duplicates. A block of 100 is small enough that an evicted isolate wastes little, and large enough that the counter costs one write per 100 creates.

From a leased block to a 7-character code
Figure 1: An isolate leases 100 IDs with one D1 update. Each ID goes through a 6-round Feistel network over 42 bits, keyed with HMAC-SHA256. Results of 62^7 or more are permuted again until they fit. The result is Base62-encoded into exactly 7 characters.

Sequential IDs would make codes guessable. A permutation fixes that: it maps every number in a range to a different number in the same range, so it never collides. mol.la uses a Feistel network, a construction that turns any round function into a permutation:

private async permute42(value: bigint): Promise<bigint> {
  let left = value >> 21n
  let right = value & HALF_MASK                  // two 21-bit halves
  for (let round = 0; round < 6; round++) {
    const next = left ^ (await this.round(round, right))
    left = right
    right = next
  }
  return (left << 21n) | right
}

The round function is HMAC-SHA256 over the round number and the right half, cut to 21 bits. The key is a Workers secret. Seven Base62 characters give 62^7, about 3.52 trillion codes. The Feistel network works on 42 bits, about 4.40 trillion values. A result outside the code space is permuted again (“cycle walking”). About 80% of IDs fit on the first pass.

Two details are specific to Workers:

  • WebCrypto is async. Each round awaits crypto.subtle.sign, so one code is six or more awaited HMACs. That takes microseconds, far below the 10 ms CPU limit. The imported key is cached in module memory, so importKey runs once per isolate.
  • Fixed vectors pin the output. A test checks the codes for fixed IDs and a fixed key, so a change to the permutation cannot quietly map old IDs to new codes.

A generated code can still hit an existing custom alias, because aliases come from the same alphabet. The primary key catches it, and the create tries the next ID, up to 32 times.

Challenge 2: one batch, three outcomes

How does a retried request get the same answer as the first one?

The web UI sends a fresh crypto.randomUUID() as the Idempotency-Key with every create. If the network fails after the server has written the link, the retry must return that link, not make a second one.

The link and the idempotency record are written in one db.batch, which D1 runs as a transaction:

try {
  await this.db.batch(statements)        // INSERT link; INSERT idempotency
  return link
} catch (err) {
  if (!isUniqueViolation(err)) throw new PlatformError('dependency', err)
  if (idem !== null) return this.replayIdempotency(link.ownerID, idem)
  throw new PlatformError('collision')
}

A unique-constraint failure has three meanings:

What collided Same request hash? Result
The idempotency key Yes Return the original link. This is the retry
The idempotency key No 409 IDEMPOTENCY_CONFLICT. The key was reused for a different request
The short code only n/a Try the next ID, or 409 ALIAS_TAKEN for a custom alias

The request hash is SHA-256 over the three body fields, in a fixed field order. Replay records last 24 hours.

Challenge 3: what a create really costs

How many D1 rows does one create write?

I measured each statement against local D1:

Statement Rows written
Insert the link 3: the row, the text primary key’s index, the purge_at index
Insert the idempotency record 3: the row, its primary key index, the ttl index
Lease a new block 1, once per 100 creates
First click on the new code 2: the stats row and its primary key index

So a create from the web UI writes 6 rows. A create through the API without a key writes 3. A takedown writes 2. A link in this schema costs more to create than to click, and the cost comes from the index entries.

Challenge 4: abuse without accounts

How do you protect a public API when the limit is a daily quota?

There is no sign-up and no API key, so limits are per client IP. There are two layers:

Layer Limit Where it runs Quota it saves
Zone WAF rate-limit rule 3 creates per 10 s per IP and data center, then a 10 s block Before the Worker Requests and D1 writes
Worker rate-limit binding 10 creates per 60 s per IP, per data center Inside the Worker, before D1 D1 writes

The WAF rule is the important one. A blocked request never runs the Worker, so a flood costs nothing. The free plan allows one such rule, with a 10-second period, a 10-second block, and cf.colo.id (the data center) as a required part of the key. The binding is the finer layer. It returns 429 with Retry-After: 60 and costs one Worker request. Cloudflare describes it as approximate and per data center by design.

Then the arithmetic. The binding lets one IP make 10 creates a minute, or 14,400 a day. At 6 rows each, that is 86,400 rows: most of the 100,000-row daily quota, from one client staying under every limit. Two such clients would stop all writes, including click counts, until 00:00 UTC.

Rate limits stop floods. They do not stop someone patient. That is the main weakness of a free-plan design with no accounts. When it happens, the result is an outage until midnight, not a bill, and the runbook’s upgrade trigger fires long before. The fixes are small: tighten the binding to a few creates a minute, or add a daily cap per IP. I have not done either yet, because real traffic is far from this point.

There is also no bot challenge. Cloudflare’s bot protection injects a JavaScript check that the site’s strict CSP (script-src 'self') would block, so Terraform turns it off.

Challenge 5: cleanup

What removes old rows?

A cron trigger runs once a day at 03:17 UTC and runs one batch:

DELETE FROM idempotency WHERE ttl < ?;
DELETE FROM links WHERE purge_at < ?;

purge_at is the link’s expires_at, or 30 days after a takedown. Deletes count as rows written too, so a day with many expiring links shows up in the write quota. Click counters for purged links are not deleted yet. Today that costs only storage.

How it is tested

The Worker tests run inside workerd, the real Workers runtime, through @cloudflare/vitest-pool-workers. D1 migrations are applied in the test setup, so every test runs against real SQLite with the real schema. The suite covers create, the idempotency replay and conflict cases, the rate limiter, redirects through the edge cache with click counting, and the cron purge. A separate test fails if the OpenAPI document drifts from the handlers. There is no load test, because any sustained load would use up the quota it was meant to measure.

What’s next

Part 4 covers deploying and running it: Terraform with its state in R2, deploys from tags, the first deploy by hand, one CSP kept in three places, backups and failure, and the bugs that only production showed.

// end of article — process exited with code 0

// share:Twitter / XLinkedIn
// llm:View as Markdown

// comments

loading...