System Design October 1, 2026 ~22 min read

// Designing a URL Shortener in AWS — part 1 of 3

Designing a URL Shortener in AWS (Part 1): The Architecture

TL;DR: mol.la is an open-source URL shortener I built in Go on AWS. The design target is 100 million new links a month with a 100:1 read-to-write ratio, a p99 redirect under 50 ms at the edge, and 99.9% availability for the redirect path. The whole design follows three rules.

First, the redirect touches as little as possible. A click is answered by CloudFront if the same code was seen in the last 5 seconds, then by Redis, and only on a double miss by DynamoDB. Click counting happens after the response, never inside it.

Second, short codes come from a leased block of integers, not a hash and not a Snowflake ID. Each Lambda execution environment takes 1,000 integers from one DynamoDB counter with a single atomic ADD, then hides the sequence behind a keyed permutation before Base62 encoding. Lambda has no stable worker ID, which is the one assumption Snowflake needs.

Third, the edge is two layers, and nothing reaches the origin without proof it came through both. Cloudflare terminates TLS and runs the firewall. CloudFront routes by path and caches redirects. A shared secret header, stamped by Cloudflare and checked by a CloudFront function, closes the direct path to the distribution.

The system is deployed and serving traffic, at hobby scale, on the smallest sizes of every managed service. The numbers in this series are the design targets that the code and the load scripts were built for, not measurements from a launch event. Where a number is a model rather than a measurement, the text says so.

This is Part 1 of a three-part series, written as a design reference: the requirements, the options I compared, and why I chose what I chose. The code is on GitHub.

  • Part 1 (this post): the problem, the short code scheme, the read path, the edge, and the whole architecture.
  • Part 2: never serving a dead link, covering the cache record, versioned tombstones, stale fallback, and takedown.
  • Part 3: writes, clicks, and the slow path, covering idempotent creation, the click pipeline, abuse controls, and testing.

The problem

What makes a URL shortener hard?

The product is small. Someone submits a long URL and gets back a seven-character code. Anyone who opens https://mol.la/{code} is sent to the long URL. That is two endpoints and one table.

The traffic shape is what makes it hard. A shortener is written once and read many times, and the reads are not spread evenly. A link posted in a popular place gets thousands of clicks a minute for an hour and then almost none. Every one of those clicks is a request that a human is waiting on, in the middle of opening a page. The redirect must be fast, it must be right, and it must stay right after the owner deletes the link or an operator takes it down.

Written out as requirements:

RequirementDetail
Write volume100 million new links a month, about 39 a second on average and 154 a second at peak.
Read volume100 reads for every write: about 3,900 redirects a second on average, 15,400 at peak.
Redirect latencyp99 under 50 ms measured at the edge.
Availability99.9% for the redirect path, measured at CloudFront. A region-wide outage is an accepted risk in the first release.
RetentionLinks live for 5 years by default, then expire. Callers can choose a shorter lifetime, down to one minute.
Custom aliasesA caller can pick a 3 to 32 character alias instead of a generated code.
Click statsA count per link, allowed to lag by minutes. Not billing-grade.
Deletion and takedownA deleted or taken-down link must stop redirecting quickly, and must never come back.
SafetyThe service never fetches, previews or probes a destination URL.

Two of these define correctness: a deleted link must stay dead, and a new link must be visible on its first click. Part 2 is about those two. The rest of this post is about the shape of the system that serves the reads cheaply.

The shape of the load

Where do 10 billion monthly redirects get answered?

The design assumes three cache tiers between a click and DynamoDB, and it assumes a hit rate for each. Figure 1 shows the model that every capacity and cost number in this series is built on.

Redirects per month reaching each tier under the design's hit-rate model
Figure 1: Where redirects are answered under the design model. CloudFront sees all 10 billion. About half reach the redirect Lambda. Redis answers 95% of those, so about 250 million reads a month reach DynamoDB. These are assumptions to be replaced by metrics, not measurements.

Three things in that picture shaped the design:

  1. Viral links and the long tail behave differently. CloudFront caches a redirect for only 5 seconds. For a link clicked a few times a day, that is always a miss. For a link clicked a thousand times a minute, it absorbs almost everything. The 50% edge hit rate is a blend of the two, and it is the number I trust least. It is also the number that matters least for correctness, because the edge cache is short by design.
  2. The database is the last resort, not the first. With Redis in front of it, DynamoDB sees about one read in forty. A viral link is one partition key, and it cannot be sharded without changing the URL that people have already shared. The mitigation is that most of its traffic never reaches the table.
  3. The per-request cost at the edge dominates the bill. Every one of the 10 billion requests is a CloudFront viewer request, hit or miss. That is the expected shape of a 100:1 read-heavy service, and it is why the edge layer gets so much attention below.

Challenge 1: where does the short code come from?

What is the single decision that everything else depends on?

Every link needs a code that is unique, short, hard to guess, and cheap to mint under concurrent writers. Whatever produces that code is on the write path of every request, so I chose it first.

I compared four options.

OptionHow it worksWhy not (or why)
A. Hash the long URLMD5 or SHA-256 of the URL, take the first 7 characters, check for a collision, retryCollisions become common at billions of rows, and every retry is a database round trip. Two people shortening the same URL also get the same code, which leaks who shortened what.
B. One incrementing counterA database row holds the next integer; every create increments itA single hot writer that every request waits on. Codes are also sequential, so anyone can enumerate every link ever created.
C. Snowflake-style IDsTimestamp, worker ID and a sequence number, packed into an integer and Base62-encodedNeeds a stable, long-lived worker ID per generator. AWS Lambda recycles execution environments on its own schedule and gives them no durable identity. Two environments that mint the same worker ID can mint the same code.
D. Leased blocks of integers, permutedOne DynamoDB counter item, advanced 1,000 at a time with an atomic ADD. Each execution environment hands out integers from its block in memory. A keyed permutation hides the sequence before Base62 encoding.One write to the counter per 1,000 codes, about 0.15 writes a second at peak. No coordination between environments and no worker identity needed. A wasted partial block after a cold start costs nothing but ID space, of which there are 3.5 trillion.

Snowflake-style IDs on Lambda compared with leased integer blocks
Figure 2: Snowflake assumes each generator has a durable worker ID, which Lambda cannot promise. Block leasing needs only one atomic counter update per block, and the permutation hides the order from clients.

I chose D. The deciding argument was Lambda itself. Snowflake is the usual production recommendation, and it is a good one on a fleet of long-lived servers. On a fleet of ephemeral functions it reintroduces the collision risk it was designed to remove. Block leasing needs only one atomic increment, which DynamoDB’s UpdateItem ADD provides without a read-modify-write race.

Three details make this safe and unguessable:

  • The permutation is a six-round Feistel network over 42 bits, using HMAC-SHA-256 from the Go standard library as the round function, with the key held in an encrypted parameter. Values that land outside the Base62 domain are cycle-walked (permuted again until they land inside). The mapping is a bijection, so two different integers can never produce the same code, and the inverse exists for anyone holding the key.
  • Seven Base62 characters give 62^7, about 3.5 trillion codes. At 1.2 billion links a year the space lasts about 2,900 years. Base62 rather than Base64 because + and / are not safe in a URL without escaping.
  • A generated code is written with the same conditional put as a custom alias. If a caller once chose the alias aB3xK9c by hand, the generated code that happens to equal it fails its condition and the allocator takes the next integer. That retry is bounded to the handful of aliases that ever collide, not a search.

The raw integer also reserves 3 high bits for a region ID, so a second region could mint codes without a cross-region coordinator. The first release uses region 0.

Challenge 2: keep the redirect short

What exactly happens inside a click?

The first design I sketched did the obvious thing: look the code up, increment a click counter on the same row, return the redirect. Under a burst, the counter increment makes the hottest row in the table also the most-written row in the table. That is the classic failure of this kind of system, and it is the one the reference literature warns about most loudly.

So the redirect is split into a hot path (what the click waits for) and a slow path (what happens after the response is sent).

Redirect hot path vs the click slow path
Figure 3: The hot path is one Redis read, one DynamoDB read only on a miss, and one click item written after the 302. Click counting happens on the slow path, fed by the table's stream.

On the hot path, GET /{code} does at most three things:

  1. Read Redis for url:{code}. On a fresh active record, that is the whole lookup.
  2. Read DynamoDB with a consistent GetItem, only on a miss or when the cached record is stale. The result populates Redis.
  3. Write one click item to a separate Clicks table, after the 302 has been written, with a 500 ms budget. If the write fails, the click is logged and lost. The redirect is never delayed or changed by it.

The link item is deliberately never written on a click. An aggregation Lambda reads the click items from the table’s stream in batches and increments a separate stats table a few seconds later. Part 3 covers that path, and the four ways of moving clicks that were compared before settling on it.

Why 302 and not 301. A 301 tells the browser the redirect is permanent, and browsers cache it indefinitely. That kills click counting, and it makes expiry, deletion and takedown impossible for anyone who has already clicked. A 302 brings the client back on almost every click. The cost of that round trip is paid by the cache tiers, not by weakening the redirect.

Caching a 302 anyway. The redirect carries Cache-Control: public, max-age=5. That looks like a contradiction, but it is not. A 302 tells the browser not to assume the destination is permanent. It does not forbid an intermediary from caching it for exactly as long as the header says. Five seconds at the edge absorbs a viral spike without every click reaching Lambda, and it bounds how long a deleted link can be served after takedown to five seconds. Not-found and error responses carry no-store, so a custom alias created a second after a miss is visible at once.

Challenge 3: three cache tiers, and why Redis over DAX

Why not the managed cache that sits in front of DynamoDB?

DynamoDB Accelerator (DAX) is the obvious cache for a DynamoDB-backed read path. It was rejected because the redirect needs more than a cached row. Part 2 explains the cache record in full, but the short version is:

NeedDAXRedis
An explicit “deleted” state that carries no destinationNo. DAX caches items; a deleted item is a miss, and a miss goes back to the table.Yes. A tombstone record with its own lifetime.
Version comparison on write, so an old populate cannot overwrite a newer tombstoneNoYes, in one atomic Lua script.
A bounded stale window for use only when DynamoDB is failingNoYes. fresh_until and stale_until are part of the record.
Provider neutralityAWS onlyAny Redis. The same adapter runs against a local Redis in development.

So the tiers are:

  • Tier 1, CloudFront. Under 5 ms. Caches a successful 302 for 5 seconds and nothing else. No stale-while-revalidate, because that would extend takedown propagation past the 5-second window.
  • Tier 2, ElastiCache for Redis. Under 10 ms. Holds active records and deletion tombstones. Sized for the hot 20% of the daily working set, about 1 GB at the design targets. The Redis client has 150 ms dial, read and write timeouts and retries disabled, so a Redis failover degrades to a DynamoDB read instead of a hung invocation.
  • Tier 3, DynamoDB. Only on a double miss or a revalidation. Reads are strongly consistent, so a newly created link is visible on its first click rather than after DynamoDB’s eventual-consistency window.

Unknown codes are cached nowhere. Negative caching was rejected because any valid alias can be created after an earlier miss, and a cached “missing” sentinel would hide it. A Bloom filter was rejected too: at 6 billion codes and a 1% false-positive rate it needs roughly a 7 GB bitset plus a correctness-sensitive update pipeline, to save reads that strict code validation and the edge rate limits already bound.

Challenge 4: the edge

Why two edge providers, and how does nothing get around them?

The redirect endpoint is public and unauthenticated, so it needs a source rate control that costs less than the abuse it stops. The original design put an AWS WAF web ACL on the CloudFront distribution. At the design volume that is 10 billion requests evaluated a month, about $6,000 in WAF request charges alone. For a public shortener that is the wrong trade.

The deployed design puts Cloudflare in front of CloudFront:

  • Cloudflare holds the DNS zone, terminates TLS in Full (strict) mode against the ACM certificate on CloudFront, runs the firewall, and rate-limits POST /api/v1/links per client IP. It also sets the real client IP in CF-Connecting-IP on every request, which the origin can trust because Cloudflare overwrites any value the client sends.
  • CloudFront routes by path. /api/* goes to API Gateway with caching disabled. /app/* and the root path go to a private S3 bucket through Origin Access Control and serve the static single-page app. Everything else, which is any bare single-segment path, goes to the redirect origin with the 5-second cache policy. The page and the API share one hostname, so the browser never sends a CORS preflight.
  • The redirect origin is a Lambda Function URL, not API Gateway. API Gateway would add about $1 per million requests on the read path, roughly $5,000 a month at 5 billion origin requests, for routing and usage plans that the redirect does not use. The Function URL costs nothing beyond the Lambda invocation. It uses IAM authorization, and CloudFront signs every origin request with Origin Access Control, so CloudFront is the only caller that can reach it.

A second edge layer creates a new hole: the *.cloudfront.net hostname is reachable by anyone, and a request sent straight to it never passes Cloudflare’s firewall. Figure 4 shows how that hole is closed.

Cloudflare stamps a secret header; a CloudFront function rejects requests without it
Figure 4: Cloudflare adds X-Origin-Verify to every request it forwards. A CloudFront viewer-request function rejects anything without the right value with a 403, before any origin is contacted, and strips the header so no origin ever sees it.

A Cloudflare Transform Rule stamps X-Origin-Verify with a secret on every forwarded request. A CloudFront function runs on every viewer request, compares the header, returns 403 with no-store on a mismatch, and deletes the header before forwarding. The secret is generated by Terraform on the first apply and stored in the state. Both the rule and the function update in one plan, so rotating it is one apply and a few seconds of 403s while the function propagates.

Two more edge details are worth knowing. CloudFront’s response headers policy sends a strict Content Security Policy as a header, because browsers ignore frame-ancestors in a meta tag, and it strips the S3 encryption headers that would otherwise reveal the AWS account ID and KMS key ARN to every viewer. And Cloudflare’s bot protections are pinned off in Terraform, because the JavaScript detections inject an inline script that the site’s CSP blocks.

Challenge 5: compute for a serverless read path

Why Lambda, and what does a VPC-attached function cost?

I compared three compute options for the redirect.

OptionWhy not (or why)
A fixed fleet behind a load balancerPredictable and fast, but it means paying for idle capacity all month for a service whose traffic arrives in spikes, plus the operations of a fleet: patching, scaling policies, and a load balancer with its own hourly charge.
Containers on FargateScales, but slower than Lambda to react to a viral spike, and the load balancer cost is the same.
Lambda (chosen)Scales to a spike in seconds, costs nothing at idle, and the Go binary on arm64 starts in about 100 to 300 ms cold. The p99 target rules out cold starts during ramps, so the redirect function is built to run with provisioned concurrency, sized from observed concurrency (peak QPS times p50 duration, about 160 at the design targets).

Reaching Redis puts the redirect function inside a VPC, in two private subnets across two Availability Zones. The write function stays outside the VPC, because its normal request path never touches Redis. Consequences:

  • No NAT gateway and no interface endpoints. No function calls the public internet. Every dependency the redirect function has outside the VPC is DynamoDB, reached through a gateway VPC endpoint, which has no hourly charge. That was not true of the first deployed version, which published clicks to Kinesis through an interface endpoint at about $7.30 per AZ-month. Part 3 explains the change.
  • One connection per execution environment. The Redis and SDK clients are created once at cold start and reused across invocations. Provisioned concurrency keeps a warm pool of them.
  • Two roles, two functions. The redirect function can read links, write click items and reach Redis. It cannot create or delete anything. The API function can create links and invoke cache invalidation. It cannot reach Redis directly or write clicks.

Every function is Go, compiled for provided.al2023 on arm64, at 128 MB. The redirect function has a 3-second timeout, and every dependency call inside it has a shorter budget: 150 ms for Redis, 1 second for the DynamoDB link read including the SDK’s three attempts, 500 ms for the click write. A hung dependency therefore leaves room to fall back or fail fast, never to time out the whole invocation.

The whole architecture

What does it look like end to end?

URL shortener architecture on AWS: Cloudflare in front of CloudFront, API Gateway and an API Lambda for writes, a redirect Lambda in a VPC with ElastiCache for Redis, DynamoDB with a Clicks table and stream, an aggregation Lambda, an invalidation Lambda for takedowns, and GitHub Actions deploying through OIDC

Figure 5: The full architecture as deployed. Cloudflare and CloudFront at the edge; a redirect Lambda in private subnets with Redis and VPC endpoints; API Gateway and an API Lambda for writes; a DynamoDB Clicks table, its stream and an aggregation Lambda on the slow path. Open full-size diagram

A click moves through it like this:

  1. The browser resolves mol.la to Cloudflare, which terminates TLS, runs the firewall, stamps the origin secret, and forwards to CloudFront.
  2. A CloudFront function checks the secret and strips it. On a cache hit, CloudFront answers the 302 itself.
  3. On a miss, CloudFront calls the redirect Lambda through its Function URL, signed with Origin Access Control.
  4. The function reads Redis. A fresh active record answers the click.
  5. On a Redis miss or a stale record, it reads DynamoDB through the gateway endpoint and populates Redis.
  6. After writing the 302, it puts one item into the Clicks table through the same gateway endpoint.

A link is created like this:

  1. POST /api/v1/links goes through CloudFront to API Gateway and the API Lambda, outside the VPC.
  2. The function validates the request, leases an integer from the Counters item if no alias was given, and writes the Links item and the optional Idempotency item in one DynamoDB transaction.

And the slow path:

  1. An aggregation Lambda consumes the Clicks table’s DynamoDB Stream in batches and increments LinkStats once per code per batch. Raw click items expire from the table after 90 days.
  2. An operator takedown soft-deletes the item and invokes the invalidation Lambda, which writes a versioned tombstone to Redis. The cached 302 at the edge ages out on its own within 5 seconds.
  3. GitHub Actions deploys by assuming an IAM role through OIDC, on a version tag only. There are no long-lived AWS keys anywhere in the pipeline.

Alarms cover the things that should wake someone up: CloudFront 5xx rate and origin p99 latency against the 50 ms target, API Gateway 5xx and p99 latency, Lambda errors and throttles on every function, Redis engine CPU and evictions, DynamoDB throttles on every table, the aggregation function’s iterator age, its dead-letter queue, and any failed cache invalidation, which is owned by security rather than platform because a missing tombstone means a taken-down link is still live. Alarms go to an SNS topic and from there to email. A budget with a monthly limit sits next to them.

What it would cost at the target

What does a 100:1 read-heavy service cost at 10 billion reads a month?

These are list-price arithmetic at the design targets, not a bill. The deployed system runs at a tiny fraction of this.

Line itemBasisApprox. per month
CloudFront10B requests, about 5 TB egress~$7,900
Lambda~5B redirect invocations plus 100M writes, 128 MB, arm64~$1,100
Lambda provisioned concurrency~160 concurrent, redirect function only~$450
ElastiCache for Rediscache.r7g.large, primary plus replica~$400
API Gateway (write path only)100M requests~$350
DynamoDB, links100M writes plus ~250M on-demand reads~$200
DynamoDB, clicks5B on-demand writes, one per redirect that reaches Lambda~$6,250
CloudWatch, Function URL, gateway endpoint~$100
CloudflareFree plan$0
Total~$16,750

CloudFront request pricing dominates because every redirect is a viewer request whether it hits or misses. Two lines moved since the original design. AWS WAF at $0.60 per million requests, about $6,000, disappeared when the firewall moved to Cloudflare. And the click line grew from about $250 for Kinesis to about $6,250 for one DynamoDB write per click, because the deployed system chose the pipeline with no fixed monthly cost over the one with the lowest cost per event. At hobby scale that is the right trade by a wide margin, and Part 3 shows the arithmetic in both directions.

What’s next

Part 2 looks inside the cache. It covers what a cache record holds and why it has three timestamps, how a versioned tombstone survives a late writer, when a stale record may be served and when it must not, and how a takedown propagates through DynamoDB, Redis and the edge in that order.

// end of article — process exited with code 0

// share:Twitter / XLinkedIn
// llm:View as Markdown

// comments

loading...