TL;DR: The redirect path serves almost every click from cache, so the cache has to be right about two things: a new link must be visible on its first click, and a deleted link must never come back. Both come down to what a cache record holds and how it is written.
- Three timestamps, not one TTL. A record is fresh for up to 24 hours, then stale for 15 minutes more, and the Redis key itself never outlives the link’s own expiry. Stale means “check DynamoDB first”, not “serve it”.
- Every write compares versions. DynamoDB increments the link’s version on every state change. A Redis write is one Lua script that refuses to overwrite a record with a higher version. A slow populate can never undo a tombstone.
- Tombstones carry no destination. A deleted link is cached as
deletedwith its version and a 24-hour lifetime. It answers 404 without touching the table. - Stale is a fallback, not a feature. A stale active record is served only when DynamoDB fails, only within the 15-minute grace, and never over a tombstone.
- Takedown finishes in Redis, not in DynamoDB. The durable soft delete happens first. The operation is not reported as done until a tombstone with that version is in the cache. The edge entry ages out on its own within 5 seconds.
This is Part 2 of a three-part series. Part 1 explains the architecture and why the redirect sits behind three cache tiers. Part 3 covers link creation, click analytics, and testing.
The invariants
What exactly does “never serve a dead link” mean?
Before writing the cache, I wrote down three invariants (rules that must always be true). Everything in this post exists to keep them true.
- A deleted link stays deleted. Once a delete or takedown is reported as successful, no later click is redirected to that destination, apart from a 302 already cached at the edge for at most 5 seconds.
- A new link is visible at once. The first click after creation, including a custom alias created a moment after someone tried it, is served, never hidden behind a cached miss.
- An expired link is not served. The link’s own
expires_atis checked on every path, in the cache and in the table, and never delegated to a background cleanup.
Invariant 1 is the one that a naive cache breaks most easily. Invariant 2 is the reason unknown codes are never cached. Invariant 3 is the reason DynamoDB TTL is treated as cleanup, not as a correctness gate.
The state
What lives in DynamoDB, and what lives in Redis?
DynamoDB is authoritative. Redis is disposable: it can be flushed at any time and rebuilt from the table by ordinary traffic. Neither of those statements is true unless the record formats make them true.
The Links item. One item per code, no sort key, so every redirect is a single-digit-millisecond GetItem on the exact key.
| Attribute | Purpose |
|---|---|
short_code | Partition key. A generated code or a custom alias. |
long_url | The validated destination. Immutable after creation. |
is_active | False once deleted. |
version | Starts at 1. Incremented atomically on every state change. This is what makes cache writes safe. |
expires_at | The link’s logical end. Checked on every read. |
purge_at | The DynamoDB TTL attribute. Equals expires_at for active links, and deleted_at plus 30 days for deleted ones. |
DynamoDB deletes expired items on a best-effort schedule, sometimes days late. So purge_at is cleanup only. The read path also requires now < expires_at and is_active, every time.
The Redis record. Key url:{code}, a small JSON document:
| Field | Purpose |
|---|---|
state | active or deleted. |
long_url | Present only on active records. A tombstone never carries a destination. |
version | Copied from the DynamoDB item at the time of the write. |
expires_at | The link’s own end, copied from the item. |
fresh_until | Serve from cache alone until this moment. At most 24 hours after populate, and never past expires_at. |
stale_until | The Redis key’s own TTL. 15 minutes after fresh_until for active records, 24 hours after deleted_at for tombstones. |
Why three timestamps instead of one TTL? Because they answer three different questions. expires_at asks whether the link is alive. fresh_until asks whether the cache may answer alone. stale_until asks whether the cache is allowed to hold anything at all. Folding them into one number forces one of two bad trade-offs: either a short TTL that sends most clicks to the table, or a long TTL that serves a deleted link for its whole length.
The lookup
What does the redirect do with each cache outcome?
The handler first checks the code against the alias rule (3 to 32 characters, letters, digits, hyphen, underscore). Anything else is a 404 before any lookup, which bounds what probing can cost. Then it reads Redis and classifies the record against the current time:
- Fresh active: return 302 with
Cache-Control: public, max-age=5. This is the common case and the only path that costs one round trip. - Deleted: return 404 with
no-store. No table read. A tombstone is a positive answer, not a miss. - Stale active: revalidate. Read DynamoDB and apply whatever it says.
- Miss, or expired by its own timestamps: read DynamoDB.
- Redis error or timeout: treat as a miss. The Redis client has a 150 ms budget and retries disabled, so a failover costs at most 150 ms before the table read starts.
After a DynamoDB read, the outcome is decided once, in one pure function, from the item’s is_active, expires_at and the current time:
- Active and not expired: populate Redis with an active record and return 302.
- Unknown or expired: return 404 and cache nothing. This is invariant 2. A cached miss would hide an alias created a second later. Expired links are not cached either, because a tombstone for a link that simply ran out would be a 24-hour entry for nothing.
- Inactive: write a tombstone and return 404. From now on the table is not consulted for this code for 24 hours.
Everything above is deterministic. The one judgement call is what to do when DynamoDB itself fails.
Failure mode 1: the database is down and the cache is old
When is it safe to serve a stale destination?
DynamoDB is very reliable, but the redirect function talks to it through a VPC endpoint with a 1-second budget, and throttles or a slow partition on a viral key are real. When the revalidation read fails, there are three options.
| Option | Problem |
|---|---|
| Return 503 | Every click on a link whose cache entry passed fresh_until in the last few minutes fails, while a perfectly good destination sits in Redis. |
| Serve the stale record, whatever it is | A deleted link whose tombstone was never written, or whose cache entry is from before the delete, redirects again. That breaks invariant 1. |
| Serve the stale record only inside a bounded window (chosen) | The stale destination is served only if the record is active, now < stale_until, and now < expires_at. A tombstone is always a 404. Anything else is 503 with no-store. |
The window is the 15 minutes between fresh_until and stale_until. Within it, the worst case is that a link deleted in the last 15 minutes, whose tombstone write also failed, is served during a DynamoDB outage. That is a double failure of two independent systems inside a short window, and the alternative is failing every revalidating click for the length of the outage. I accepted it, and the handler logs a warning on every stale serve so the count is visible.
What the window never does is override a tombstone. If Redis says deleted, the answer is 404 no matter what DynamoDB is doing. The tombstone is a positive fact, and the stale path only ever fills in for a missing one.
Failure mode 2: the slow writer
Can a cache populate undo a delete?
Yes, if writes are plain SETs. The redirect function and the invalidation function both write the same key, and they are not coordinated. Consider a click and a takedown that overlap:
The redirect read the item while it was still active. The takedown ran, wrote the DynamoDB update and the Redis tombstone. Then the redirect’s populate, carrying the old active record, arrived and overwrote the tombstone. The link is dead in the table and alive in the cache for the next 24 hours. Every click is a cache hit, so nothing ever asks the table again.
The fix is a version on every record and a compare on every write. DynamoDB increments version in the same conditional update that flips is_active. Both the redirect and the invalidation function copy that version into the record they write. The write itself is one Lua script:
local existing = redis.call('GET', KEYS[1])
if existing then
local ok, decoded = pcall(cjson.decode, existing)
if ok and type(decoded) == 'table' and decoded['version'] then
if tonumber(decoded['version']) > tonumber(ARGV[2]) then
return 0 -- stored record is newer: keep it
end
end
end
redis.call('SET', KEYS[1], ARGV[1], 'PX', tonumber(ARGV[3]))
return 1
Redis runs a script from start to finish before anything else, so the read and the write cannot be separated by another writer. A separate GET followed by a SET in Go would reopen exactly the race in the figure. Equal versions are allowed to overwrite, so a retried write is idempotent (safe to run more than once, with the same result), and an active populate carrying the same version as the stored active record simply refreshes its timestamps.
The rule I took from this is the same one that decided the seat reservation system’s admission script: a guard and the action it guards must live in the same atomic step. A check followed by a separate write is a race, and under load every race is lost sooner or later.
Failure mode 3: the delete that only half happened
What if the table is updated but the cache is not?
A delete has two writes in two systems, and the second one can fail: the invalidation function can time out, Redis can be mid-failover, the invoke can be throttled. If the operation reports success anyway, the link is dead in DynamoDB and alive in Redis for up to 24 hours. That is invariant 1 broken by a partial failure rather than by a race.
The order is deliberate, and each step is idempotent:
- Identity comes from AWS, not from a flag. The admin CLI calls
GetCallerIdentityand uses the returned ARN as the actor. There is no--actoroption to forge. A reason is mandatory. - The durable delete is a conditional update.
is_activebecomes false andversionbecomesold + 1, on the condition that the item is still active with the old version. A retry after a partial failure finds the item already inactive and returns the stored deletion, with its version and timestamps, instead of failing. A retry by a different actor is refused. - Invalidation is synchronous. The API role and the operator role are the only principals allowed to invoke the invalidation function, and they invoke it with
RequestResponseand wait. The payload is{short_code, version, deleted_at}. Success means the tombstone with that version is in Redis. An older invalidation can never overwrite a newer record, because of the script above. - A failed invalidation is a failed delete. The operation returns 503, records an audit event with outcome
invalidate_failed, and raises the alarm that security owns. The caller retries. The retry skips the state change, reissues the same tombstone, and returns success once it lands. The item is never reported deleted while a known positive cache entry may remain. - The edge takes care of itself. CloudFront caches a 302 for 5 seconds and nothing else. There is no invalidation API call, no distribution-wide purge, and no
stale-while-revalidatethat could extend the window.
The same service runs an owner’s own delete when authenticated ownership returns. In the current release, create and stats are public, so there is no owner to authorise, and the only delete is the operator’s.
Why not DAX, and why not a simpler Redis
Was all of this necessary?
Part 1 listed the reasons DAX could not replace Redis here: no explicit deleted state, no version comparison, no bounded stale window. It is worth restating them from this side, because each one maps to an invariant.
| Invariant | What the cache must support | DAX | Redis with this record |
|---|---|---|---|
| A deleted link stays deleted | A positive “deleted” answer that survives late writers | No: a deleted item is just a miss, and a miss goes to the table | Tombstone plus version compare |
| A new link is visible at once | No negative caching of misses | Item cache only, so misses are not cached, which is fine | Unknown codes are never written |
| An expired link is not served | The record carries the link’s own expiry and the read checks it | The item carries it, but the cache TTL is separate | expires_at in the record, checked on every read, and the key TTL is capped at it |
A simpler Redis layout would also have worked for the first two rows, using two keys per code or a plain SET with a short TTL. The third row and Failure mode 2 are what forced the versioned single record. Once every write is one script that compares versions, the tombstone, the stale window and the takedown retry all fall out of the same mechanism.
Two configuration details finish the picture. Redis runs with in-transit and at-rest encryption and an AUTH token held in an encrypted parameter. And eviction is allowed: any record can be evicted under memory pressure, because eviction is just a miss, and a miss goes to DynamoDB. That is the opposite of the seat reservation cache, where eviction would have lost the source of truth. Here the source of truth is the table, and the cache is only ever a copy.
How it is tested
How do you convince yourself a cache is right?
- The redirect decision is a pure function with a table test for every branch: active and not expired, active and expired, inactive, unknown.
- The handler is tested against in-memory fakes of the store, the cache and the publisher, for every outcome in Figure 2: fresh, stale with a healthy table, stale with a failing table inside and outside the grace window, tombstone with a failing table, miss, and a Redis error.
- The Redis adapter is tested against a fake Redis server in-process. The tests cover the version script directly: a lower version is rejected, an equal version is accepted, a tombstone survives a late active write, and the key TTL never exceeds the link’s own expiry.
- The takedown path is tested for partial failure: a delete whose invalidation fails returns a dependency error, records the failed-invalidation audit event, and succeeds on retry without a second state change.
None of these tests needs a cloud account, a network, or a running emulator. Part 3 explains why the code is structured to make that possible.
What’s next
Part 3 follows a link from creation to its click count. It covers how a create request is validated without ever fetching the destination, how an idempotency key and a generated code are written in one transaction, how the click stream is kept off the redirect and aggregated in batches, what the abuse controls are for a public API with no accounts, and how the whole thing is tested in layers without a cloud account.