# Designing a URL Shortener in AWS (Part 2): Never Serve a Dead Link

> Source: https://www.sumonselim.com/url-shortener-part2-never-serve-a-dead-link
> Author: Muhammad Sumon Molla Selim
> Published: 2026-10-02
> Tag: System Design
> Summary: A cache that is fast but wrong is worse than no cache. Part 2 explains the Redis record behind every redirect: three timestamps instead of one TTL, a version number on every write, tombstones that carry no destination, the narrow case in which a stale URL may still be served, and how an operator takedown propagates through DynamoDB, Redis and the edge in that order.

**TL;DR:** The redirect path serves almost every click from cache, so the cache has to be right about two things: a new link must be visible on its first click, and a deleted link must never come back. Both come down to what a cache record holds and how it is written.

- **Three timestamps, not one TTL.** A record is fresh for up to 24 hours, then stale for 15 minutes more, and the Redis key itself never outlives the link's own expiry. Stale means "check DynamoDB first", not "serve it".
- **Every write compares versions.** DynamoDB increments the link's version on every state change. A Redis write is one Lua script that refuses to overwrite a record with a higher version. A slow populate can never undo a tombstone.
- **Tombstones carry no destination.** A deleted link is cached as `deleted` with its version and a 24-hour lifetime. It answers 404 without touching the table.
- **Stale is a fallback, not a feature.** A stale active record is served only when DynamoDB fails, only within the 15-minute grace, and never over a tombstone.
- **Takedown finishes in Redis, not in DynamoDB.** The durable soft delete happens first. The operation is not reported as done until a tombstone with that version is in the cache. The edge entry ages out on its own within 5 seconds.

This is Part 2 of a three-part series. [Part 1](https://www.sumonselim.com/url-shortener-part1-architecture-for-redirects) explains the architecture and why the redirect sits behind three cache tiers. [Part 3](https://www.sumonselim.com/url-shortener-part3-writes-clicks-and-the-slow-path) covers link creation, click analytics, and testing.

## The invariants

**What exactly does "never serve a dead link" mean?**

Before writing the cache, I wrote down three invariants (rules that must always be true). Everything in this post exists to keep them true.

1. **A deleted link stays deleted.** Once a delete or takedown is reported as successful, no later click is redirected to that destination, apart from a 302 already cached at the edge for at most 5 seconds.
2. **A new link is visible at once.** The first click after creation, including a custom alias created a moment after someone tried it, is served, never hidden behind a cached miss.
3. **An expired link is not served.** The link's own `expires_at` is checked on every path, in the cache and in the table, and never delegated to a background cleanup.

Invariant 1 is the one that a naive cache breaks most easily. Invariant 2 is the reason unknown codes are never cached. Invariant 3 is the reason DynamoDB TTL is treated as cleanup, not as a correctness gate.

## The state

**What lives in DynamoDB, and what lives in Redis?**

DynamoDB is authoritative. Redis is disposable: it can be flushed at any time and rebuilt from the table by ordinary traffic. Neither of those statements is true unless the record formats make them true.

**The Links item.** One item per code, no sort key, so every redirect is a single-digit-millisecond `GetItem` on the exact key.

| Attribute | Purpose |
| :-- | :-- |
| `short_code` | Partition key. A generated code or a custom alias. |
| `long_url` | The validated destination. Immutable after creation. |
| `is_active` | False once deleted. |
| `version` | Starts at 1. Incremented atomically on every state change. This is what makes cache writes safe. |
| `expires_at` | The link's logical end. Checked on every read. |
| `purge_at` | The DynamoDB TTL attribute. Equals `expires_at` for active links, and `deleted_at` plus 30 days for deleted ones. |

DynamoDB deletes expired items on a best-effort schedule, sometimes days late. So `purge_at` is cleanup only. The read path also requires `now < expires_at` and `is_active`, every time.

**The Redis record.** Key `url:{code}`, a small JSON document:

| Field | Purpose |
| :-- | :-- |
| `state` | `active` or `deleted`. |
| `long_url` | Present only on active records. A tombstone never carries a destination. |
| `version` | Copied from the DynamoDB item at the time of the write. |
| `expires_at` | The link's own end, copied from the item. |
| `fresh_until` | Serve from cache alone until this moment. At most 24 hours after populate, and never past `expires_at`. |
| `stale_until` | The Redis key's own TTL. 15 minutes after `fresh_until` for active records, 24 hours after `deleted_at` for tombstones. |

![The lifetime of one cache record](https://www.sumonselim.com/images/articles/url-shortener/cache-record-timeline.svg "Figure 1: An active record is fresh for up to 24 hours, then stale for 15 minutes, then gone. The key's TTL is capped at the link's own expiry, so the record cannot outlive the link. A tombstone lives 24 hours and carries no URL.")

Why three timestamps instead of one TTL? Because they answer three different questions. `expires_at` asks whether the link is alive. `fresh_until` asks whether the cache may answer alone. `stale_until` asks whether the cache is allowed to hold anything at all. Folding them into one number forces one of two bad trade-offs: either a short TTL that sends most clicks to the table, or a long TTL that serves a deleted link for its whole length.

## The lookup

**What does the redirect do with each cache outcome?**

![What the redirect handler does with each cache outcome](https://www.sumonselim.com/images/articles/url-shortener/cache-lookup-flow.svg "Figure 2: The cache read has five outcomes. Only fresh and deleted are answered from Redis alone. Stale, miss and a Redis error all go to DynamoDB, whose answer decides what to cache and what to return. A read failure has one narrow fallback.")

The handler first checks the code against the alias rule (3 to 32 characters, letters, digits, hyphen, underscore). Anything else is a 404 before any lookup, which bounds what probing can cost. Then it reads Redis and classifies the record against the current time:

- **Fresh active:** return 302 with `Cache-Control: public, max-age=5`. This is the common case and the only path that costs one round trip.
- **Deleted:** return 404 with `no-store`. No table read. A tombstone is a positive answer, not a miss.
- **Stale active:** revalidate. Read DynamoDB and apply whatever it says.
- **Miss, or expired by its own timestamps:** read DynamoDB.
- **Redis error or timeout:** treat as a miss. The Redis client has a 150 ms budget and retries disabled, so a failover costs at most 150 ms before the table read starts.

After a DynamoDB read, the outcome is decided once, in one pure function, from the item's `is_active`, `expires_at` and the current time:

- **Active and not expired:** populate Redis with an active record and return 302.
- **Unknown or expired:** return 404 and cache nothing. This is invariant 2. A cached miss would hide an alias created a second later. Expired links are not cached either, because a tombstone for a link that simply ran out would be a 24-hour entry for nothing.
- **Inactive:** write a tombstone and return 404. From now on the table is not consulted for this code for 24 hours.

Everything above is deterministic. The one judgement call is what to do when DynamoDB itself fails.

## Failure mode 1: the database is down and the cache is old

**When is it safe to serve a stale destination?**

DynamoDB is very reliable, but the redirect function talks to it through a VPC endpoint with a 1-second budget, and throttles or a slow partition on a viral key are real. When the revalidation read fails, there are three options.

| Option | Problem |
| :-- | :-- |
| Return 503 | Every click on a link whose cache entry passed `fresh_until` in the last few minutes fails, while a perfectly good destination sits in Redis. |
| Serve the stale record, whatever it is | A deleted link whose tombstone was never written, or whose cache entry is from before the delete, redirects again. That breaks invariant 1. |
| **Serve the stale record only inside a bounded window (chosen)** | The stale destination is served only if the record is `active`, `now < stale_until`, and `now < expires_at`. A tombstone is always a 404. Anything else is 503 with `no-store`. |

The window is the 15 minutes between `fresh_until` and `stale_until`. Within it, the worst case is that a link deleted in the last 15 minutes, whose tombstone write also failed, is served during a DynamoDB outage. That is a double failure of two independent systems inside a short window, and the alternative is failing every revalidating click for the length of the outage. I accepted it, and the handler logs a warning on every stale serve so the count is visible.

What the window never does is override a tombstone. If Redis says `deleted`, the answer is 404 no matter what DynamoDB is doing. The tombstone is a positive fact, and the stale path only ever fills in for a missing one.

## Failure mode 2: the slow writer

**Can a cache populate undo a delete?**

Yes, if writes are plain `SET`s. The redirect function and the invalidation function both write the same key, and they are not coordinated. Consider a click and a takedown that overlap:

![A late cache populate overwrites a tombstone unless writes compare versions](https://www.sumonselim.com/images/articles/url-shortener/tombstone-vs-late-populate.svg "Figure 3: Left: a redirect reads the active item from DynamoDB, a takedown writes the tombstone, then the redirect's populate lands and replaces it. The deleted link redirects for up to 24 hours. Right: every write runs one Lua script that rejects a lower version, so the tombstone survives.")

The redirect read the item while it was still active. The takedown ran, wrote the DynamoDB update and the Redis tombstone. Then the redirect's populate, carrying the old active record, arrived and overwrote the tombstone. The link is dead in the table and alive in the cache for the next 24 hours. Every click is a cache hit, so nothing ever asks the table again.

**The fix is a version on every record and a compare on every write.** DynamoDB increments `version` in the same conditional update that flips `is_active`. Both the redirect and the invalidation function copy that version into the record they write. The write itself is one Lua script:

```lua
local existing = redis.call('GET', KEYS[1])
if existing then
  local ok, decoded = pcall(cjson.decode, existing)
  if ok and type(decoded) == 'table' and decoded['version'] then
    if tonumber(decoded['version']) > tonumber(ARGV[2]) then
      return 0                                  -- stored record is newer: keep it
    end
  end
end
redis.call('SET', KEYS[1], ARGV[1], 'PX', tonumber(ARGV[3]))
return 1
```

Redis runs a script from start to finish before anything else, so the read and the write cannot be separated by another writer. A separate `GET` followed by a `SET` in Go would reopen exactly the race in the figure. Equal versions are allowed to overwrite, so a retried write is idempotent (safe to run more than once, with the same result), and an active populate carrying the same version as the stored active record simply refreshes its timestamps.

The rule I took from this is the same one that decided the seat reservation system's admission script: **a guard and the action it guards must live in the same atomic step.** A check followed by a separate write is a race, and under load every race is lost sooner or later.

## Failure mode 3: the delete that only half happened

**What if the table is updated but the cache is not?**

A delete has two writes in two systems, and the second one can fail: the invalidation function can time out, Redis can be mid-failover, the invoke can be throttled. If the operation reports success anyway, the link is dead in DynamoDB and alive in Redis for up to 24 hours. That is invariant 1 broken by a partial failure rather than by a race.

![An operator takedown, end to end](https://www.sumonselim.com/images/articles/url-shortener/takedown-flow.svg "Figure 4: The admin CLI takes its identity from STS, soft-deletes the item with a conditional update, then synchronously invokes the invalidation function, which writes the versioned tombstone. Only then is the audit event recorded as a success. If invalidation fails, the call returns 503 and must be retried.")

The order is deliberate, and each step is idempotent:

1. **Identity comes from AWS, not from a flag.** The admin CLI calls `GetCallerIdentity` and uses the returned ARN as the actor. There is no `--actor` option to forge. A reason is mandatory.
2. **The durable delete is a conditional update.** `is_active` becomes false and `version` becomes `old + 1`, on the condition that the item is still active with the old version. A retry after a partial failure finds the item already inactive and returns the stored deletion, with its version and timestamps, instead of failing. A retry by a different actor is refused.
3. **Invalidation is synchronous.** The API role and the operator role are the only principals allowed to invoke the invalidation function, and they invoke it with `RequestResponse` and wait. The payload is `{short_code, version, deleted_at}`. Success means the tombstone with that version is in Redis. An older invalidation can never overwrite a newer record, because of the script above.
4. **A failed invalidation is a failed delete.** The operation returns 503, records an audit event with outcome `invalidate_failed`, and raises the alarm that security owns. The caller retries. The retry skips the state change, reissues the same tombstone, and returns success once it lands. The item is never reported deleted while a known positive cache entry may remain.
5. **The edge takes care of itself.** CloudFront caches a 302 for 5 seconds and nothing else. There is no invalidation API call, no distribution-wide purge, and no `stale-while-revalidate` that could extend the window.

The same service runs an owner's own delete when authenticated ownership returns. In the current release, create and stats are public, so there is no owner to authorise, and the only delete is the operator's.

## Why not DAX, and why not a simpler Redis

**Was all of this necessary?**

Part 1 listed the reasons DAX could not replace Redis here: no explicit deleted state, no version comparison, no bounded stale window. It is worth restating them from this side, because each one maps to an invariant.

| Invariant | What the cache must support | DAX | Redis with this record |
| :-- | :-- | :-- | :-- |
| A deleted link stays deleted | A positive "deleted" answer that survives late writers | No: a deleted item is just a miss, and a miss goes to the table | Tombstone plus version compare |
| A new link is visible at once | No negative caching of misses | Item cache only, so misses are not cached, which is fine | Unknown codes are never written |
| An expired link is not served | The record carries the link's own expiry and the read checks it | The item carries it, but the cache TTL is separate | `expires_at` in the record, checked on every read, and the key TTL is capped at it |

A simpler Redis layout would also have worked for the first two rows, using two keys per code or a plain `SET` with a short TTL. The third row and Failure mode 2 are what forced the versioned single record. Once every write is one script that compares versions, the tombstone, the stale window and the takedown retry all fall out of the same mechanism.

Two configuration details finish the picture. Redis runs with in-transit and at-rest encryption and an AUTH token held in an encrypted parameter. And eviction is allowed: any record can be evicted under memory pressure, because eviction is just a miss, and a miss goes to DynamoDB. That is the opposite of the seat reservation cache, where eviction would have lost the source of truth. Here the source of truth is the table, and the cache is only ever a copy.

## How it is tested

**How do you convince yourself a cache is right?**

- **The redirect decision is a pure function** with a table test for every branch: active and not expired, active and expired, inactive, unknown.
- **The handler is tested against in-memory fakes** of the store, the cache and the publisher, for every outcome in Figure 2: fresh, stale with a healthy table, stale with a failing table inside and outside the grace window, tombstone with a failing table, miss, and a Redis error.
- **The Redis adapter is tested against a fake Redis server in-process.** The tests cover the version script directly: a lower version is rejected, an equal version is accepted, a tombstone survives a late active write, and the key TTL never exceeds the link's own expiry.
- **The takedown path is tested for partial failure:** a delete whose invalidation fails returns a dependency error, records the failed-invalidation audit event, and succeeds on retry without a second state change.

None of these tests needs a cloud account, a network, or a running emulator. Part 3 explains why the code is structured to make that possible.

## What's next

[Part 3](https://www.sumonselim.com/url-shortener-part3-writes-clicks-and-the-slow-path) follows a link from creation to its click count. It covers how a create request is validated without ever fetching the destination, how an idempotency key and a generated code are written in one transaction, how the click stream is kept off the redirect and aggregated in batches, what the abuse controls are for a public API with no accounts, and how the whole thing is tested in layers without a cloud account.
