Skip to main content

Architecting a Production URL Shortener at 100K QPS Scale

NR Tech Studio Team
NR Tech Studio Team NR Tech Studio
18 min read

A production URL shortener operating at 100,000 read queries per second faces a classic distributed systems dilemma: sub-5ms redirection latency requires aggressive caching and non-blocking reads, while enterprise compliance demands accurate click analytics, automated abuse mitigation, and strictly unique key generation. When traffic spikes or viral links saturate individual cache keys, trivial hash-and-check schemes collapse under database write locks and collision cascades.

Building a resilient shortening infrastructure requires decoupling ingestion from redirect evaluation, treating key generation as a distributed consensus problem, and isolating telemetry into asynchronous log-structured pipelines. Toy implementations relying on random Base62 generation against an auto-incrementing SQL primary key introduce single points of failure, partition hot-spots, and unpredictable tail latencies under real-world loads.

This architectural guide provides the end-to-end engineering blueprint for an enterprise URL shortening platform. It covers mathematical capacity planning, distributed token-range coordination using etcd, multi-region edge caching strategies, deep database benchmark comparisons, and zero-latency clickstream telemetry with Apache Kafka and ClickHouse.

System Scoping and Capacity Estimates: Sizing a 100K QPS Service

Before writing a line of code or deploying infrastructure, you must size the system across computational, network, and storage dimensions. When you design a url shortening service at tier-one scale, read traffic dwarfs write traffic. We establish an architectural baseline based on an asymmetric 100:1 read-to-write ratio.

Traffic and Bandwidth Math

Assuming a sustained baseline of 100,000 read queries per second (QPS) and a 100:1 read-to-write ratio:

  • Write Throughput: 100,000 / 100 = 1,000 write QPS. Peak write bursts are calculated at 2.5x, or 2,500 write QPS.
  • Read Throughput: 100,000 read QPS baseline, scaling to a peak traffic factor of 2x (200,000 read QPS).
  • Ingress Bandwidth: Each write payload contains the long URL (average 500 bytes), custom alias requests, and metadata (200 bytes). At 2,500 peak write QPS: 2,500 * 700 bytes = 1.75 MB/sec (14 Mbps).
  • Egress Bandwidth: Redirection responses emit HTTP 302 headers containing the Location header, cookie tracking markers, and cache directives (roughly 600 bytes total payload). At 200,000 peak read QPS: 200,000 * 600 bytes = 120 MB/sec (960 Mbps).

5-Year Storage Capacity Projections

Storage sizing dictates whether the service can reside entirely in memory or if it requires multi-tiered cold storage. Over a standard 5-year retention lifecycle:

  • Total Generated URLs: 1,000 writes/sec * 86,400 sec/day * 365 days * 5 years = 157.68 billion records.
  • Record Schema Footprint:
    • Short Key (7 chars Base62): 7 bytes
    • Original Long URL: 500 bytes (average)
    • Created Timestamp (Unix Epoch, int64): 8 bytes
    • Expiration Timestamp (Unix Epoch, int64): 8 bytes
    • User/Account UUID: 16 bytes
    • Status/Flags (int8): 1 byte
    • Row Overhead/Index Pointers: ~60 bytes
    • Total per Record: ~600 bytes
  • Raw Persistent Storage: 157.68 billion * 600 bytes = 94.6 TB.
  • Replication Overhead: Factoring in a multi-region Raft or quorum replication factor of 3 (RF=3), persistent disk demands reach 283.8 TB without accounting for indexes or telemetry logs.
Metric Baseline Value Peak Multiplier Peak Engineering Spec
Write QPS 1,000 req/sec 2.5x 2,500 req/sec
Read QPS 100,000 req/sec 2.0x 200,000 req/sec
Network Egress 60 MB/sec 2.0x 120 MB/sec (0.96 Gbps)
5-Year Data Volume 94.6 TB N/A 283.8 TB (Replication Factor = 3)
Working Set Cache (80/20) 3.45 TB / day 1.0x 3.45 TB RAM in Redis cluster

The 80/20 Memory Cache Sizing Rule: Roughly 20% of generated URLs account for 80% of daily read redirection volume. Daily read volume equals 100,000 * 86,400 = 8.64 billion requests. If 20% of the active daily working set requires in-memory hot caching, and each cache entry consumes 1 KB in Redis with hash table overheads, the caching layer must allocate at least 3.45 TB of dedicated RAM across the global Redis cluster to prevent memory evictions from thrashing the underlying database.

Architectural Design Constraints

  • Redirection Latency: p95 < 5ms, p99 < 15ms globally.
  • Availability: 99.999% uptime (less than 5.26 minutes of unscheduled downtime annually).
  • Predictable Key Format: Uniform 7-character string output supporting standard HTTP URI encoding without reserved character escaping.
  • Durability: Zero link loss once an HTTP 201 Created acknowledgment is returned to the client.

End-to-End URL Shortener System Design and Request Lifecycles

A high-performance url shortener system design relies on a strict separation of concerns between state mutation (link creation) and state resolution (link redirection). Coupling these execution pathways inside the same application runtime leads to thread pool starvation when traffic spikes on hot URLs.

To design a url shortener capable of global low-latency redirects, the edge layer must resolve links as close to the user as possible, reserving core application services strictly for cache misses and write ingestions.

[ Client Browser / API Consumer ] ── (Anycast DNS / Cloudflare Edge) ──┐ 302 Redirect (Cache Hit) 5ms p99
 │
 ▼
 ┌─────────────────────────────┐
 │ Global L4/L7 ALB Cluster │
 │ (Envoy / TLS Termination)│
 └──────────────┬──────────────┘
 │
 ┌────────────────────────────────────────────────────┴────────────────────────────────────────┐
 │ │
 ▼ (Read Path: GET /{slug}) ▼ (Write Path: POST /v1/links)
 ┌───────────────────────────┐ ┌───────────────────────────┐
 │ Redirection API Workers │ │ Link Ingestion API │
 │ (Go / Fiber / Actix-Web) │ │ (Go / Tokio Engine) │
 └─────────────┬─────────────┘ └─────────────┬─────────────┘
 │ │
 ┌─────────┴─────────┐ ┌───────┴───────┐
 ▼ ▼ ▼ ▼
┌───────────────┐ ┌───────────────┐ ┌───────────────┐ ┌───────────────┐
│ Redis L1/L2 │ │ Storage Read │ │ Key Gen (KGS) │ │ Threat & Phish│
│ Cache Cluster │ │ Replicas │ │ Lease Client │ │ Async Scanner │
│ (Sub-1ms Read)│ │ (ScyllaDB) │ │ (etcd ranges) │ │ (SafeBrowsing)│
└───────┬───────┘ └───────────────┘ └───────┬───────┘ └───────────────┘
 │ │
 ▼ (Async Metric Stream) ▼
┌────────────────────────────────────────┐ ┌───────────────┐
│ Kafka Event Stream (analytics.clicks)│ │ ScyllaDB Core │
└───────────────────┬────────────────────┘ │ Primary Node │
 ▼ └───────────────┘
┌────────────────────────────────────────┐
│ ClickHouse Real-time Analytics Engine │
└────────────────────────────────────────┘

Detailed Component Request Lifecycles

  1. Write Request Flow (Creation):
    • The client sends a POST /v1/links request with payload {"url": "https://long-domain.com/path", "custom_alias": null}.
    • The API Gateway terminates TLS, enforces token-bucket rate limiting based on the client IP and API key, and routes the request to an Ingestion Service instance.
    • The Ingestion Service extracts an unused 7-character key from its local in-memory token buffer, previously acquired via lease from the Key Generation Service.
    • The URL format and scheme are sanitized. The payload is concurrently dispatched to an asynchronous security scanning pipeline to inspect for phishing, malware, and redirect loops.
    • The record is committed to the primary distributed storage cluster (e.g. ScyllaDB) with monotonic write consistency.
    • The new mapping is immediately injected into the local Redis cluster using a write-through or warm-up strategy to eliminate cold misses on instant shares.
    • An HTTP 201 Created status is returned to the caller with the shortened URL entity.
  2. Read Request Flow (Redirection):
    • The client requests GET /a8X9k1z via browser navigation.
    • The edge Anycast layer directs traffic to the geographically closest point of presence (PoP). If the edge proxy holds the key in memory, it terminates the request with an HTTP 302 instantly.
    • On an edge miss, the request traverses the application load balancer to the Redirection Service.
    • The service performs an in-memory lookup against the primary Redis cluster. On a cache hit (98%+ probability in production), the service prepares the HTTP 302 redirect.
    • On a cache miss, the service queries the read replica of the persistent data store, populates Redis using a singleflight mutex to prevent cache stampedes, and forms the HTTP response.
    • Before closing the connection, the Redirection Worker dispatches a lightweight, non-blocking telemetry event to an Apache Kafka topic (analytics.clicks) containing the request headers, IP address, timestamp, and referer.
    • The HTTP 302 response with the Location: https://long-domain.com/path header is returned to the user agent.

Architectural Segregation of Concerns: Never place ingestion handlers and redirect endpoints within the same process thread pools. An influx of heavy POST payloads requiring JSON validation, regex sanitization, and database durability checks must never exhaust worker threads tasked with resolving high-throughput HTTP 302 GET redirects.

Key Generation Service (KGS) vs Base62: Eliminating Distributed Collision

A critical challenge when teams examine tinyurl system design lies in generating high-entropy, short, collision-free tokens under horizontal scaling. Naive solutions calculate a cryptographic hash (such as MD5 or SHA-256) of the destination URL and truncate it to the first 7 characters. This approach introduces unacceptable flaws at enterprise scale.

Why Hash Truncation Collapses

A 7-character Base62 string provides 62^7 = 3,521,614,606,208 (3.52 trillion) permutations. However, MD5 produces a 128-bit space. Truncating an MD5 hash down to 42 bits of entropy guarantees hash collisions under the birthday paradox after roughly sqrt(62^7) = 1.87 million keys. Resolving collisions requires the application to append a salt or counter and re-query the database in an iterative loop:

Loop until insertion success:
 1. candidate_key = Truncate(Hash(url + counter))
 2. SELECT key FROM urls WHERE key = candidate_key
 3. If exists: counter++, continue
 4. Else: INSERT INTO urls..

This design introduces unpredictable tail latencies, database read contention, and eventual deadlock during concurrent writes of duplicate URLs.

Pre-Generated Key Pools with Distributed Consensus

To eliminate runtime collision detection entirely, the system must treat key assignment as a sequence consumption problem. A dedicated Key Generation Service (KGS) pre-generates sequences or manages monotonic counters, vending blocks of unallocated keys to application workers using distributed leases managed by etcd or Apache ZooKeeper.

Rather than making a distributed network call for every single write, worker nodes acquire a discrete range (e.g. node 1 claims [1,000,000 to 1,999,999]; node 2 claims [2,000,000 to 2,999,999]). The worker loads this range into an atomic in-memory counter, converts the integer sequence directly to Base62, and yields keys locally with sub-microsecond latency and zero cross-node locks.

Generation Strategy Collision Probability Runtime Latency Coordination Overhead Crash Recovery Behavior
MD5 / SHA-256 Truncation High (Birthday Paradox) Unbounded (O(N) db checks) None (Stateless) Idempotent recalculation
Snowflake IDs (64-bit) Zero (Unique bitfields) Sub-millisecond Clock sync required (NTP drift) Safe if clock strictly monotonic
Distributed KGS (etcd leases) Mathematically Zero Nanoseconds (Local RAM) Lease renewal per 1M keys Abandon unassigned block on crash

Worker Crash Failure Semantics: A frequent objection in tiny url system design interviews is: What happens if a worker node crashes with 800,000 unallocated keys in its assigned range? The answer is simple: you let those keys go. The total Base62 address space (3.52 trillion keys) is vast. Discarding even tens of millions of unused keys across catastrophic node restarts over five years burns less than 0.001% of the total key capacity while guaranteeing absolute uniqueness without requiring complex multi-node transaction recovery.

Production Base62 Implementation in Go

The following Go implementation demonstrates a production-grade, thread-safe Base62 token encoder and decoder operating over a sequential range engine, incorporating atomic operations for zero lock contention.

package main

import (
 "errors"
 "fmt"
 "strings"
 "sync/atomic"
)

const (
 base62Alphabet = "0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ"
 base = uint64(len(base62Alphabet))
 tokenLength = 7
)

var (
 ErrRangeExhausted = errors.New("assigned key range exhausted")
 ErrInvalidToken = errors.New("token contains invalid Base62 characters")
)

// TokenRangeManager holds an atomically safe sequence allocated from etcd leases.
type TokenRangeManager struct {
 current uint64
 max uint64
}

func NewTokenRangeManager(start, max uint64) *TokenRangeManager {
 return &TokenRangeManager{
 current: start,
 max: max,
 }
}

// NextID atomically reserves the next 64-bit integer in the assigned lease range.
func (m *TokenRangeManager) NextID() (uint64, error) {
 for {
 val:= atomic.LoadUint64(&m.current)
 if val > m.max {
 return 0, ErrRangeExhausted
 }
 if atomic.CompareAndSwapUint64(&m.current, val, val+1) {
 return val, nil
 }
 }
}

// EncodeBase62 transforms an unsigned 64-bit integer into a left-padded 7-character string.
func EncodeBase62(num uint64) string {
 if num == 0 {
 return strings.Repeat("0", tokenLength)
 }

 var sb strings.Builder
 for num > 0 {
 rem:= num % base
 sb.WriteByte(base62Alphabet[rem])
 num = num / base
 }

 // Reverse the encoded characters
 bytes:= []byte(sb.String())
 for i, j:= 0, len(bytes)-1; i < j; i, j = i+1, j-1 {
 bytes[i], bytes[j] = bytes[j], bytes[i]
 }

 res:= string(bytes)
 if len(res) < tokenLength {
 res = strings.Repeat("0", tokenLength-len(res)) + res
 }
 return res
}

// DecodeBase62 converts a Base62 token back into a 64-bit sequence integer.
func DecodeBase62(token string) (uint64, error) {
 var result uint64
 for i:= 0; i < len(token); i++ {
 char:= token[i]
 idx:= strings.IndexByte(base62Alphabet, char)
 if idx == -1 {
 return 0, ErrInvalidToken
 }
 result = result*base + uint64(idx)
 }
 return result, nil
}

func main() {
 // Simulate a worker allocated range [100,000,000 to 100,000,010] by etcd
 manager:= NewTokenRangeManager(100000000, 100000010)

 for i:= 0; i < 5; i++ {
 seq, err:= manager.NextID()
 if err!= nil {
 panic(err)
 }
 slug:= EncodeBase62(seq)
 decoded, _:= DecodeBase62(slug)
 fmt.Printf("Seq: %d => Base62: %s => Decoded: %d\n", seq, slug, decoded)
 }
}

Storage Engine Trade-Offs: DynamoDB vs ScyllaDB vs Partitioned PostgreSQL

When selecting the persistence layer for your tinyurl design, the data access pattern must dictate the engine choice. A URL shortener represents an archetypal point-lookup workload: reads and writes access data strictly by a single key (the 7-character slug), with virtually no requirements for complex joins, multi-table transactions, or foreign key constraints.

Storage Engine Comparison Matrix

Evaluation Metric Partitioned PostgreSQL Amazon DynamoDB ScyllaDB (Wide-Column)
Read Latency (p99) 8ms to 25ms 4ms to 12ms (Single-digit with DAX) Sub-2ms (Kernel-bypass / Seastar)
Write Throughput Scale Limited by Primary IOPS Virtually infinite (Auto-partitioning) Linear scale across bare-metal nodes
Partition Hot-Spotting High (B-Tree index bloat) Throttles on single partition partition key Mitigated via token ring hashing
Hosting Economics Low at small scale, high at petabyte High (Read/Write capacity unit costs) Optimal (Maximum IOPS per hardware dollar)
Maintenance Burden Vacuuming, failovers, reindexing Zero (Fully managed serverless) Moderate (Cluster maintenance, node repair)

Why Relational Databases Break Down at Scale

Relational databases such as PostgreSQL or MySQL function well in early-stage architectures. However, as the table size scales past 50 billion rows (representing tens of terabytes of index data), performance degrades:

  • B-Tree Index Eviction: When the secondary B-Tree indexes on short_key exceed available RAM, every index lookup triggers an expensive random NVMe SSD read. Read latency degrades from under 2ms to over 30ms.
  • Write Amplification during Re-Indexing: High-throughput concurrent writes cause index node splits, write amplification, and severe table vacuum bloat.
  • Manual Sharding Fragility: Sharding PostgreSQL requires maintaining an application-level routing layer (e.g. using consistent hashing on the short key) or adopting third-party coordinators like Citus, increasing operational failure surfaces.

The Case for ScyllaDB or Managed DynamoDB

For globally distributed scale, distributed wide-column stores (ScyllaDB/Apache Cassandra) or managed distributed key-value databases (Amazon DynamoDB) are the superior choice. ScyllaDB is implemented in C++ using an asynchronous, shared-nothing thread-per-core architecture based on the Seastar framework. It bypasses the Linux kernel page cache entirely using direct I/O (O_DIRECT), preventing garbage collection pauses and ensuring deterministic sub-2ms read latencies under 100K+ QPS.

Data Partitioning Strategy: In ScyllaDB, designate the 7-character short_key as the primary partition key. The cluster applies a Murmur3 hash to the token to place records uniformly across the distributed token ring, completely avoiding partition skew even under massive write concurrency.

-- Production ScyllaDB / Cassandra Table Schema
CREATE KEYSPACE url_shortener WITH replication = {
 'class': 'NetworkTopologyStrategy',
 'us-east-1': 3,
 'eu-central-1': 3
};

CREATE TABLE url_shortener.url_mappings (
 short_key varchar,
 destination_url varchar,
 account_id uuid,
 created_at timestamp,
 expires_at timestamp,
 is_active boolean,
 PRIMARY KEY (short_key)
) WITH default_time_to_live = 0
 AND compaction = {'class': 'LeveledCompactionStrategy'}
 AND comment = 'Core URL redirection mapping lookup store';

Redirect Mechanics: HTTP 301 vs 302 and Edge Redis Caching Architecture

The choice of HTTP status code fundamentally alters network path behavior, origin load, and analytics fidelity when you design url redirection infrastructure.

HTTP 301 vs HTTP 302 / 307: The Architectural Trade-Off

  • HTTP 301 (Moved Permanently): The HTTP 301 status signals to intermediate proxies and client browsers that the requested URI has permanently migrated to the target address. The browser caches this association locally in its persistent disk cache. Subsequent navigations to the short link never hit your application origin or edge CDN. While this minimizes your infrastructure compute load, it completely blinds your analytics pipeline: you will register zero click metrics, geo-telemetry, or user-agent tracking for repeat visits. Furthermore, if a destination URL must be revoked due to malware or updated after an error, the change cannot be forced onto clients with cached 301 records.
  • HTTP 302 (Found) / HTTP 307 (Temporary Redirect): These status codes instruct the client to preserve the original method and check the target resource on every invocation. The client does not cache the destination permanently. Every click traverses your edge or origin infrastructure, allowing complete telemetry capture, instant revocation of malicious links, and dynamic expiration validation at the cost of handling inbound HTTP requests on every single user click.

Production Standard: Enterprise URL shorteners standardize exclusively on HTTP 302 (or HTTP 307 to preserve HTTP verbs on API endpoints). If downstream caching is acceptable for non-tracked links, set an explicit Cache-Control: private, max-age=90 header, allowing client browsers to cache the resolution for only 90 seconds.

Mitigating Cache Stampedes: Probabilistic Early Expiration

When a globally trending link expires from an in-memory cache, hundreds of concurrent requests for that key miss simultaneously and hit the persistent database layer at the exact same millisecond. This phenomenon, known as a cache stampede or the thundering herd problem, can saturate database connection pools and take down read replicas.

To prevent this, production caching layers implement singleflight mutexes or probabilistic early expiration using the XFetch algorithm.

import math
import random
import time
import redis

class ResilientURLCache:
 def __init__(self, redis_client: redis.Redis, delta: float = 1.0):
 self.client = redis_client
 self.delta = delta # Beta factor configuring refresh aggressiveness

 def get_url(self, short_key: str, fallback_db_fetcher) -> str:
 """
 Implements the XFetch probabilistic early expiration algorithm.
 Prevents cache stampedes by recomputing keys in the background 
 before they hard-expire in Redis.
 """
 cache_key = f"url:{short_key}"
 # Hash structure stores: url, ttl_remaining, compute_time_ms
 data = self.client.hgetall(cache_key)

 if data:
 dest_url = data.get(b"url").decode("utf-8")
 ttl = self.client.ttl(cache_key)
 compute_time = float(data.get(b"compute_time", 0.05))
 
 # XFetch condition: -(compute_time * beta * ln(random())) > ttl
 # If true, refresh early before expiration occurs
 if ttl > 0 and -(compute_time * self.delta * math.log(random.random())) > ttl:
 # Re-fetch in background or handle proactively
 self._refresh_key(cache_key, short_key, fallback_db_fetcher)
 
 return dest_url

 # Complete cache miss: fetch from persistent database
 return self._refresh_key(cache_key, short_key, fallback_db_fetcher)

 def _refresh_key(self, cache_key: str, short_key: str, fetcher) -> str:
 start_time = time.time()
 dest_url = fetcher(short_key)
 compute_time = time.time() - start_time

 if dest_url:
 pipe = self.client.pipeline()
 pipe.hset(cache_key, mapping={
 "url": dest_url,
 "compute_time": compute_time
 })
 pipe.expire(cache_key, 86400) # 24-hour default TTL
 pipe.execute()
 
 return dest_url

Edge Caching Architecture Checklist

  • Active-Active Multi-Region Cluster: Deploy Redis instances in all operating regions, synchronizing updates via an event stream or relying on local cache warming to isolate failure domains.
  • Memory Eviction Policy: Enforce volatile-lru (least recently used among keys with an explicit expiration) or allkeys-lru to protect Redis against out-of-memory panics.
  • Singleflight Request Coalescing: Implement Go’s singleflight.Group inside your redirection worker services so that only one database query executes per missing key, even when 1,000 requests hit an unprimed worker concurrently.
  • Connection Pooling: Maintain persistent TCP connections to Redis with pre-allocated connection pools to avoid three-way handshake overhead on high-throughput read paths.

Decoupled Click Telemetry and Enterprise Abuse Prevention Pipelines

A high-performance URL shortener must execute two asynchronous workflows during a redirect without increasing user-perceived latency: logging analytical events and preventing the platform from distributing malicious exploits or phishing links.

Zero-Overhead Asynchronous Clickstream Telemetry

Recording user IP addresses, geolocations, referral domains, and device fingerprints inside the critical path of an HTTP redirect is an anti-pattern. Directly querying analytics databases or inserting click rows synchronously adds tens of milliseconds of latency to every redirect and causes total service failure if the analytics database undergoes maintenance.

Instead, the redirection worker emits an asynchronous event to an Apache Kafka or AWS Kinesis topic over a local non-blocking producer buffer. Specialized ingestion consumers consume the stream in micro-batches and persist raw records into a columnar database optimized for real-time analytics, such as ClickHouse.

  1. Edge Event Ingestion: The redirection worker yields the HTTP 302 response to the client immediately. In a detached background goroutine, it dispatches an Avro or Protocol Buffer binary record to the analytics.clicks Kafka topic.
  2. Partitioning Strategy: Kafka messages are partitioned using the short_key as the partition key. This guarantees that all clicks for a given link land in the same Kafka partition, enabling efficient sliding-window aggregations without cross-partition shuffling.
  3. Micro-Batch Loading to ClickHouse: A stateless consumer group consumes events using Kafka Engine tables in ClickHouse or an intermediate Vector/FluentBit aggregator, inserting batches of 50,000 records at a time into a MergeTree table.
  4. Materialized Views: ClickHouse continuously computes analytical rollups (e.g. clicks per country, referer counts, daily active totals) via SummingMergeTree materialized views, allowing enterprise dashboards to query billions of click events in milliseconds without scanning raw click logs.
-- Production ClickHouse Analytical Schema for URL Telemetry
CREATE TABLE default.link_clicks (
 short_key LowCardinality(String),
 clicked_at DateTime64(3, 'UTC'),
 ip_address IPv4,
 country_code LowCardinality(FixedString(2)),
 user_agent String,
 referer String
)
ENGINE = MergeTree()
PARTITION BY toYYYYMM(clicked_at)
ORDER BY (short_key, clicked_at);

-- Pre-aggregated Real-time Click Counters
CREATE TABLE default.link_clicks_hourly (
 short_key LowCardinality(String),
 hour_bucket DateTime,
 total_clicks SimpleAggregateFunction(sum, UInt64)
)
ENGINE = SummingMergeTree()
ORDER BY (short_key, hour_bucket);

CREATE MATERIALIZED VIEW default.mv_link_clicks_hourly TO default.link_clicks_hourly AS
SELECT
 short_key,
 toStartOfHour(clicked_at) AS hour_bucket,
 count() AS total_clicks
FROM default.link_clicks
GROUP BY short_key, hour_bucket;

Automated Abuse, Phishing, and Crawling Mitigations

Public URL shorteners are prime targets for malicious actors seeking to mask phishing domains, bypass email spam filters, and distribute malware. A production service must defend itself proactively across multiple security perimeters:

  • Edge Rate Limiting: Use a distributed token-bucket algorithm running in Redis or Envoy to limit link creation per IP address and account ID (e.g. maximum 50 links per minute for authenticated users, 5 links per minute for anonymous users).
  • Deterministic Blacklist Regex Scanning: Incoming destination URLs are matched against a high-performance Bloom filter loaded in memory containing millions of known malicious domains and IP ranges.
  • Asynchronous Threat Verification: During the link creation lifecycle, the URL is submitted to an asynchronous evaluation queue consumed by security workers. These workers query external threat intelligence feeds (Google Safe Browsing API, PhishTank) and follow redirection hops to ensure destination pages do not perform drive-by payload downloads.
  • Dynamic Safe Link Gating: If a link is flagged by the security pipeline post-creation, its status flag in ScyllaDB and Redis is immediately flipped to SUSPENDED. Subsequent requests bypass the HTTP 302 redirect and instead render a warning landing page alerting the visitor to a compromised or dangerous link.

Frequently Asked Questions

Should a URL shortener use HTTP 301 or HTTP 302 redirects?

Use HTTP 301 permanent redirects if you want browsers to cache the destination URL locally and minimize origin traffic. Use HTTP 302 or 307 temporary redirects if your business requires capturing precise click analytics, user tracking, and real-time link expiration validation on every request.

How does a Key Generation Service prevent duplicate keys if a worker crashes?

A Key Generation Service mitigates worker crashes by allocating distinct token ranges using distributed consensus leases via etcd or ZooKeeper. If a worker dies mid-batch, unassigned keys in its volatile memory are discarded rather than reused, completely eliminating duplicates across the distributed system.

How many characters should a short URL slug contain?

A 7-character Base62 string (using characters a-z, A-Z, and 0-9) provides 62 to the power of 7, or approximately 3.5 trillion unique link combinations. At an ingestion rate of 1,000 new URLs per second, this slug length guarantees over 110 years of unique identifiers without collisions.

How do you protect a public URL shortener from malware and phishing abuse?

Integrate an asynchronous security inspection pipeline using message queues like Kafka. Every destination URL is verified against threat intelligence APIs such as Google Safe Browsing and checked with heuristic regex scanners before being permanently flagged as safe for downstream public redirects.

Designing a production URL shortener capable of sustaining 100,000+ read QPS requires shifting your focus away from simplistic key-shortening formulas toward fundamental distributed systems principles. By replacing runtime hash collision strategies with a Key Generation Service managing distributed etcd lease ranges, you eliminate database lock contention entirely. Moving away from monolithic relational schemas to kernel-bypass wide-column stores like ScyllaDB guarantees deterministic sub-millisecond persistence under heavy write bursts.

On the read path, pairing edge Anycast routing with stampede-resistant Redis caching and decoupled Kafka-to-ClickHouse analytics pipelines ensures that global redirections resolve in single-digit milliseconds without degrading telemetry fidelity. These architectural patterns transform what is often treated as a trivial interview puzzle into a resilient, enterprise-grade distributed infrastructure platform.

References & Further Reading