Backend, APIs, and System Design

The State Machine Pattern: A Backend Engineer's Quiet Hero

A state machine is a pattern where a record can only exist in one of a defined set of states, and transitions between those states are explicit, validated, and logged. I use it any time a business object moves through a lifecycle. It turns implicit, scattered conditional logic into a single source of truth for how things change.

April 21, 2025 · 12 min read
Backend, APIs, and System Design

The Multi Tenant Database: One Schema or Many?

Multi-tenant database design is the set of architectural patterns for storing data from multiple customers (tenants) in a shared database infrastructure. The three primary patterns are shared tables (all tenants in the same tables, distinguished by a tenant_id column), schema per tenant (each tenant has a separate PostgreSQL schema within the same database), and database per tenant (each tenant has a separate database). Each pattern has different implications for data isolation, query complexity, operational overhead, and cost.

April 19, 2025 · 12 min read
Backend, APIs, and System Design

Zero Downtime Database Migrations: A Step By Step Guide

A zero downtime database migration is a sequence of small, backward compatible changes that lets the application run normally throughout. The trick is never doing the schema change and the code change in the same deploy. You expand, you backfill, you migrate reads, you stop writing the old way, then you contract. Each step is reversible. None of them is dramatic.

April 18, 2025 · 13 min read
Backend, APIs, and System Design

The Replication Lag Problem: How to Detect and Defend

Replication lag is the delay between a write being committed on the primary database and that write becoming visible on a replica. In most systems it stays under a second and nobody notices. When it grows to seconds or minutes, features that read from replicas show stale data in ways that look like bugs. I treat replication lag as an observable metric with alerting, not a background condition to ignore.

April 15, 2025 · 12 min read
Backend, APIs, and System Design

The Write Heavy Workload: A Different Set of Tradeoffs

A write heavy workload is one where the database spends most of its time accepting and persisting new data rather than serving reads. The tradeoffs are different from read heavy systems: indexes slow you down, normalization costs more, and the bottleneck is usually I/O throughput rather than query complexity. I approach it with a different toolset than I use for read heavy work.

April 13, 2025 · 11 min read
Backend, APIs, and System Design

The Read Heavy Workload: Strategies That Move the Needle

A read heavy workload is any system where reads outnumber writes by a significant ratio, typically ten to one or higher. The strategies that help are not all equal. I separate the ones that move the needle from the ones that look good in a talk but cost more to run than the problem they solve.

April 12, 2025 · 12 min read
Backend, APIs, and System Design

The Edge: When to Move Logic Off Your Origin

The edge is the network of CDN nodes distributed globally where request processing can happen before traffic reaches the origin server. Moving logic to the edge reduces latency for global users, offloads work from the origin, and can enforce security rules closer to the requester. Not all logic belongs at the edge: operations that require database access, full Node.js APIs, or shared state must stay at or near the origin.

April 6, 2025 · 12 min read
Backend, APIs, and System Design

The Outbox Pattern: A SaaS Reliability Cheat Code

The outbox pattern solves the dual write problem by storing side effects in the same database transaction as the business event, then delivering them asynchronously from a separate relay process. I use it whenever a service needs to publish to a message broker or call a webhook without risking a partially applied state change that leaves data inconsistent.

March 29, 2025 · 12 min read
Backend, APIs, and System Design

The N+1 Query Problem: Detection, Prevention, and Refactoring

The N+1 query problem occurs when an application executes one query to retrieve a list of N records, then executes N additional queries to retrieve related data for each record, producing N+1 total queries where one query with a join would have sufficed. At small scale, the performance impact is invisible. At production scale, N+1 queries are one of the most common causes of slow API endpoints and database overload. Detection requires query logging; prevention requires joins or data loaders; refactoring requires identifying which ORM calls produce multiple queries and replacing them with eager loading.

March 24, 2025 · 12 min read
Backend, APIs, and System Design

Webhooks vs Polling vs Server Sent Events vs WebSockets

Webhooks, polling, server sent events, and WebSockets are four ways to move data between a server and a client when something changes. Each has a different cost, reliability profile, and implementation complexity. I pick between them based on who initiates the connection, how often the data changes, and whether the channel needs to carry data in both directions.

March 20, 2025 · 13 min read
Backend, APIs, and System Design

The Public API Decision: When to Build One, When to Resist

A public API is a formal contract between your product and the outside world. Once live it carries a stability obligation that shapes every engineering decision downstream. I treat the decision to build one as a product launch, not a feature release, because the cost of a poorly designed public API compounds for years while the cost of waiting one more quarter almost never does.

March 19, 2025 · 11 min read
Backend, APIs, and System Design

The API Versioning Strategy That Survives Real World Use

API versioning is the contract between your product and every customer who has integrated with it. Break the contract without a plan and you break integrations. Ignore versioning entirely and you are unable to evolve the API. The strategy that survives real world use is the one that makes breaking changes explicit, gives customers a migration window, and keeps the surface area manageable.

March 15, 2025 · 12 min read