Security, Auth, and Compliance

The Threat Model: How to Build One in Two Hours

A threat model is a structured analysis of what could go wrong with a system, who might cause it, and how likely and damaging each scenario is. The point is not to document every possible attack. The point is to surface the areas most at risk so the team can prioritize security work against actual threats rather than generic best practices checklists.

May 28, 2025 · 13 min read
Security, Auth, and Compliance

The Permission System That Scales With Your B2B Customers

A B2B permission system that scales requires a layered model: system roles, organization roles, and resource level permissions. Flat role lists collapse under enterprise requirements. I build these with RBAC as the baseline, ABAC for attribute driven rules, and a clear separation between platform level and customer configurable permissions. That combination handles the first customer and the fiftieth without a rewrite.

May 22, 2025 · 12 min read
Security, Auth, and Compliance

The Bug Bounty Decision: When You Are Ready, When You Are Not

A bug bounty program invites security researchers to find vulnerabilities in your system in exchange for payment or recognition. Run one when you are ready and it strengthens your security posture. Run one before you are ready and you are paying people to find problems you already knew existed. The readiness decision is the decision.

May 15, 2025 · 12 min read
Security, Auth, and Compliance

Vulnerability Disclosure Programs: Why Even Small Teams Need One

A vulnerability disclosure program is a public document that tells security researchers how to report issues to you, what they can expect, and what is in scope. It is not a bug bounty. It does not have to pay anyone. It exists so that when a researcher finds a problem in your product, they have a path that does not end in a public tweet. The setup is hours. The protection is real.

May 14, 2025 · 11 min read
Security, Auth, and Compliance

The Single Tenant Argument: When Enterprise Customers Demand It

Single tenancy is a deployment model where a customer gets their own dedicated infrastructure rather than sharing a multitenant environment. Enterprise customers demand it for data isolation, regulatory compliance, or internal security policy. The decision to offer it is a product and pricing decision as much as a technical one, and the teams that handle it well have a clear answer ready before the first customer asks.

May 12, 2025 · 11 min read
Security, Auth, and Compliance

The Security Gap: How One Missing SOC 2 Control Kills Your Enterprise Deal

The security gap that kills an enterprise deal is almost never a fundamental security failure. It is a single missing SOC 2 control, a gap in access logging, an unwritten incident response plan, or a data retention policy that does not exist. Enterprise procurement teams work from checklists. The first item that cannot be answered is the item that stalls the deal, sometimes permanently.

May 6, 2025 · 12 min read
Backend, APIs, and System Design

Why Your Service Should Have Two Health Checks Not One

A service needs two health checks because liveness and readiness answer different questions. Liveness asks should this process be killed and restarted. Readiness asks should this process receive new traffic. A single endpoint that mixes the two will cause unnecessary restarts during transient dependency failures and is one of the most common quiet sources of self inflicted outages.

May 4, 2025 · 12 min read
Backend, APIs, and System Design

The Health Check Endpoint: Less Trivial Than It Looks

A health check endpoint is an API route that load balancers, orchestrators, and monitoring systems call to determine whether a service instance is ready to receive traffic. The trivial implementation, returning 200 OK immediately, is nearly useless for detecting real service failures. The correct implementation tests the dependencies that the service needs to function (database connectivity, cache availability, critical external APIs) and returns a structured response that distinguishes between healthy, degraded, and unhealthy states.

May 3, 2025 · 12 min read
Backend, APIs, and System Design

Timeouts: The Setting Most Engineers Get Wrong

A timeout is the maximum time your code is willing to wait before giving up and doing something else. Most engineers set them once, forget them, and discover they were wrong during an incident. The correct value depends on the operation, the SLA of the dependency, and what happens to the user or the data when the timeout fires. Getting it right requires thinking about all three, not just copying a number from a tutorial.

May 2, 2025 · 11 min read
Backend, APIs, and System Design

The Backend Engineer's Reading List for 2026

The backend engineer's reading list is shorter than most people think. Five to eight books cover the foundational ideas that show up in every production system. The rest is reading documentation and learning from production incidents. This list is the five to eight, with notes on when each one is most useful and which chapter to start with.

April 30, 2025 · 12 min read
Backend, APIs, and System Design

The Boring API: Why Predictability Beats Cleverness

A boring API is one where every developer who touches it can predict the behavior of any endpoint they have not seen before, based on what they learned from the first three. Consistent naming, consistent error responses, consistent pagination, consistent authentication. The boring API is not a technical achievement. It is a communication achievement. And it is the one that survives five years of customer integrations.

April 29, 2025 · 12 min read
Backend, APIs, and System Design

Why Logical Deletes Are Almost Always a Mistake

A logical delete, also called a soft delete, is a pattern where rows are never removed from the database. Instead, a deleted_at timestamp or is_deleted flag marks them as gone. The data stays. The complexity grows. Most teams adopt this pattern early and regret it late, when every query needs a WHERE deleted_at IS NULL and migrations become a negotiation with historical junk.

April 26, 2025 · 12 min read
Backend, APIs, and System Design

The Soft Delete Trap: A Pattern That Catches Up With Teams

Soft deletes mark a record as deleted with a flag or timestamp instead of removing it from the database. The pattern is popular because it feels safe and recoverable. In my experience it is safe at first and expensive later, as the deleted rows accumulate, queries get slower, and the team discovers that every feature touching that table needs to filter out the deleted rows or it shows deleted data.

April 25, 2025 · 11 min read
Backend, APIs, and System Design

Time Series Data in SaaS: When to Pull in TimescaleDB or InfluxDB

Time series data is data where the timestamp is the primary key and queries are almost always based on ranges: give me the values between now and thirty days ago. Postgres handles it adequately at low volumes. TimescaleDB extends Postgres with automatic partitioning and compression for time ordered data. InfluxDB is purpose built for metrics at high ingestion rates. Knowing which to reach for saves weeks of wasted infrastructure work.

April 22, 2025 · 12 min read
Backend, APIs, and System Design

The State Machine Pattern: A Backend Engineer's Quiet Hero

A state machine is a pattern where a record can only exist in one of a defined set of states, and transitions between those states are explicit, validated, and logged. I use it any time a business object moves through a lifecycle. It turns implicit, scattered conditional logic into a single source of truth for how things change.

April 21, 2025 · 12 min read
Backend, APIs, and System Design

The Multi Tenant Database: One Schema or Many?

Multi-tenant database design is the set of architectural patterns for storing data from multiple customers (tenants) in a shared database infrastructure. The three primary patterns are shared tables (all tenants in the same tables, distinguished by a tenant_id column), schema per tenant (each tenant has a separate PostgreSQL schema within the same database), and database per tenant (each tenant has a separate database). Each pattern has different implications for data isolation, query complexity, operational overhead, and cost.

April 19, 2025 · 12 min read
Backend, APIs, and System Design

Zero Downtime Database Migrations: A Step By Step Guide

A zero downtime database migration is a sequence of small, backward compatible changes that lets the application run normally throughout. The trick is never doing the schema change and the code change in the same deploy. You expand, you backfill, you migrate reads, you stop writing the old way, then you contract. Each step is reversible. None of them is dramatic.

April 18, 2025 · 13 min read
Backend, APIs, and System Design

The Replication Lag Problem: How to Detect and Defend

Replication lag is the delay between a write being committed on the primary database and that write becoming visible on a replica. In most systems it stays under a second and nobody notices. When it grows to seconds or minutes, features that read from replicas show stale data in ways that look like bugs. I treat replication lag as an observable metric with alerting, not a background condition to ignore.

April 15, 2025 · 12 min read
Backend, APIs, and System Design

The Write Heavy Workload: A Different Set of Tradeoffs

A write heavy workload is one where the database spends most of its time accepting and persisting new data rather than serving reads. The tradeoffs are different from read heavy systems: indexes slow you down, normalization costs more, and the bottleneck is usually I/O throughput rather than query complexity. I approach it with a different toolset than I use for read heavy work.

April 13, 2025 · 11 min read
Backend, APIs, and System Design

The Read Heavy Workload: Strategies That Move the Needle

A read heavy workload is any system where reads outnumber writes by a significant ratio, typically ten to one or higher. The strategies that help are not all equal. I separate the ones that move the needle from the ones that look good in a talk but cost more to run than the problem they solve.

April 12, 2025 · 12 min read
Backend, APIs, and System Design

The Edge: When to Move Logic Off Your Origin

The edge is the network of CDN nodes distributed globally where request processing can happen before traffic reaches the origin server. Moving logic to the edge reduces latency for global users, offloads work from the origin, and can enforce security rules closer to the requester. Not all logic belongs at the edge: operations that require database access, full Node.js APIs, or shared state must stay at or near the origin.

April 6, 2025 · 12 min read
Backend, APIs, and System Design

The Outbox Pattern: A SaaS Reliability Cheat Code

The outbox pattern solves the dual write problem by storing side effects in the same database transaction as the business event, then delivering them asynchronously from a separate relay process. I use it whenever a service needs to publish to a message broker or call a webhook without risking a partially applied state change that leaves data inconsistent.

March 29, 2025 · 12 min read
Backend, APIs, and System Design

The N+1 Query Problem: Detection, Prevention, and Refactoring

The N+1 query problem occurs when an application executes one query to retrieve a list of N records, then executes N additional queries to retrieve related data for each record, producing N+1 total queries where one query with a join would have sufficed. At small scale, the performance impact is invisible. At production scale, N+1 queries are one of the most common causes of slow API endpoints and database overload. Detection requires query logging; prevention requires joins or data loaders; refactoring requires identifying which ORM calls produce multiple queries and replacing them with eager loading.

March 24, 2025 · 12 min read
Backend, APIs, and System Design

Webhooks vs Polling vs Server Sent Events vs WebSockets

Webhooks, polling, server sent events, and WebSockets are four ways to move data between a server and a client when something changes. Each has a different cost, reliability profile, and implementation complexity. I pick between them based on who initiates the connection, how often the data changes, and whether the channel needs to carry data in both directions.

March 20, 2025 · 13 min read