Backend, APIs, and System Design

Designing an API That Customers Will Not Curse In Five Years

An API that customers will not curse in five years is the result of decisions made at week one. Stable shapes. Additive evolution. Consistent naming. Clear errors. Cursor based pagination. Versioning that respects existing integrations. The discipline is unglamorous. The result is an API that customers integrate once and keep using rather than tolerating.

May 21, 2026 · 12 min read
Backend, APIs, and System Design

Database Partitioning: Strategies and Pitfalls

Database partitioning is the technique of splitting one logical table into multiple physical pieces. The pieces can be queried as one but are stored and indexed separately. The pattern fits very large tables where queries naturally filter by the partition key. The pattern hurts tables where the queries do not naturally filter. The right key and the right time are the two decisions that matter most.

May 21, 2026 · 12 min read
Backend, APIs, and System Design

Database Indexes: A Practical Primer for SaaS Engineers

A database index is a data structure that speeds up query lookups. The right index turns a query from seconds to milliseconds. The wrong index slows writes and wastes storage without helping reads. Most SaaS performance problems are missing indexes. The fix is mechanical. The discipline is to recognize which indexes to add and which to skip. The primer is small. The impact is large.

May 19, 2026 · 12 min read
Backend, APIs, and System Design

CRDTs in Production: A Real World Look

CRDTs, Conflict-Free Replicated Data Types, are data structures designed to support concurrent edits across multiple replicas without coordination. They allow each replica to update independently and converge to the same state. CRDTs are the right tool for real time collaboration, offline first apps, and specific eventual consistency scenarios. They are the wrong tool for almost everything else.

May 19, 2026 · 12 min read
Backend, APIs, and System Design

CQRS in Practice: When the Complexity Earns Its Keep

CQRS, Command Query Responsibility Segregation, is the architectural pattern of separating the write side and the read side of a system into different models. The pattern adds real complexity. It pays back when reads and writes have dramatically different patterns or scales. For most SaaS the simpler single model approach is the right call. For specific cases CQRS is the only sane option.

May 19, 2026 · 11 min read
Backend, APIs, and System Design

CockroachDB and TiDB: When Distributed SQL Pays Off

Distributed SQL is the category of databases that provide the familiar SQL interface with horizontal scaling and multi region replication. CockroachDB and TiDB are the leaders. They solve the problem of scaling a single database past what Postgres or MySQL can do on a single node. The trade off is operational complexity that most SaaS does not need. The right call depends on whether your scale or geography demands it.

May 19, 2026 · 12 min read
Backend, APIs, and System Design

Choosing a Cache: Redis vs Memcached vs In Memory

In memory cache lives inside your application process. It is the fastest and simplest cache but does not share across processes. Memcached is a distributed key value cache with simple semantics and extreme performance. Redis is a distributed cache and data structure server with richer features. Most modern SaaS picks Redis for the breadth. Memcached and in memory each win in specific cases that justify their constraints.

May 19, 2026 · 12 min read
Backend, APIs, and System Design

CDN Strategy for a Global SaaS in 2026

A CDN strategy in 2026 covers static assets, public API responses, dynamic page caching, edge functions, image optimization, DDoS protection, and global routing. The right CDN reduces origin load, improves user perceived performance, and gives you a defensive layer in front of the application. The wrong CDN is just a passthrough that adds latency and cost.

May 19, 2026 · 12 min read
Backend, APIs, and System Design

Building Internal APIs vs Public APIs: Different Disciplines

An internal API connects services inside your own system. The consumers are your own engineers. The contract can change quickly. The optimization is for development velocity and operational efficiency. A public API is consumed by customers and partners. The contract is binding for years. The optimization is for stability and ergonomics. Treating one like the other produces brittle internal coupling or unstable customer integrations. The disciplines are different.

May 19, 2026 · 12 min read
Backend, APIs, and System Design

Building APIs That Survive Five Years of Customer Change Requests

An API that survives five years of customer change requests has stable shapes, additive evolution, deliberate versioning, and a clear contract between you and your customers. The discipline starts at design time and never stops. The teams that respect it ship APIs that customers integrate once and keep using. The teams that do not respect it ship APIs that customers integrate and complain about for the entire relationship.

May 19, 2026 · 12 min read
Backend, APIs, and System Design

Batch Processing vs Streaming: The Choice Founders Often Conflate

Batch processing runs over a bounded window of data on a schedule. Streaming processes events continuously as they arrive. Batch is simpler, cheaper, and right for almost every SaaS analytics workload. Streaming is right when the latency of the output matters in seconds rather than minutes, or when the data has no natural batch boundary. The wrong choice often doubles the operational cost and provides no business benefit.

May 18, 2026 · 12 min read
Backend, APIs, and System Design

Background Jobs at Scale: Inngest, Trigger, Cron, and Beyond

Inngest and Trigger.dev are managed durable execution platforms designed for serverless and Next.js stacks. They handle retries, scheduling, fan out, and step level state for multi step jobs. Cron is still the right tool for simple scheduled work. The combination of a managed durable executor for complex flows and a queue for simple background work covers most modern SaaS workloads cleanly.

May 18, 2026 · 13 min read
Backend, APIs, and System Design

Async Job Failure Recovery: Patterns That Actually Work

Async jobs fail. Patterns that recover them well share five qualities. Idempotency, exponential backoff with jitter, dead letter queues with alerting, structured retries, and reconciliation jobs. Together they handle the vast majority of failure modes a SaaS will see. The teams that ship them ride out incidents that would sink the teams that skipped them.

May 17, 2026 · 11 min read
Backend, APIs, and System Design

API Gateway Patterns for SaaS: Kong, Tyk, AWS API Gateway Compared

An API gateway sits in front of your services and handles auth, rate limiting, routing, and observability before the request reaches your code. Kong and Tyk are the open source leaders. AWS API Gateway is the managed default for teams already on AWS. Each has a clear best fit. The teams that pick wrong end up either paying for features they do not use or self hosting a service they do not have capacity to operate.

May 17, 2026 · 11 min read
Backend, APIs, and System Design

API Documentation That Developers Actually Read

API documentation that developers read has three layers. A quick start that lets them make their first call in five minutes. A reference that covers every endpoint with copy paste examples. A cookbook that shows how to combine endpoints to do the common things. The teams that ship all three keep developers engaged. The teams that ship only the reference watch developers drift to the support inbox.

May 17, 2026 · 11 min read
Backend, APIs, and System Design

ACID vs BASE: When Each Belongs in Your Architecture

ACID is the guarantee that a database transaction either fully succeeds or fully fails. BASE is the guarantee that a system remains available even when individual nodes disagree for a short time. Most production systems run both, in different layers. Picking the wrong one for the wrong workload costs you either data integrity or uptime.

May 17, 2026 · 12 min read
Backend, APIs, and System Design

Why Your Service Should Have Two Health Checks Not One

A service needs two health checks because liveness and readiness answer different questions. Liveness asks should this process be killed and restarted. Readiness asks should this process receive new traffic. A single endpoint that mixes the two will cause unnecessary restarts during transient dependency failures and is one of the most common quiet sources of self inflicted outages.

May 4, 2025 · 12 min read
Backend, APIs, and System Design

The Health Check Endpoint: Less Trivial Than It Looks

A health check endpoint is an API route that load balancers, orchestrators, and monitoring systems call to determine whether a service instance is ready to receive traffic. The trivial implementation, returning 200 OK immediately, is nearly useless for detecting real service failures. The correct implementation tests the dependencies that the service needs to function (database connectivity, cache availability, critical external APIs) and returns a structured response that distinguishes between healthy, degraded, and unhealthy states.

May 3, 2025 · 12 min read
Backend, APIs, and System Design

Timeouts: The Setting Most Engineers Get Wrong

A timeout is the maximum time your code is willing to wait before giving up and doing something else. Most engineers set them once, forget them, and discover they were wrong during an incident. The correct value depends on the operation, the SLA of the dependency, and what happens to the user or the data when the timeout fires. Getting it right requires thinking about all three, not just copying a number from a tutorial.

May 2, 2025 · 11 min read
Backend, APIs, and System Design

The Backend Engineer's Reading List for 2026

The backend engineer's reading list is shorter than most people think. Five to eight books cover the foundational ideas that show up in every production system. The rest is reading documentation and learning from production incidents. This list is the five to eight, with notes on when each one is most useful and which chapter to start with.

April 30, 2025 · 12 min read
Backend, APIs, and System Design

The Boring API: Why Predictability Beats Cleverness

A boring API is one where every developer who touches it can predict the behavior of any endpoint they have not seen before, based on what they learned from the first three. Consistent naming, consistent error responses, consistent pagination, consistent authentication. The boring API is not a technical achievement. It is a communication achievement. And it is the one that survives five years of customer integrations.

April 29, 2025 · 12 min read
Backend, APIs, and System Design

Why Logical Deletes Are Almost Always a Mistake

A logical delete, also called a soft delete, is a pattern where rows are never removed from the database. Instead, a deleted_at timestamp or is_deleted flag marks them as gone. The data stays. The complexity grows. Most teams adopt this pattern early and regret it late, when every query needs a WHERE deleted_at IS NULL and migrations become a negotiation with historical junk.

April 26, 2025 · 12 min read
Backend, APIs, and System Design

The Soft Delete Trap: A Pattern That Catches Up With Teams

Soft deletes mark a record as deleted with a flag or timestamp instead of removing it from the database. The pattern is popular because it feels safe and recoverable. In my experience it is safe at first and expensive later, as the deleted rows accumulate, queries get slower, and the team discovers that every feature touching that table needs to filter out the deleted rows or it shows deleted data.

April 25, 2025 · 11 min read
Backend, APIs, and System Design

Time Series Data in SaaS: When to Pull in TimescaleDB or InfluxDB

Time series data is data where the timestamp is the primary key and queries are almost always based on ranges: give me the values between now and thirty days ago. Postgres handles it adequately at low volumes. TimescaleDB extends Postgres with automatic partitioning and compression for time ordered data. InfluxDB is purpose built for metrics at high ingestion rates. Knowing which to reach for saves weeks of wasted infrastructure work.

April 22, 2025 · 12 min read