SaaS Architecture and Scaling

Time Zones, Locales, and Currencies: The Three Horsemen of SaaS Apocalypse

Time zones, locales, and currencies are the three infrastructure concerns that look solved until your first international customer logs in. Each one has a correct pattern and a dozen incorrect patterns that compound over years. I've seen teams lose months to retrofitting these after the fact. The right design stores UTC, renders locally, stores amounts in minor units, delegates tax, and separates formatting from logic.

November 27, 2024 · 12 min read
SaaS Architecture and Scaling

The Customer Configuration Problem: How SaaS Companies Handle It Badly

Customer configuration is the set of settings, preferences, and customizations that make a SaaS product behave differently for each customer. The naive approach is a settings table with key/value pairs. The production ready approach is a structured configuration schema with validation, versioning, and a clear model for how configuration changes propagate. Most SaaS companies start with the naive approach and pay for it at scale.

November 26, 2024 · 12 min read
SaaS Architecture and Scaling

Trunk Based Development for SaaS Teams

Trunk based development is a source control practice where every engineer integrates to a single main branch at least once per day. Long lived feature branches are replaced by short lived branches of one to two days maximum, combined with feature flags to hide incomplete work. I've seen this practice cut integration incidents in half and reduce deployment anxiety on every team I've introduced it to.

November 25, 2024 · 11 min read
SaaS Architecture and Scaling

The User Impersonation Feature: Building It Securely

User impersonation lets an internal support or admin user act as a customer user inside the product, seeing exactly what the customer sees. It is a powerful support tool and a serious security surface. The secure implementation logs every impersonation session, requires explicit authorization, is visible to the customer in their audit trail, and cannot be used to modify billing or escalate permissions. I have built this feature multiple times and the security requirements are consistent.

November 23, 2024 · 12 min read
SaaS Architecture and Scaling

The Reporting Layer: A SaaS Story of Patience and OLAP

A reporting layer is the part of the system that answers questions about data in aggregate: how many users did X last month, what is the revenue trend by cohort, which customers are approaching a usage limit. It is separate from the transactional layer because the query shapes are different and the two layers have different scaling needs. Most SaaS products bolt reporting onto the OLTP database until it hurts, then migrate under pressure.

November 21, 2024 · 12 min read
SaaS Architecture and Scaling

The Search Problem: Why Adding It Late Always Hurts

Search in a SaaS product is the ability for users to find records across the product's data by entering natural language or structured queries. It is more complex than it looks because it requires a separate data model, an indexing pipeline, relevance tuning, and a query interface that maps what users type to what the system has stored. Teams that skip the design work ship search that disappoints and then spend months fixing it.

November 19, 2024 · 11 min read
SaaS Architecture and Scaling

The Notification System: A Bigger Project Than Founders Realize

A notification system is the layer that decides when to tell a user something, through which channel, and what to say. It sounds simple. In practice it involves preference management, delivery guarantees, channel routing, templating, suppression rules, and an audit trail. Most SaaS products build it in pieces across six different engineers over three years. I have seen what that looks like. It is not good.

November 19, 2024 · 12 min read
SaaS Architecture and Scaling

Workflow Engines: When You Need Temporal, When You Need Cron

A workflow engine runs multi step processes that have state, branching logic, and the need to survive failures and restarts. Cron schedules recurring single tasks. A job queue runs discrete background work. The three tools overlap in marketing but not in design. Choosing a workflow engine for a cron job is overengineering. Choosing a cron job for a multi step business process is a reliability disaster.

November 16, 2024 · 12 min read
SaaS Architecture and Scaling

The Operator Dashboard: A SaaS Founder's Forgotten Asset

The operator dashboard is the internal facing interface that lets your team view customer accounts, impersonate users for support, trigger manual actions, override system state, and monitor the health of the product in real time. It is not the same as the customer facing UI. It is a different surface built for your team. Most founders build it as a scattered collection of admin routes and database queries. I have seen what the good version looks like, and the gap is significant.

November 14, 2024 · 11 min read
SaaS Architecture and Scaling

Why Your SaaS Should Have a Job Queue From Day One

A job queue decouples work that does not belong in the request thread from the moment the user triggers it. Email, file processing, third party API calls, reports, and scheduled tasks all belong in a queue. Starting without one is a choice that costs you twice: first in reliability incidents, then in the refactor. Starting with one costs almost nothing on a modern stack.

November 12, 2024 · 12 min read
SaaS Architecture and Scaling

Webhooks: The Reliable Pattern That Most Companies Get Wrong

A webhook is an HTTP callback that a SaaS product sends to a customer endpoint when an event occurs. It is the primary mechanism for real time event notification in B2B integrations. The failure modes are well understood and consistently ignored: no retries, no signatures, no ordering guarantees, no dead letter handling. I have seen the same mistakes on products at every scale, from early stage to post-IPO.

November 11, 2024 · 12 min read
SaaS Architecture and Scaling

The Tenant Aware Permission System: A SaaS Engineer's Guide

A tenant aware permission system controls what actions a user can take within a specific organization, not just across the product. A user can be an admin in one tenant and a viewer in another. Permissions are scoped to the tenant context, and the enforcement layer always knows which tenant it is operating in. This is the design that B2B SaaS requires and the design that many teams skip in favor of a simpler global role model that does not scale past the first enterprise customer.

November 10, 2024 · 12 min read
SaaS Architecture and Scaling

The First Time a User Costs You Money: SaaS Unit Economics for Engineers

SaaS unit economics are the revenue and cost calculations at the level of a single customer: what do they pay, what does it cost to serve them, and what is the gross margin contribution from their subscription. Engineers rarely think in these terms, but every architectural decision affects the cost to serve. Understanding the unit economics from an engineering perspective allows architects to make decisions that improve the product's business viability, not just its technical correctness.

November 10, 2024 · 12 min read
SaaS Architecture and Scaling

The Database You Did Not Think You Needed: When to Add Redis, Elasticsearch, or ClickHouse

Postgres can do a lot. So can MySQL. But there are three specific scaling problems where a specialized database outperforms a relational one by an order of magnitude: caching and session storage (Redis), full text and faceted search (Elasticsearch), and analytical queries over large datasets (ClickHouse). The signal to add a specialized database is when your primary database queries for these use cases start affecting application performance.

November 3, 2024 · 12 min read
SaaS Architecture and Scaling

The Hidden Cost of Eventual Consistency: A SaaS Postmortem

Eventual consistency is the consistency model where writes to a distributed system are guaranteed to propagate to all nodes, but not immediately. In the window between a write and full propagation, different parts of the system may see different values. This model offers significant performance advantages over strong consistency (lower latency, higher availability) but introduces a category of bugs that are subtle, hard to test, and specifically expensive in SaaS products where customers expect their data to be accurate in real time.

November 1, 2024 · 12 min read
SaaS Architecture and Scaling

The Stateless API: Building Backends That Scale Horizontally

A stateless API holds no session or user specific data in server memory between requests. Each request carries everything the server needs to process it: a token, a tenant ID, whatever context is required. The server can be restarted, replaced, or multiplied without affecting sessions that are already in flight. This is the prerequisite for horizontal scaling, blue and green deploys, and restarts with zero downtime.

October 28, 2024 · 11 min read
SaaS Architecture and Scaling

The Modular Monolith: How to Buy Yourself Two Years

The modular monolith is a single deployable application that is internally organized into well defined modules with explicit interfaces between them. Each module owns its data and business logic; no module accesses another module's database tables directly. This structure gives teams the deployment simplicity of a monolith while building the internal boundaries that make future extraction to microservices or separate services tractable. The modular monolith is the right architecture for most products until they reach genuine scale, which is 10 to 50 times more users than most teams expect.

October 23, 2024 · 12 min read
Hiring Developers, Freelancers, and Agencies

The Hiring Funnel for Engineering Roles: From Applied to Signed

The engineering hiring funnel is the sequence of stages from initial application to signed offer: application review, recruiter screen, technical phone screen, technical assessment, team interviews, reference checks, and offer. The conversion rates at each stage determine how many candidates must enter the top of the funnel to produce one signed hire. Understanding the funnel allows companies to identify the stages where strong candidates are being lost unnecessarily and to design each stage to be both evaluative and compelling to the candidates who pass.

October 20, 2024 · 12 min read
Hiring Developers, Freelancers, and Agencies

Why Most Tech Recruiting Firms Send You The Wrong People

Most tech recruiting firms send you the wrong people because their incentive is to close the placement, not to find the right match. They are optimized for speed and volume, not for the kind of slow, specific evaluation that a technical hire actually requires. The fee structure rewards the close, not the outcome, and the candidate pool is whoever responded to an email, not whoever would be best for the role.

October 18, 2024 · 11 min read
Hiring Developers, Freelancers, and Agencies

The Recruiter Engineer Relationship: How to Make It Work

The recruiter engineer relationship works when the recruiter understands what the engineer actually does and the engineer understands what the recruiter actually needs. In my experience watching this relationship from the technical side, the failures almost always come from the recruiter treating the engineer as a keyword match rather than a person with a specific set of preferences. The ones that work feel like a conversation between two people who respect each other's expertise.

October 17, 2024 · 12 min read
Hiring Developers, Freelancers, and Agencies

The Annual Engineering Review That Engineers Actually Find Useful

The annual engineering review is the most important management conversation of the year and the one most commonly wasted. Most engineering reviews are filled with competency ratings, generic feedback, and calibration theater. The reviews that engineers find useful are specific, honest, and focused on trajectory. The format matters far less than the honesty.

October 14, 2024 · 12 min read
Hiring Developers, Freelancers, and Agencies

The Senior Developer Test: Architecture, Tradeoffs, Communication

When I need to know whether a developer is actually senior, I do not look at the resume. I ask three questions that cannot be rehearsed: one about a system they designed under constraints, one about a decision they regret, and one about how they explain something hard to someone who does not code. The answers tell me everything the LinkedIn profile will not.

October 10, 2024 · 12 min read
Hiring Developers, Freelancers, and Agencies

Why I Stopped Asking Whiteboard Coding Questions

I stopped asking whiteboard coding questions because they answered the wrong question. I was testing whether a developer could perform a kind of algorithmic recall under social pressure, when what I actually needed to know was whether they could build, communicate, and make judgment calls on a real product. The trial project replaced the whiteboard and the hiring quality improved immediately.

October 9, 2024 · 11 min read
Hiring Developers, Freelancers, and Agencies

The First Thirty Days: Setting a New Developer Up for Success

The first thirty days with a new developer determine whether the engagement will be productive or frustrating. The developers who become valuable contributors quickly are almost always in environments where expectations are clear, context is provided proactively, and feedback is given early. The developers who take months to become productive are almost always in environments where they are expected to figure things out on their own. The difference is in the setup, not in the developer.

October 8, 2024 · 12 min read