Performance Optimization

API Response Times: How to Track What Matters

API response time discipline is built on three numbers per endpoint. The p50, the p95, and the p99. The p50 tells you what most users experience. The p95 tells you what the unlucky users experience. The p99 tells you what the worst customer event looks like. Averages lie. Percentiles do not. The teams that track all three by endpoint and by customer catch problems before users complain.

May 17, 2026 · 11 min read
Performance Optimization

The Three Hour Performance Audit Every Team Should Run Quarterly

A three hour performance audit is a structured quarterly review covering database query health, API endpoint percentiles, front end bundle size, and caching layer effectiveness. I run these with client teams as a repeating calendar item. The goal is not perfection. It is catching the regressions that compound quietly over a quarter and turning each one into a tracked task before a customer notices.

July 16, 2025 · 12 min read
Performance Optimization

The Performance Regression That Hides in CI

A CI performance regression is a latency or throughput degradation introduced by a code change that passes functional tests but is never caught because the pipeline has no budget check wired to the build. It ships quietly, accumulates over weeks, and surfaces only when a customer complains or a metric dashboard finally gets reviewed. The fix is a budget gate, not more manual review.

July 14, 2025 · 12 min read
Performance Optimization

Why Memoization Is a Trap When Misused

Memoization caches the result of a function so future calls return instantly. Used carefully on expensive, deterministic computations, it is a real win. Used as a reflex on cheap functions or as a fix for rerendering issues, it adds complexity, increases memory, hides the real problem, and produces stale data bugs that are nearly impossible to reproduce in development.

July 11, 2025 · 12 min read
Performance Optimization

The Hot Path: Finding and Optimizing It

The hot path is the code that executes on every request or on the most performance critical requests in a system. It is the code where a 1ms improvement has the largest impact on the overall system latency, and where a 5ms regression is immediately visible in p99 latency metrics. Finding the hot path requires profiling, not intuition: experienced engineers are wrong about which code is slow as often as they are right. Optimizing the hot path requires removing unnecessary work, deferring work to background tasks, and caching the results of expensive operations.

July 9, 2025 · 12 min read
Performance Optimization

The Slow Query Log: A Discipline Every SaaS Team Should Practice

The slow query log is a database level feature that records every query exceeding a configured threshold, typically 100 to 500 milliseconds. I treat it as a standing discipline, not a firefighting tool. Reviewed weekly, it surfaces the queries that will degrade under load before users see the effect. Teams that skip this step discover the same queries in production under pressure.

July 4, 2025 · 12 min read
Performance Optimization

The Garbage Collection Tax: A Backend Story

Garbage collection tax is the latency added to request processing when the runtime's garbage collector pauses execution to reclaim memory. In Node.js and JVM based services, GC pauses are the most common source of latency spikes that are not explained by slow database queries or external API calls. The tax is paid on every GC cycle but is invisible in average latency metrics. It shows up as p99 and p999 latency outliers that are much higher than p50.

July 1, 2025 · 12 min read
Performance Optimization

The HTTP Caching Strategy That Most Teams Get Wrong

HTTP caching is the mechanism by which browsers and CDNs store responses and serve them without contacting the origin server. The correct caching strategy depends on whether the content changes frequently, whether it is user specific, and what the acceptable staleness window is. Most teams either cache too aggressively (serving stale user specific data) or too conservatively (disabling caching for static assets that could safely be cached for months). The difference between a correct and incorrect caching strategy is measurable in page load time, infrastructure costs, and the correctness of what users see.

June 27, 2025 · 12 min read
Performance Optimization

The Edge Rendering Bet: When It Pays Off

Edge rendering executes server side rendering logic at CDN edge nodes distributed globally, rather than at a single origin server. The promise is lower latency for geographically distributed users. The reality depends on the workload: pages that are mostly static or can be cached benefit significantly from edge rendering. Pages that require database queries, authenticated sessions, or dynamic content may not benefit, and may perform worse due to the limitations of the edge runtime.

June 23, 2025 · 12 min read
Performance Optimization

The Real Numbers Behind a Fast Web App in 2026

The real numbers behind a fast web app are not goals you pick from a blog post. They are the thresholds at which users stop noticing load time, at which search engines reward you, and at which churn starts to drop. I track six metrics per app: LCP, INP, CLS, API p95 latency, time to first byte, and bundle size. Every one has a concrete target and a measurement method.

June 16, 2025 · 12 min read
Performance Optimization

Why Your App Got Slower After You Added Users

An app slows down as it grows because the work per request scales with the size of the data, not the number of features. Queries that scanned ten rows now scan ten million. The fix is rarely a rewrite. It is finding the three or four queries and code paths that grew with the data and fixing those.

June 15, 2025 · 12 min read