Scalability advice tends to arrive in two unhelpful flavours. One says to build for a million users immediately, which wastes a year you may not have. The other says not to worry until you have a problem, which is fine until the problem is your data model.
The useful frame is reversibility. Some decisions are cheap to change later; some are extremely expensive. Spend your early effort on the second category and stay deliberately casual about the first.
Expensive to reverse
- Your data model and the identifiers other systems will hold references to.
- Tenancy and authorisation boundaries: retrofitting multi-tenancy is a rewrite in disguise.
- Whether operations are idempotent, particularly anything touching money.
- Public API contracts, once a third party depends on them.
- Audit and event history: data you did not record at the time cannot be reconstructed later.
Cheap to reverse
- Which cloud region or instance type you run on.
- Whether a workload is a service or a module inside the monolith, provided the boundaries are clean.
- Caching strategy, queue technology, and most of your infrastructure choices.
- Anything behind a well-defined interface with tests around it.
Sorted this way, the work becomes tractable. Spend a week on the data model and the authorisation boundaries. Do not spend a month choosing a message broker.
Design the seams carefully. Behind a good seam, almost any decision can be replaced.
Scale is usually operational, not computational
Most systems that fall over under growth do not run out of CPU. They run out of human capacity to operate them: nobody can tell which service caused the latency, deploys have become risky, and the only person who understands the migration path is on leave.
Which is why observability, infrastructure as code, and a deployment pipeline are scalability features. They are the difference between a system that can be grown by a team and one that can only be grown by its original authors.
The load test that matters
Before a seasonal peak, rehearse it. Not a synthetic benchmark, but the actual traffic shape at a multiple of your worst hour, against production-scale data, with the team watching the dashboards they will use on the day. On the Northwind replatform we tested at fifteen times the previous peak and found two issues that would each have been an outage in November.
Scale, in practice, is far less about clever engineering than about deciding early which mistakes you are prepared to live with, and then rehearsing the day you find out.
About the author
Bonolo Mathabela
Bonolo co-founded Bonang Technologies to prove a specific point: that a small, senior team with clear standards ships better software than a large one with none. As COO they own how the studio runs: how work is estimated, how it is staffed and reviewed, and what is allowed to reach production.