When we start a cloud engagement we ask for two things before any architecture diagram: twelve months of billing data, and access to the monitoring. The bill is usually more honest about how the system is designed than the documentation is.
Idle staging environments running at production size. Data transfer charges revealing that two services which talk constantly are in different regions. Storage growing linearly with no lifecycle policy, meaning nothing is ever deleted because nobody decided what should be. Each line is a design decision that was made once, probably in a hurry, and never revisited.
The savings are rarely heroic
There is a persistent assumption that cutting cloud cost requires re-architecture. In our experience the first forty percent comes from unglamorous work:
- Right-sizing workloads that were provisioned for a launch-day estimate nobody revisited.
- Scheduling non-production environments to sleep outside working hours.
- Storage lifecycle rules moving cold objects to cheaper tiers automatically.
- Deleting orphaned volumes, unattached addresses, and the environments of projects that ended.
- Committing to reserved capacity for the baseline that has clearly not moved in a year.
On the Sentinel migration this class of work accounted for a 41% reduction in run cost. None of it involved rewriting an application.
Attribution before optimisation
You cannot manage what you cannot attribute. Before optimising anything, every resource should carry tags for team, environment, and service, and the bill should be breakable down along those lines. Once an engineering lead can see what their service costs per month, behaviour changes without a mandate from finance.
Cost is a non-functional requirement. Treat it like latency: measured, budgeted, and alerted on.
Design for the shape of your load
The deeper savings come from matching architecture to actual traffic shape. A retail platform with an eleven-times seasonal peak wants different economics from an internal tool with a flat weekday profile. Static generation at the edge for content that changes daily; autoscaling for the genuinely variable parts; reserved baseline for the floor that never moves.
This is why we treat cost as an architectural input rather than an operations problem discovered later. The cheapest system is usually also the simplest one, and simplicity is not something you can retrofit into a design that assumed money was free.
A quarterly ritual worth keeping
Once a quarter, put the bill and the reliability report on the same page and review them together. Cost and resilience trade against each other constantly, and reviewing either in isolation produces bad decisions. Teams that do this stop being surprised by invoices, which turns out to be most of the benefit.
About the author
Angel Maile
Angel co-founded Bonang Technologies after a decade spent building software inside organisations where the technology decisions and the commercial ones were made in separate rooms. The company exists to close that gap: engineering that starts from what the business is actually trying to achieve, and leadership that can hold both conversations at once.