Why cloud bills drift
Cloud pricing is pay-as-you-go, which is its biggest strength and its most common trap. Nobody signs a purchase order when an engineer launches a large test server, leaves a database running over a long weekend, or turns on verbose logging. Costs creep up a little at a time until the monthly invoice gets attention.
Controlling that drift is less about a single tool and more about habits. The FinOps Foundation defines FinOps as an operational framework and cultural practice for getting the most business value from technology, built on collaboration between engineering, finance and business teams. Its six principles include that teams work together, that everyone takes ownership of their own technology usage, that cost data should be accessible, timely and accurate, and that organisations should take advantage of the cloud’s variable cost model. The work runs through three repeating phases:
- Inform: see and allocate what you are spending, forecast it, and report on it.
- Optimise: find and prioritise ways to improve efficiency, from shrinking oversized resources to changing contracts.
- Operate: put the chosen changes into practice, then loop back.
AWS’s Well-Architected cost optimisation pillar covers similar ground: practise cloud financial management, build expenditure and usage awareness, use cost-effective resources, match supply to demand, and optimise over time. The practical controls below follow that order.
1. Make costs visible with tags
You cannot manage what you cannot attribute. Agree a small, mandatory set of tags (or labels) for every resource, for example environment, application, owner and cost-centre, and enforce them in your infrastructure-as-code templates.
On AWS, tags only show up in cost reports once they are activated as cost allocation tags in the Billing console. User-defined and AWS-generated tags must each be activated separately, and newly activated tags can take up to 24 hours to appear. AWS also advises against putting sensitive information in tags.
Once tagging is in place, reports show what each application and environment costs. That is the precondition for giving teams responsibility for their own spending.
2. Set budgets and alerts, and know what they do not do
All three major providers let you set budgets with alerts on both actual and forecasted spending. The key fact is easy to miss: a budget alert does not stop spending by itself.
- AWS: AWS Budgets supports cost, usage, and reservation or Savings Plans utilisation and coverage budgets, with alerts by email or Amazon SNS. Budget data is updated up to three times a day, and AWS warns that charges can exceed a threshold before you are notified. Budget actions can respond automatically, for example by applying an IAM policy that prevents new resources being provisioned.
- Azure: in Microsoft Cost Management, crossing a threshold triggers notifications, but resources are not affected and consumption is not stopped. Cost data is typically available within 8 to 24 hours, and budgets are evaluated every 24 hours. For subscription and resource-group budgets, you can call an action group to run automation.
- Google Cloud: an alerts-only budget does not cap usage. The default thresholds are 50%, 90% and 100% of the budget. Connecting a Pub/Sub topic gives programmatic notifications several times a day, which you can use to trigger automated responses. Because cost data arrives with a delay, Google recommends setting budgets below the funds you actually have available.
Practical set-up: one budget per environment or application, a forecast alert well before 100%, an actual-spend alert at 100%, and alerts sent to a channel the owning team actually reads. Automate hard responses only for environments where stopping things is safe, such as sandboxes.
3. Rightsize before you do anything else
Over-provisioning is the most common waste: servers sized for a peak that never comes, or for a guess made at launch. The provider tools analyse real usage and suggest smaller, cheaper configurations or idle resources to remove.
AWS Compute Optimizer, for example, must be opted into. It then analyses 14 days of CloudWatch metrics by default, extendable to 93 days with a paid enhanced-metrics option. It covers EC2 instances, Auto Scaling groups, EBS volumes, Lambda functions, RDS and Aurora databases, Fargate services, NAT gateways and more. Azure and Google Cloud have their own recommendation services; check what each covers in your account.
Treat recommendations as a starting point. Check memory as well as CPU, test the smaller size in a non-production environment, and change one thing at a time. Rightsize before buying commitments, so you do not lock in a discount on capacity you never needed.
4. Scale with demand
Fixed capacity sized for peak load means paying for idle machines most of the day. Autoscaling adds capacity when load rises and removes it when load falls. With Amazon EC2 Auto Scaling, you set a minimum, a maximum and a desired number of instances, and scaling policies adjust within those limits. There is no extra fee for Auto Scaling itself; you pay for the resources it runs. The other major providers offer comparable autoscaling for groups of virtual machines.
For spiky or low-traffic workloads, also consider whether a managed or serverless service that scales down when idle would fit better than always-on servers.
5. Switch off non-production when nobody is using it
Development, test and demo environments often run around the clock but are used only during working hours. Scheduling them to stop overnight and at weekends is one of the simplest savings available.
Azure offers auto-shutdown for virtual machines at a set time, with optional notification beforehand; note that the default time zone is UTC. Be careful about what “stopped” means. Azure’s billing states show that a VM that is Stopped but still allocated to a host continues to be billed for compute. Only a Deallocated VM stops compute charges, and disks and some networking resources continue to cost money either way.
The same thinking applies everywhere: stopping a server rarely removes its storage costs, so delete unattached disks, old snapshots and unused IP addresses as well.
6. Commit to steady usage, and only steady usage
Once usage is rightsized and predictable, commitment discounts trade flexibility for a lower rate:
| Provider | Option | Terms (from official docs) |
|---|---|---|
| AWS | Savings Plans | Commit to an hourly spend for 1 or 3 years; savings of up to 72% on compute; all, partial or no upfront payment |
| Azure | Savings plans and reservations | 1- or 3-year hourly commitment; usage above it is billed at pay-as-you-go rates; unused commitment in an hour does not roll over; savings plans cannot be cancelled or refunded |
| Google Cloud | Committed use discounts | Resource-based: 1 or 3 years, up to 55% for most machine series (up to 70% for memory-optimised), cannot be cancelled. Spend-based “flexible” commitments: for example 28% (1 year) or 46% (3 years) on general-purpose compute |
Two rules keep commitments safe. First, cover only your stable baseline, the usage you are confident will still be there at the end of the term, and leave peaks on demand. Second, monitor utilisation: AWS Budgets can alert you when Savings Plans or reservation utilisation drops below a target, which signals you over-committed.
Azure notes that reservations give deeper discounts for a specific resource, size and region, while savings plans apply more flexibly across services and regions. Choose based on how likely your architecture is to change.
7. Use spare capacity for interruptible work
EC2 Spot Instances use spare capacity at prices below On-Demand, but AWS can reclaim them with a two-minute interruption notice. They suit batch jobs, data analysis, CI builds and other work that can be retried. Spot usage is not covered by Savings Plans, so the two discounts do not stack. Check your provider’s equivalent before relying on it for production capacity.
8. Put storage on the right tier, and let it expire
Storage costs grow quietly because data is rarely deleted. Most providers offer cheaper tiers for data that is accessed less often, and rules to move or delete data automatically.
With Amazon S3 Lifecycle, transition rules move objects to cheaper classes after a set age (for example to Standard-IA after 30 days, or to Glacier Flexible Retrieval after a year), and expiration rules delete them. Rules apply to existing objects as well as new ones. Watch the fine print: transitions incur per-request charges, and expiring objects from a storage class that has a minimum storage duration can incur a charge. Model the request and minimum-duration costs before moving large numbers of small objects.
Also review backup and snapshot retention, log retention periods, and old container images.
9. Design to reduce data transfer (egress)
Data transfer is the cost most often left out of estimates. AWS’s guidance on data transfer modelling notes that charges depend on the source, destination and volume of traffic. The companion guidance on choosing components suggests:
- Avoid moving data between regions unless there is a clear reason; inter-region transfer typically incurs charges.
- Use VPC endpoints to reach provider services privately rather than over the public internet.
- Keep heavy-traffic resources in the same availability zone as the NAT gateway they use, to avoid cross-zone charges.
- Use a content delivery network to serve content closer to users.
- Monitor network flows with tools such as VPC Flow Logs and cost reports, so surprises are caught early.
Multi-zone and multi-region designs are often worth their transfer costs for resilience. The point is to make that a conscious decision, priced in advance.
Making it stick
One-off clean-ups decay. What lasts is a routine:
- Weekly: each team reviews its own cost dashboard and any budget alerts.
- Monthly: review rightsizing and idle-resource recommendations; check commitment utilisation.
- Quarterly: revisit architecture choices (data transfer, storage tiers, managed versus self-run services) and plan commitment renewals.
- Always: enforce tags and schedules in code, so new resources start out compliant.
Sources
- FinOps Foundation: FinOps Framework
- FinOps Foundation: FinOps Phases
- AWS Well-Architected Framework: Cost Optimization Pillar
- AWS Billing: Organizing and tracking costs using AWS cost allocation tags
- AWS Cost Management: Managing your costs with AWS Budgets
- Microsoft Learn: Create and manage budgets (Cost Management)
- Google Cloud: Create, edit, or delete budgets and budget alerts
- AWS: What is AWS Compute Optimizer?
- AWS: What is Amazon EC2 Auto Scaling?
- Microsoft Learn: Auto-shutdown a VM
- Microsoft Learn: States and billing status of Azure Virtual Machines
- AWS: What are Savings Plans?
- Microsoft Learn: What are savings plans?
- Google Cloud: Committed use discounts overview
- AWS: Amazon EC2 Spot Instances
- AWS: Managing the lifecycle of objects (Amazon S3)
- AWS Well-Architected: COST08-BP01 Perform data transfer modeling
- AWS Well-Architected: COST08-BP02 Select components to optimize data transfer cost