Request a callbackBook a call
← All posts

How I Cut a $200K/Year Cloud Bill by More Than 70%: A Solo GCP to Azure and AWS Migration

TL;DR
  • AccioJob's cloud run rate, projected at more than $200K a year, fell to under $60K, a cut of more than 70%, after a GCP to Azure and AWS migration I ran solo, with zero downtime and no external DevOps support.
  • The order did more than any single change: credits first ($200K from Microsoft for Startups, $100K from AWS Activate), then workload placement, then per-request managed services moved to reserved capacity, then storage tiering and model routing, then the DevOps queue retired.
  • The 70% comes from architecture and survives credit expiry. The $300K of credits paid for the destination clouds while both ran, and I do not count them as savings.
The bill, annualised
Before, on GCP
Compute and per-request managed services
about $120K a year
Databases and storage
about $40K a year
AI and LLM inference
about $40K a year
Deploys, infrastructure changes, incidents
queued behind one DevOps role
Projected run rate
more than $200K a year
After, on Azure and AWS
Steady workloads
self-hosted on reserved capacity
Databases and storage
right-sized and tiered by access
Inference
routed to the cheapest model that does the job
Deploys, infrastructure changes, incidents
owned by cross-trained engineers
Run rate
under $60K a year
More than 70% off the annual run rate at list price, with $300K of credits kept out of the calculation
The before split is approximate. The after total is the published result; I have not split it by line, because the migration changed what the lines were.

How do you cut a $200K/year cloud bill by more than 70%?

By changing what you pay for, in a fixed order: secure credits, place each workload where its pricing is best, move steady workloads off per-request billing onto reserved capacity, tier storage, route model calls, and retire the DevOps queue. That took AccioJob's run rate from more than $200K a year to under $60K.

I did it at AccioJob, where I am Head of Engineering, and I did it alone: a full migration from GCP to Azure and AWS with no external DevOps support and zero downtime. The stack on the other side is ordinary: Terraform, Docker, Kubernetes, PostgreSQL and Redis. Nothing below depends on a tool you cannot buy or download today.

Most teams run the levers in reverse. They start with instance right-sizing, which is per-resource work with per-resource payoff, spend a quarter on it, and conclude that cloud cost is the price of doing business. Right-sizing came after the architecture work here, and it was the smallest of the changes.

Treat the numbers as one estate's, not a benchmark. The split of the bill, the credit amounts and the timeline will differ for you. What transfers is the order of the changes, and the rule that every change is measured at list price before anyone calls it a saving.

$200K+
projected annual run rate on GCP before the migration
under $60K
annual run rate afterwards on Azure and AWS, at list price
$300K
cloud credits: $200K Microsoft for Startups, $100K AWS Activate
0
external DevOps engineers needed after the migration
Four published numbers from one migration. The first two are run rates, the third paid for the move, and the fourth is why the saving held.

Where was the $200K a year going?

Into four places: compute and managed services billed per request (about $120K a year), over-provisioned databases and storage not tiered by access (about $40K), AI inference whose cost grew in a straight line with usage (about $40K), and a DevOps function that every deploy and incident queued behind. The split is approximate; the projected total was more than $200K.

Per-request pricing was the biggest line and the most avoidable. The infrastructure had grown organically on GCP, and managed services that bill per request are the right choice when you do not yet know your load. By the time I looked, much of that load was steady: the same services, busy all day, paying an elasticity premium for elasticity they never used.

The inference line was the fastest-growing. Its cost scaled in a straight line with usage, and nothing routed easy requests to a cheaper model. The databases were over-provisioned for the load they carried. And the DevOps role was a cost twice over: a salary, and a queue that slowed every change that touched infrastructure.

Before any of this was fixable it had to be visible. I mapped three months of billing data line by line before changing anything. That step produces no saving on its own, which is why most teams skip it, and every later decision depended on it.

Line, beforeApproximate share of the run rateWhat was wrong with itWhat replaced it
Compute and per-request managed servicesabout $120K a yearSteady load billed per requestSelf-hosted equivalents on reserved capacity, placed by price and credits
Databases and storageabout $40K a yearDatabases over-provisioned; storage not tiered by accessRight-sized databases and aggressive storage tiering
AI and LLM inferenceabout $40K a yearCost linear with usage, no routing between modelsModel routing: a cheap model by default, a larger one when needed
DevOps functionone full-time roleEvery deploy and incident queued behind one personEngineers cross-trained to own deploys and infrastructure

Why secure cloud credits before changing any code?

Because credits pay for the destination while both clouds are running, and that overlap is usually the most expensive part of a migration. With both applications in before the first workload moved, $200K from Microsoft for Startups and $100K from AWS Activate, no deadline forced a rushed cutover.

The applications were pitches, not forms. What carried them was high-volume production AI usage already running, named workloads with their current volumes, and a specific account each provider would keep after the credits ran out. Providers award large credits to workloads they expect to keep. An application that shows real traffic and a migration date gives them something to underwrite; a projection deck does not.

The published programmes have moved since. Microsoft now advertises up to $150,000 on its Microsoft for Startups page, less than the $200K I received. AWS Activate lists up to $200,000 on its Portfolio tier, which needs an Organization ID from an Activate Provider, for pre-Series B companies founded in the last ten years. The full walkthrough is in how to get cloud credits.

One rule kept the credits honest: I did not count them as savings. The 70% is the list-price run rate after the re-architecture. The credits covered bills during the move and bought time to do it properly, and a cost model that only works while someone else pays for compute has a cliff in it.

The credit applications
What went into the applications
  • Production usage graphs, not projectionsVolume in the provider's own units
  • Named AI workloads with their monthly inference volume
  • The specific workloads that would move, and whenA migration with a date beats a general subsidy
  • A projection of the account after the credits run outProviders underwrite the account, not the credit period
  • A partner route where you have oneAWS Portfolio requires an Activate Provider Org ID
  • An explicit request for the higher tierThe default offer is rarely the ceiling
Every item answers the provider's one question: will this company still be spending money here in three years?

In what order did the migration run, and how did it avoid downtime?

In four phases over about four months: audit and credit applications, workload placement, steady workloads moved to reserved capacity, then cross-training and retiring the DevOps role. Within each move, stateless services went first behind a traffic-shifting layer, the primary datastore went last, and a rollback path stayed live until the destination had proven itself.

Placement was decided per workload, by price and credit position, not by preference. Loosely coupled workloads with their own data can sit on a second cloud with little coupling cost, because they talk to the rest of the system across an API rather than a shared database. Tightly coupled pieces stayed together, because splitting a chatty service from its database across two clouds would have spent the credit on cross-cloud traffic.

Zero downtime came from sequencing, not tooling. Stateful systems were replicated and cut over with a short write pause rather than an outage, and the rollback path stayed warm until the events that worried me had passed. The egress, dual-running and identity costs of a move like this are broken down in GCP to AWS migration cost.

Budget for the data to move more than once. A migration that cannot take a long outage copies data at least three times: a bulk sync, one or more delta syncs, and a final catch-up at cutover. That makes deleting data nobody needs, before the first sync, the cheapest saving in the project: it shrinks the transfer bill and the destination's storage bill at the same time.

The migration, month by month
  1. Month 1
    Audit and credits

    Three months of billing data mapped line by line, and both credit applications submitted before any code moved.

  2. Month 2
    Workload placement

    Each service assigned a target cloud by pricing and credit position, not by preference.

  3. Month 3
    Managed to reserved

    Steady workloads moved to self-hosted equivalents on committed capacity. Spiky workloads deliberately left alone.

  4. Month 4
    Cross-training

    Application engineers took ownership of deploys and infrastructure, and the dedicated DevOps role was retired.

Each phase paid for the next: the audit made placement possible, the credits paid for the overlap, and a simpler estate made the last phase safe.
The order inside each move
  1. 1
    Stateless servicesfirst

    Moved behind a traffic-shifting layer, so traffic could shift in steps and shift back.

  2. 2
    Queues, schedulers and cachesnext

    Moved in dependency order while the stateless tier was already serving from the destination.

  3. 3
    Primary datastorelast

    Replicated, then cut over with a short write pause instead of an outage.

  4. 4
    Rollback pathkept warm

    The source stayed available until the destination had carried real traffic through the events that worried me.

Moving the database first is the mistake that turns a migration into an incident. Everything else in this list exists to make that last step small.

Which cloud cost changes removed the most money?

Moving steady workloads off per-request managed services onto self-hosted equivalents on reserved capacity removed the most. The other changes were placing each workload where its pricing and credits were best, removing per-gigabyte network paths, aggressive storage tiering, and model routing for inference. Right-sizing mattered least.

The test for each workload was simple. If it runs all day at a predictable load, per-request pricing charges a premium for elasticity it never uses, and reserved capacity is cheaper. The published ceilings show the size of that premium: AWS lists up to 72% off on-demand for EC2 Instance Savings Plans and up to 66% for Compute Savings Plans (AWS), and Azure lists up to 72% for reserved virtual machines (Azure). Spiky workloads stayed serverless, because moving them would have cost more.

Storage and network were the quiet lines. Rarely read data moved to colder classes: on AWS, S3 Standard is $0.023 per GB-month against $0.00099 for Glacier Deep Archive, per the rates in the AWS cost optimization checklist. Per-gigabyte network paths were removed wherever a cheaper route existed, because a per-gigabyte meter never shows up in a CPU graph.

Inference got its own treatment. Routing sends each request to the cheapest model that handles it well and keeps the expensive model for the requests that need it, which breaks the one-to-one link between usage and cost. Routing, caching and batching are covered in LLM inference cost optimization.

Managed or reserved?
Should this workload stay on per-request pricing?
Steady, predictable load
Move to reserved capacity

You are paying an elasticity premium for elasticity you never use. This is where the large savings were.

Spiky, unpredictable load
Stay serverless

Per-request pricing is cheaper here. Moving it would raise cost and add operational work.

Steady but tiny
Leave it alone

The engineering hours cost more than the saving. Not every line item is worth touching.

Two of the three branches say do nothing, which is why blanket advice to go serverless, or to go reserved, tends to disappoint.

How did removing the DevOps function save money without slowing delivery?

By cross-training application engineers to own deploys and infrastructure, so the role stopped being a queue. Every deploy, infrastructure change and incident had routed through one person, which was expensive and slow. Once the estate was simpler, engineers could own their own changes, and delivery held because the bottleneck left with the line item.

The order matters. Cross-training came last, after the architecture had removed most of what a specialist had been needed for. Asking application engineers to own a sprawling estate of per-request services and hand-tuned databases would have moved the queue, not removed it.

It is the same lean-team logic as a separate change I made to the engineering organisation, restructuring 20 people into a cross-functional team of seven and taking its run rate from $60K to $12K a month. A small team with broad skills can move without coordinating across silos, which is also why one person could run the migration.

The caveat: this works when your infrastructure is simple enough for application engineers to own. If you run multi-region Kubernetes under strict compliance requirements, keep the specialists.

What did the migration save, and what did the credits pay for?

The run rate fell from more than $200K a year, projected, to under $60K, a cut of more than 70% at list price, and the dedicated DevOps role went to zero. Separately, $300K of credits covered the bills on the destination clouds while they lasted. Only the first number is a saving; the credits were runway.

Why keep them apart? A saving that depends on credits disappears the day they expire, and the board sees the bill jump. A saving from architecture survives expiry. The credits were worth more used this way: they paid for the months when both clouds ran, on workloads that were already cheaper at list price.

Against the projected run rate, more than $140K a year stopped leaving the company, and it accumulates: every month the bill stays down is money that never leaves again. The other effect was slower to see. With the queue gone, infrastructure changes stopped waiting on one person, and cost became something engineers could see and act on.

Report a cut the same way if you want it to survive scrutiny: the run rate before and the run rate after, both at list price, with credits shown as a separate line. A cut that only exists net of credits will not survive the first review after they expire, and the team that reported it loses the credibility it needs for the next change.

Saving versus runway
pick
The 70% cut
Architecture: survives expiry
  • Run rate from more than $200K to under $60K a year
  • Measured at list price, with no credits applied
  • Came from what the estate pays for, not from a discount
  • Holds after every credit is spent
The $300K of credits
Runway: expires
  • $200K Microsoft for Startups, $100K AWS Activate
  • Paid destination bills during the move
  • Bought time to migrate without a deadline
  • Not counted in the 70%
Two numbers, reported separately on purpose. Adding them together would describe one year's balance sheet, not an operating cost.

What would I do differently next time?

Instrument list-price spend and tag every resource from day one. For a while the bill looked healthy because credits were absorbing it, and reconstructing the true unit economics afterwards took real time. The rule I would keep: buy commitments only after the architecture stops moving, never in the middle of a migration.

The tagging lesson is specific. I spent longer than I should have working out which team owned which line item, and every hour of that was an hour not spent cutting. On a fresh estate, tags are nearly free; on an old one, reconstructing ownership is the most tedious work in cost engineering.

The second lesson is about ownership. Nobody is promoted for lowering the cloud bill, so nobody owns it, and it grows until someone senior sees a number they cannot explain. Give one senior engineer the bill as a metric with a target. The return beats most features you could ship in the same quarter, and unlike most features it keeps paying every month.

How can you apply this playbook to your own cloud bill?

Run the same steps in order: map three months of billing by line, apply for credits before moving anything, classify every workload as steady, spiky or tiny, move steady ones to reserved or self-hosted capacity, tier storage and route model calls, then give one engineer the bill with a target.

Do not start with a migration. Most of the saving people expect from changing clouds (commitments, storage lifecycle rules, network paths, log retention) is available on the cloud you are on. Migrate only for a structural reason: a pricing model that fits your workload better, a service you need, a compliance requirement, or credits large enough to fund the move several times over, which is the position I was in.

Size the prize before spending engineering time on it. Put your current bill into the cloud cost calculator to see where the per-request and per-gigabyte lines sit. If the bill belongs to a product you have not built yet, the AI product cost estimator prices the build and its monthly hosting, with cloud credit applications as an option. Below about $2,000 a month, set billing alerts and revisit next quarter; the engineering time costs more than the saving.

If you would rather have the audit and the re-architecture done for you, that is cloud cost optimization: a fixed-fee audit that maps every leak, or a savings share where I implement the changes and am paid from the verified savings, so if your bill does not drop, I do not get paid.

The playbook
The cost playbook, in order
  • Export three months of billing and map every line to an ownerNothing is fixable until it is visible
  • Apply for credits before any code movesThey pay for the overlap while both clouds run
  • Classify each workload: steady, spiky, or steady but tinyOnly the first class moves to reserved capacity
  • Move steady workloads to reserved or self-hosted capacityAWS and Azure publish reserved discounts of up to 72%
  • Tier storage and remove per-gigabyte network paths
  • Route model calls by taskA cheap model by default, a larger one when needed
  • Retire queues by giving engineers ownership of their changesOnly once the estate is simple enough to own
  • Give one senior engineer the bill as a metric with a targetThe only step that makes the others repeat
The first two steps cost almost nothing and make every later step calm instead of rushed.

Cutting a cloud bill: common questions

→How much can cloud cost optimization save?

It depends on how much of your bill is steady load on per-request pricing, which is where large cuts come from. My cut was more than 70% of the annual run rate, from more than $200K to under $60K, and it came from re-architecture and a migration. It excludes the $300K of credits, which I treat as runway rather than savings.

→How long does a cloud cost migration take?

Mine ran in four phases over about four months: audit and credits, workload placement, steady workloads moved to reserved capacity, then cross-training. A mid-size GCP to AWS move typically takes three to five months end to end, with two to four of them running both clouds, which is why credits secured first matter so much.

→How do you get $300K in cloud credits?

By leading with production usage rather than a plan. I secured $200K from Microsoft for Startups and $100K from AWS Activate by showing high-volume production AI workloads, their current volumes and the account each provider would keep afterwards. Microsoft now advertises up to $150,000, and AWS Activate Portfolio up to $200,000 with an Activate Provider Org ID.

→Is multi-cloud worth the operational complexity?

Only when workloads split along natural boundaries and each is placed by price and credits. Independent workloads with their own data can run on a second cloud with little coupling cost. A tightly coupled application split across clouds pays cross-cloud transfer on every request and doubles the on-call surface, which usually costs more than the credit it chased.

→Do cloud credits count as savings?

Not in my accounting. Credits expire, and a cost model that only works while someone else pays for compute breaks on the day they run out. Track two numbers every month, credited spend and list-price spend, and manage the business on the second. My 70% is a list-price figure; the $300K of credits sat on top as runway.

→What is the biggest mistake teams make with cloud costs?

Starting with instance right-sizing. It is per-resource work with per-resource payoff, it decays with every deploy, and it was the smallest lever in my migration. Moving steady workloads off per-request pricing, removing per-gigabyte network paths and tiering storage remove whole line items, and credits applied for first pay for the time to do it.

Take this into your own chat

Open the article in your assistant with one click and ask it how this applies to your product.

Keep reading

See it in production: how a $200K a year cloud bill was cut by more than 70%

Ready to talk numbers?

Twenty minutes, straight to the engineer. No sales rep, no deck.