Engineering

Migrating an app from AWS to GCP

Editorial · Reveneau · May 7, 2026

Migrating an app from AWS to GCP

A cloud migration gets described internally as "lift and shift" (copying the system to the new cloud without changes) more often than it should be, because that description makes the project sound smaller than it actually is. Moving an application from AWS to GCP affects cost structure, reliability assumptions, security model, and the daily experience of every engineer who has to operate the system afterward. Treating it as a pure infrastructure swap is how migrations end up over budget and full of surprises.

In our work, the migrations that go well share one trait: the team spent real time deciding what the new system should look like before anyone changed a load balancer. The ones that go badly all made the same early mistake, which was to copy the AWS architecture onto GCP resource-for-resource and hope the bill and the incident count would take care of themselves. This post is the checklist we use with engineering teams: compute, storage, networking, IAM, managed services, cost, and the cutover itself.

Decide what actually needs to move

Not every piece of an AWS-based architecture deserves to be carried over unchanged. Some of it exists because of AWS-specific behaviours or past accidents, not because it is the right design. Before moving anything, it is worth auditing the current architecture with a real question in mind: if we were building this from scratch on GCP today, would we build it this way? Components that pass that test should move largely as-is. Components that only exist to work around an AWS limitation are a chance to simplify, not just relocate.

Compute is usually the clearest place to apply that question. An EC2 fleet with custom autoscaling logic maps cleanly to Compute Engine managed instance groups, so if the design is sound, move it. But many teams are running containers on ECS or Fargate and never checked again whether that design still fits. On GCP the equivalent conversation is GKE or Cloud Run, and the honest answer is sometimes "neither, exactly." Cloud Run in particular removes a class of infrastructure you may have been managing by hand. If a service is stateless and request-driven, the migration is the right moment to remove the cluster you were constantly managing rather than rebuild it on the new provider.

This step gets skipped constantly because it feels slower than just starting the migration. It is the single biggest factor in whether a migration finishes on budget or turns into slow, difficult work that lasts several quarters. Write the target architecture down first, service by service, and mark each one as "move as-is," "re-architect," or "retire." That single document does more to keep the project on schedule than any tooling choice you make later.

Cost modeling is not just comparing list prices

AWS and GCP pricing pages look comparable on paper, which leads teams to assume the migration is cost-neutral or even a savings. Actual costs depend heavily on usage patterns that do not map directly between providers. Egress pricing, sustained-use discounts, and how each platform prices its managed database and storage tiers all behave differently under real production load. The only reliable way to model this is to estimate against your actual traffic and storage patterns, not the listed price of comparable instance types.

Two areas surprise teams more than the rest. The first is data egress. If your system moves a lot of traffic between zones, out to users, or across to another provider, egress can become one of your largest costs without anyone noticing, and the rules differ from what you learned on AWS. The second is the discount model. AWS Reserved Instances and Savings Plans are a commitment you make up front. GCP applies sustained-use discounts automatically the longer an instance runs, and offers committed-use discounts on top for workloads you can predict. That difference changes how you plan capacity, so model your steady-state usage before you sign anything. Build the estimate against a real month of your own metrics, not a spreadsheet of comparable instance sizes, and you will avoid the most common post-migration budget surprise.

Data migration strategy determines your downtime

Moving the application code is usually the easy part. Moving the data, especially for a system with a large, actively-written database, is where migrations get risky. A one-time cutover, where you take the system offline, copy everything, and start it again on GCP, is the simplest approach and the one most likely to cause real customer-facing downtime.

A safer pattern for anything with meaningful traffic is a phased approach: replicate data continuously to the new environment, run both systems in parallel for a validation period, and move traffic once the new environment has proven it works under real load. It takes longer to execute but removes most of the risk of a bad cutover.

Storage affects this plan. Object storage tends to be the easy part: an S3 bucket has a clear counterpart in Cloud Storage, and bulk transfer of static assets is a common, well-understood task. Databases are where the real decisions are. RDS or Aurora on Postgres or MySQL moves naturally to Cloud SQL, and for something running on DynamoDB the closest fit is usually Firestore or Bigtable, which are not direct replacements and may require you to rethink the data model. Whatever the target, keep a way to roll back for the whole validation period. In our work we do not delete the AWS environment the day traffic moves. We keep it running and able to accept writes again until the new stack has handled real production load through a full business cycle, including the traffic peak you only see once a week or once a month.

IAM is where migrations fail without anyone noticing

AWS's IAM model and GCP's IAM model share a name and little else in terms of structure. AWS builds permissions around policies attached to users, roles, and resources with fairly granular control at each layer. GCP's model is built around a resource hierarchy, projects nested under folders nested under an organization, with permissions inherited down that hierarchy by default.

Teams that treat these as equivalent tend to end up with permission gaps or, more dangerously, permissions that are broader than intended because inheritance behaved differently than expected. This deserves its own dedicated review pass, separate from the general infrastructure migration, because getting it wrong is a security problem, not just an inconvenience.

The mistake we see most often is granting a role too high in the hierarchy. A permission set at the project level applies to every resource inside it, so a role that felt narrow on AWS can end up far wider on GCP once inheritance is in play. Permissions between services are different too. On AWS you rely on instance roles and assumed roles. On GCP the unit is the service account, and mapping each workload to a dedicated service account with only the permissions it needs is work you should do deliberately, not by translating your old policies line by line. Map the identities first, grant the minimum, then test that each service can reach exactly what it needs and nothing more.

Networking and managed service equivalents need real research, not assumptions

VPC design, load balancer behavior, and DNS all differ enough between AWS and GCP that assuming a one-to-one mapping exists is a mistake. The same goes for managed services: an AWS-managed queue or cache does not always have a direct GCP equivalent with the same guarantees, and sometimes the right answer is a genuinely different architecture rather than the "closest" service on the other platform.

Networking has a few surprises in its structure. A GCP VPC is global by default, with subnets scoped to regions, which works differently from the region-bound VPCs you build on AWS. Load balancing and DNS behave differently enough that you should treat them as new designs, not translations. On managed services the match is exact for some parts and approximate for others. SQS has a reasonable counterpart in Pub/Sub, though the delivery and ordering semantics are not identical, so read them carefully rather than assuming your consumers will behave the same. A cache built on ElastiCache maps to Memorystore. The queue and streaming cases are where choosing the "closest service" causes outages, because the guarantees your code depended on without anyone noticing may not exist on the new platform. Read the semantics for every managed service on the critical path, and where they differ, decide on purpose rather than discovering it in production.

Reducing the risk of the cutover

The cutover is one moment, but the confidence to run it comes from everything before it. Set up the full GCP environment and run it in parallel for a while, without customers using it: copy real traffic to it, compare its behavior against production, and let it fail where no customer is affected. Move traffic in small parts rather than all at once, using DNS or a load balancer to move a small percentage first and watch your metrics before you send more. Keep the AWS side able to take traffic back for the entire period. A cutover you can reverse in minutes is a cutover you can run without fear.

Where this goes right

We treat cloud migrations as an architecture decision first and an infrastructure task second, which is the only way to avoid creating the same problems again on new infrastructure. If your team is weighing a migration like this, it is exactly the kind of engineering-heavy work we take on through staff augmentation or a dedicated full product build engagement, depending on how much of the surrounding system needs to change alongside the move.

Migration is one of the areas where AI-generated code helps most, and one where it is most dangerous to trust. Translating configuration and rewriting service calls is mechanical, well represented in training data, and fast, so a large part of the routine work becomes faster.

The risk is that migration failures are almost never in the routine work. They come from small differences in behaviour between the two platforms in a case nobody tested: a timeout default, an ordering guarantee, a permission model that is nearly but not quite the same. Generated code will copy your structure accurately and has no way to know which of those differences will matter to you. Let it do the translation. Do not let it do the verification.

Common questions

Is a cloud migration just a lift and shift?

Rarely. Treating a migration as a pure lift and shift, copying everything unchanged, usually means keeping infrastructure decisions that were made because of AWS-specific behaviours, which often do not make sense on GCP. The bigger, harder decision is which components should move as-is and which should be re-architected to fit the new platform properly.

What are the main considerations when moving between cloud platforms?

Cost modeling differences (list prices look similar but usage patterns and egress costs do not convert directly), a real data migration strategy that minimizes downtime, IAM model differences between the two providers, networking reconfiguration, and finding the right managed service equivalents rather than assuming a one-to-one mapping exists.

Why would a team move from AWS to GCP in the first place?

The reasons vary by team: cost structure on their real usage, a better fit for specific managed services, or a chance to simplify an architecture that grew around AWS-specific workarounds. The migration is worth it when it lets you remove infrastructure you were managing by hand, not just move the same system to a new provider.

How do AWS compute services map to GCP?

An EC2 fleet with custom autoscaling maps cleanly to Compute Engine managed instance groups. Containers on ECS or Fargate match GKE or Cloud Run, and for stateless request-driven services Cloud Run often lets you remove the cluster you were managing by hand rather than rebuild it.

What is the GCP equivalent of S3, RDS, and DynamoDB?

An S3 bucket has a clear counterpart in Cloud Storage, and RDS or Aurora on Postgres or MySQL moves naturally to Cloud SQL. DynamoDB is the harder case: the closest fits are Firestore or Bigtable, which are not direct replacements and may require you to rethink the data model.

How is GCP IAM different from AWS IAM?

AWS builds permissions around policies attached to users, roles, and resources with granular control at each layer. GCP is built around a resource hierarchy, projects nested under folders nested under an organization, with permissions inherited down that hierarchy by default. Treat them as different systems, because a role that felt narrow on AWS can end up far wider on GCP once inheritance is in play.

How do you handle service-to-service permissions on GCP?

On AWS you rely on instance roles and assumed roles. On GCP the unit is the service account, so map each workload to a dedicated service account with only the permissions it needs, granted deliberately rather than translated line by line from your old policies.

What networking differences should I plan for?

A GCP VPC is global by default with subnets scoped to regions, which works differently from the region-bound VPCs you build on AWS. Load balancing and DNS behave differently enough that you should treat them as new designs, not direct translations. The same caution applies to managed queues and caches on the critical path: read the delivery and ordering semantics for each service before assuming the closest equivalent behaves the same way.

How do managed queues and caches translate to GCP?

A cache on ElastiCache maps to Memorystore, and SQS has a reasonable counterpart in Pub/Sub, though the delivery and ordering semantics are not identical. The queue and streaming cases are where choosing the closest-looking service causes outages, so read the semantics for every service on the critical path before you commit.

How do you reduce the risk of the cutover?

Set up the full GCP environment and run it in parallel first, without customers using it: copy real traffic to it and compare behavior against production. Move traffic in small parts using DNS or a load balancer, watch your metrics, and keep the AWS side able to take traffic back for the whole window so you can reverse in minutes.

How long does an AWS to GCP migration take?

The timeline is set by how much you re-architect versus move as-is, and it depends mostly on two things: the up-front decision on what should move versus be rebuilt, and the parallel-run validation window for the data. A phased cutover takes longer to execute than a one-time switch, and that extra time is what removes most of the risk.