“Moving to the cloud” is often described as if it were a single decision with a single outcome: cheaper, faster, more reliable. It is none of those automatically. The cloud is a way of renting infrastructure by the hour, and whether that turns out cheaper, faster or more reliable depends almost entirely on how the environment is designed and operated.
This guide explains what cloud architecture actually involves, how to decide whether you need it, how migrations are structured, and where security, backups and cost fit in. It sticks to what the major providers document themselves, and avoids the confident percentages and savings claims that circulate in vendor marketing.
What cloud architecture covers
Cloud architecture is the design of the environment an application runs in. It is broader than choosing a provider, and it usually includes:
- Compute: virtual machines, containers, or serverless functions — and how they scale.
- Networking: how services connect, what is exposed to the internet, and what is kept private.
- Data: databases, object storage and caching, and how each is backed up.
- Identity and access: who and what may do which things.
- Delivery: how code and infrastructure changes reach production (CI/CD and Infrastructure as Code).
- Observability: logs, metrics and alerts, so problems are seen before customers report them.
- Resilience and recovery: what happens when a component, a zone or a region fails.
- Cost controls: budgets, tagging and regular review of what is being paid for.
None of these is an afterthought. They interact — a decision about how data is stored affects security, recovery time and cost at once — which is why they are best decided together, at the design stage.
Do you need the cloud at all?
Often, yes; but not always in the form vendors present. A small business website, a brochure site or a modest internal tool may run perfectly well on managed hosting or a single virtual server, with far less to configure and far less to go wrong. A complicated cloud architecture built for a workload that does not need it adds cost and operational effort without adding anything the business can feel.
Cloud infrastructure tends to earn its complexity when:
- Demand varies significantly, so capacity should grow and shrink rather than sit idle.
- Downtime is expensive enough to justify redundancy across zones or regions.
- The application needs managed services — queues, managed databases, identity, analytics — that would be costly to run yourself.
- Several environments (development, staging, production) need to be created and repeated reliably.
- Compliance or customer requirements call for audit logging, encryption controls and documented access.
The honest recommendation for a new product is usually to start with the simplest architecture that meets today’s needs and leaves a clear route to grow. Kubernetes, microservices and multi-cloud designs each carry real operational cost and are worth adopting when a specific requirement calls for them, not in advance.
The six design pillars
The major providers publish design guidance that is broadly consistent. The AWS Well-Architected Framework organises it into six pillars, and Microsoft’s and Google’s equivalents cover very similar ground under slightly different names.
| Pillar | The question it asks |
|---|---|
| Operational excellence | Can we deploy, monitor and improve this reliably? |
| Security | Are data and systems protected, and is access controlled? |
| Reliability | Does it recover from failure and meet the availability we need? |
| Performance efficiency | Are we using the right resources, and do they scale with demand? |
| Cost optimisation | Are we paying for what we use and no more? |
| Sustainability | Are we minimising the environmental impact of the workload? |
These pillars pull against each other, and good design is largely about deliberate trade-offs. Higher availability costs more. Tighter security can slow delivery. The useful discipline is to make those trade-offs explicitly — with the business owner in the room — rather than discovering them during an outage or when the first invoice arrives.
Shared responsibility
One misunderstanding causes a disproportionate share of cloud problems: the belief that the provider secures everything. Every major provider operates a shared responsibility model. In AWS’s wording, the provider is responsible for security of the cloud — the data centres, hardware and core services — while the customer is responsible for security in the cloud: their data, accounts, access controls and configuration.
Where the line falls shifts with the type of service. With a virtual machine you are responsible for the operating system and everything above it; with a managed database the provider handles more, and you still own access, data and configuration. A publicly readable storage bucket, an over-permissive access role or an unpatched server are customer-side problems regardless of how secure the provider’s data centre is.
Cloud migration: the seven strategies
Migration is not one activity. AWS’s prescriptive guidance describes seven common migration strategies, often called the “7 Rs”. The vocabulary is AWS’s, but the choices apply on any platform.
| Strategy | What it means | Typically suits |
|---|---|---|
| Retire | Switch off what is no longer needed | Unused or duplicated systems found during assessment |
| Retain | Leave it where it is, for now | Systems with hard dependencies or recent investment |
| Rehost | Move as-is onto cloud virtual machines (“lift and shift”) | Fast moves with minimal change, or a first stage |
| Relocate | Move virtualised workloads at the hypervisor level | Large virtualised estates, mostly in enterprise contexts |
| Repurchase | Replace with a different product, often SaaS | Commodity software such as email or CRM |
| Replatform | Move with targeted changes, such as a managed database | Reducing operational work without a rewrite |
| Refactor / re-architect | Redesign to use cloud-native services | Applications where scale or agility justify the effort |
Two practical points follow. First, a single application rarely needs the most ambitious strategy — rehosting first and modernising later is frequently the lower-risk route, provided the second step actually happens. Second, the assessment stage that decides which strategy applies to which system is where most of the value in planning sits. Migrating something that should have been retired is pure waste.
What a sensible migration plan includes
- An inventory of applications, data, dependencies and owners.
- A strategy chosen per workload, not one for the whole estate.
- A target architecture, with security and networking settled before anything moves.
- A data migration approach, including how much downtime the cutover can tolerate.
- A rehearsal — a full trial migration into a test environment.
- A rollback plan, so a failed cutover is an inconvenience rather than an emergency.
- Post-migration tuning and clean-up, including switching off the old environment.
Lock-in and portability
Every platform choice creates some dependency. Using a provider’s proprietary database, queueing or serverless services makes a workload quicker to build and harder to move; using only generic virtual machines keeps it portable and gives up much of what makes the cloud useful. Neither extreme is right for everyone, and the trade-off is worth deciding consciously.
Regulation is starting to address this in Europe. The EU Data Act introduces rules on switching between cloud providers, including phasing out switching charges. The European Commission’s explainer sets out the timetable. Whether or not it applies to you, it reflects a commercial reality: keep your data exportable, keep your infrastructure defined as code, and know what leaving would involve before you sign up.
DevOps and CI/CD
DevOps is the practice of bringing development and operations together so software is built, tested, released and run as a single continuous process. CI/CD — continuous integration and continuous delivery — is its most visible mechanism: every change is built and tested automatically, and a passing change can be released through a repeatable, automated pipeline instead of a manual checklist.
The reason to care is not fashion; it is that manual releases are slow, inconsistent and stressful, and stressful releases get put off, which makes each one larger and riskier. Automation makes small, frequent releases safe.
The research programme DORA measures delivery performance with four widely used metrics:
- Deployment frequency — how often changes reach production.
- Lead time for changes — how long a change takes from commit to production.
- Change failure rate — how often a release causes a problem that needs remediation.
- Time to recover — how quickly service is restored after a failed release.
These are worth tracking as a team’s own trend over time, not as a scoreboard against other organisations. The metrics work together: shipping faster while failure rates climb is not progress.
A sound baseline pipeline
- Every change goes through version control and a reviewed pull request.
- Builds and automated tests run on each change.
- Dependency and vulnerability checks run in the pipeline.
- Deployment to staging is automatic; production deployment is one controlled step.
- Releases can be rolled back quickly.
- Secrets are held in a secrets manager, never in the repository.
Infrastructure as Code
Infrastructure as Code (IaC) means describing servers, networks, databases and permissions in files that are kept in version control and applied by tooling, rather than created by clicking through a console. Terraform and OpenTofu are common cross-platform tools; CloudFormation (AWS) and Bicep (Azure) are provider-specific options; Pulumi lets you write infrastructure in general-purpose languages.
The benefits are practical:
- Environments can be recreated identically, so staging really resembles production.
- Changes are reviewed, recorded and reversible, like any other code change.
- The environment is documented by definition — the code is the description.
- Recovery from a serious failure can mean re-running the code rather than remembering what was configured.
- Configuration drift — where hand-made changes make environments diverge — is easier to detect.
IaC is not automatically appropriate for everything. For a single small server, the overhead may exceed the benefit. It earns its place as soon as there are multiple environments, more than one person making changes, or a need to rebuild reliably.
Storage, serverless and scaling
Choosing storage
Cloud providers offer several storage types that suit different jobs: object storage for files, images, backups and static assets; block storage attached to servers; file storage for shared filesystems; and managed databases for structured data. Choosing between them is a question of access pattern, durability needs and cost, and the cheapest option for archives is rarely appropriate for data that is read constantly.
When serverless fits
Serverless computing — functions such as AWS Lambda, Azure Functions or Google Cloud Functions — runs code in response to events and charges for execution rather than for idle servers. It suits intermittent or spiky workloads, event-driven processing, scheduled jobs and lightweight APIs.
It is not a universal answer. Functions have hard limits — AWS Lambda, for example, caps a single invocation at 15 minutes — so long-running processes need a different home. Steady, high-volume workloads can be cheaper on ordinary servers or containers, and heavily serverless designs can be harder to test locally and to move between providers.
Scaling without waste
Auto-scaling adds and removes capacity as demand changes, and load balancers spread traffic across it. The important part is setting sensible limits: a minimum that keeps the service responsive, and a maximum that stops a runaway process or an attack from producing an unlimited bill. A content delivery network (CDN) in front of the application reduces load and latency for static and cacheable content.
Cloud security
Cloud security follows from the shared responsibility model: the provider gives you strong tools, and it is your configuration that determines whether they are used. A dependable baseline includes:
- Multi-factor authentication on every human account, and especially on the root or global administrator account.
- Least privilege: each person and service gets only the permissions it needs, and unused permissions are removed.
- No shared credentials, and no long-lived access keys where short-lived credentials are available.
- Encryption in transit and at rest, with keys managed deliberately.
- Private networking for databases and internal services, with only the necessary entry points exposed.
- Storage locked down by default — public access is an explicit, reviewed exception.
- Audit logging switched on and retained, with alerts for high-risk changes.
- Regular patching of operating systems, containers and dependencies.
- Separate accounts or projects for production and non-production environments.
- Secrets held in a secrets manager, not in code or configuration files.
Many of these are straightforward to set up when an environment is built and awkward to retrofit later, which is the case for designing security in from the beginning. Independent guidance such as the CIS Benchmarks provides platform-specific configuration checklists that are useful for review.
Backups and disaster recovery
Two numbers should drive every recovery decision, and both are business decisions rather than technical ones:
- Recovery Point Objective (RPO): how much data you can afford to lose, measured in time. An RPO of one hour means losing up to an hour of recent data is tolerable.
- Recovery Time Objective (RTO): how long you can afford to be offline before the impact becomes unacceptable.
Tighter targets cost more. AWS’s disaster recovery guidance describes four broad approaches, and the pattern holds on other platforms too:
| Approach | How it works | Recovery speed and cost |
|---|---|---|
| Backup and restore | Regular backups stored elsewhere; rebuild when needed | Slowest recovery, lowest cost |
| Pilot light | Core components kept running at minimum size in a second location | Faster; moderate cost |
| Warm standby | A scaled-down but working copy in a second location | Faster still; higher cost |
| Multi-site active/active | The workload runs in several locations simultaneously | Fastest; highest cost and complexity |
Most small and mid-sized businesses are well served by the first or second option, matched to a realistic RTO. Multi-region active/active is rarely justified unless an hour of downtime costs more than the additional infrastructure and engineering.
Backups that will actually work
- Follow the widely cited 3-2-1 approach: three copies, on two different types of storage, with one held off-site or in a separate account.
- Keep at least one copy immutable or otherwise protected from deletion, as a defence against ransomware and accidental removal.
- Test restores on a schedule. A backup that has never been restored is an assumption, not a safeguard.
- Back up configuration and infrastructure code as well as data.
- Write the recovery procedure down, and make sure more than one person can follow it.
Cost optimisation
Cloud costs are variable by design, which is both the appeal and the risk. The practice of managing them is often called FinOps, and the FinOps Foundation frames it as a collaboration between engineering, finance and the business, rather than a one-off clean-up.
The levers that consistently matter are unglamorous:
- Visibility first: tag resources by project, environment and owner so every cost has an explanation.
- Budgets and alerts, so an unexpected spike is noticed within days rather than at month end.
- Right-sizing: matching instance and database sizes to measured usage instead of guesses.
- Switching off what is unused — idle development environments, orphaned disks, old snapshots, unattached addresses.
- Commitment discounts (such as reserved capacity or savings plans) for steady, predictable workloads, once usage is understood.
- Choosing storage classes to match how often data is actually accessed.
- Watching data transfer charges, which are easy to overlook when designing across regions.
Commitments deserve particular care: they reduce unit prices in exchange for a promise to pay, so they are only a saving if the usage genuinely continues. Committing before the workload is understood turns a flexible cost into a fixed one.
Mistakes worth avoiding
- Assuming the cloud is automatically cheaper, more secure or more reliable.
- Lifting and shifting an oversized on-premises setup without right-sizing it.
- Choosing Kubernetes, microservices or multi-cloud before any requirement calls for them.
- Configuring everything by hand in a console, with no record of what was done.
- Using one all-powerful account for everyone, or leaving MFA off the administrator account.
- Leaving storage publicly accessible without a deliberate reason.
- Taking backups but never testing a restore.
- Having no budgets or alerts, and discovering costs only on the invoice.
- Migrating everything, including systems that should have been retired.
- Treating the migration as finished at cutover, with no tuning or clean-up afterwards.
Questions to ask a cloud partner
- Why this architecture for our workload, and what simpler alternative did you consider?
- Which parts of security remain our responsibility, and how are they documented?
- What will this cost to run each month, and what assumptions is that estimate based on?
- How will infrastructure be defined and versioned, and can we take the code with us?
- What are the recovery targets, and how will restores be tested?
- How does the migration handle cutover, and what is the rollback plan?
- What happens after launch — monitoring, patching, and cost review?
A partner who answers these plainly, including where the honest answer is “you don’t need that yet”, is generally more useful than one who leads with the largest architecture.
Frequently asked questions
Not automatically. Cloud pricing rewards workloads that are sized properly and can scale with demand. A steady workload on oversized resources can cost more than a fixed server. Estimate costs before you move, set budgets and alerts, and review usage after launch.
Final thoughts
Good cloud architecture is mostly a series of ordinary decisions made early and deliberately: what the workload needs, who is responsible for what, how changes are released, how data is recovered, and what it should cost each month.
The platform features are the easy part. What separates a dependable environment from a fragile and expensive one is whether those decisions were made on purpose, written down, and revisited as usage changed.
If you are planning a migration or a new environment, start with the requirements — recovery targets, compliance needs, expected demand and budget — and let the architecture follow from them, not the other way around.


