Writing · Cloud strategy
Own the Boring, Rent the Weird: A Workload-Level Infrastructure Strategy
How I choose between cloud, owned infrastructure and a deliberate hybrid one workload at a time.
“Are we a cloud company?” is usually the wrong infrastructure question.
It turns a portfolio of different workloads into an identity. A public API, a steady batch pipeline, a specialized model-training cluster, an industrial control workload and an archive do not share the same economics or failure model. Giving them one company-wide answer may simplify a slide. It does not simplify the system.
The practical choice sits at workload level: which capabilities to rent, which to operate directly and what evidence would justify moving the boundary later.
The false company-wide choice
Cloud computing has concrete properties. NIST’s widely used definition includes on-demand self-service, resource pooling, rapid elasticity and measured service. Those properties can create real option value when demand or technical direction is uncertain.
Owned infrastructure has concrete properties too: direct control of capacity, locality, hardware choice and depreciation across a useful life. Those properties can create advantage when utilization is steady and the organization can operate the system well.
Neither list decides the architecture. The decision depends on the workload and the people who will run it.
A useful strategy starts by separating these categories:
- commodity application and platform services;
- baseline compute with predictable utilization;
- burst capacity and experiments;
- specialized or scarce hardware;
- storage and data movement;
- control-plane services;
- workloads tied to a physical location or device fleet.
Different answers within one company are not inconsistency. They are evidence that the architecture reflects the work.
Use five questions, in order
Infrastructure debates often start with cost because cost appears quantifiable. The decision should start one level earlier. I use five questions:
- Shape: What does demand look like, including time, locality and failure tolerance?
- Capability: Is the requirement commodity, specialized, fast-moving or tied to a physical environment?
- Competence: Which operating model can the organization support reliably over the intended lifecycle?
- Control: Which capacity, data, interface or recovery rights must remain directly governable?
- Review: Which assumption would cause the boundary to move?
Cost belongs inside all five. It is not a separate answer. A lower unit price can be irrelevant when the organization lacks the competence to operate the capacity, when data movement dominates the workflow or when the design removes an option the product will need.
This sequence also prevents false precision. Before the workload and operating boundary are defined, a detailed cost model is usually a detailed comparison of different systems.
Start with workload shape
Before comparing providers or server quotes, describe the workload in operational terms:
- Is demand steady, periodic, bursty or unknown?
- How much headroom is required, and how quickly must capacity arrive?
- Can work be queued or scheduled?
- Is the workload coupled to a physical site, device or data source?
- What happens when capacity is unavailable?
- How quickly is the technology changing?
- What capability is scarce inside the organization?
The answer should include a time horizon. A workload that is unpredictable during product discovery may later become a stable baseline. A formerly predictable service may become volatile after a new integration or market launch.
Architecture is a decision made with expiring assumptions.
Predictable and variable utilization
Predictability changes the value of ownership.
For a stable baseline workload, owned capacity can be planned and used continuously. But the real comparison must include resilience headroom, maintenance, spares, power, cooling, networking and the people required to operate it. A utilization chart that ignores those requirements is not total cost.
For a variable workload, rented capacity can absorb peaks without committing to permanent headroom. But “elastic” is not a magic property of an invoice. The application must be able to scale, state must be managed, capacity may need reservations, and rapid growth can still hit provider limits.
Queueable work creates another option. If the output is needed within hours rather than milliseconds, demand can sometimes be shaped instead of capacity being scaled. AWS’s Well-Architected cost guidance explicitly treats throttles, buffers and queues as ways to manage demand and supply. That is an architectural choice, not merely a purchasing tactic.
| Workload property | Pressure toward rented capacity | Pressure toward owned capacity |
|---|---|---|
| Uncertain or bursty demand | Avoid early commitment; add capacity when needed. | Weak unless spare capacity has another productive use. |
| Stable, baseline-heavy demand | Useful when managed services remove meaningful operating burden. | Stronger when utilization and lifecycle are well understood. |
| Queueable work | Use flexible capacity and time-based pricing where appropriate. | Fill otherwise idle capacity without affecting interactive workloads. |
| Local physical dependency | Useful for coordination, fleet management and off-site resilience. | Strong for deterministic local response and disconnected operation. |
| Specialized, fast-changing hardware | Buy access while the requirement is moving. | Consider ownership only after sustained demand and operating ability are proven. |
Commodity versus specialized capability
The title “Own the boring, rent the weird” is intentionally provocative. “Boring” does not mean unimportant. It means stable, understood and available from multiple sources. “Weird” means scarce, fast-moving or hard to predict.
Renting weird capability can buy learning. A team can test whether a specialized accelerator, managed database or model service changes the product before committing to a long hardware or platform lifecycle. The premium is partly the price of the option not to be right yet.
Ownership becomes more attractive when the workload is durable, differentiation depends on the capability, utilization is high enough to justify dedicated capacity, and the team can operate it without neglecting the product.
The reverse is also true. A provider-native service may look inexpensive because the hard operating work has been externalized. Replacing it with generic infrastructure can create an internal platform product that nobody intended to own.
Operational competence is part of the price
Owned infrastructure includes patching, physical security, monitoring, failed disks, firmware, spares, power, networking, backup, recovery and incident response. Renting does not remove operations; it changes the layer. The team still owns identity, configuration, architecture, data protection, cost controls, service limits and provider incidents.
Do not compare a cloud bill with hardware acquisition cost. Compare two complete operating models over the period in which the decision is expected to remain valid.
Ask who carries the pager, who can diagnose a failure outside normal hours, how replacements arrive, how changes are tested and what happens when the person who designed the platform leaves. If the answer is “the infrastructure team,” include that team’s capacity as a constrained portfolio resource.
Rent uncertainty. Own stable advantage. Revisit the boundary when either one changes.
Data gravity is a constraint, not a slogan
Where is data created? Where is it transformed? Who consumes the result? How much must move, how quickly, and under which contractual or regulatory conditions?
These questions matter more than a general statement that “data has gravity.” A connected product may need local decisions during an outage while still benefiting from cloud-based fleet analysis. A large dataset may stay near a specialized compute environment while a small derived result moves elsewhere. An archive and an interactive working set may have different homes.
Record transfer time, egress exposure, synchronization semantics and failure behavior. Avoid using data location as a permanent excuse: compression, aggregation, changed retention or a different interface can alter the boundary.
Use cloud to buy option value
Early in a product or workload lifecycle, uncertainty is often more expensive than unit compute. The team does not yet know demand, performance shape, data volume, hardware fit or which part of the system will become differentiating.
Rented capability can shorten that learning loop. The mistake is allowing an experiment to become permanent architecture without a review point.
Every experiment should record boundary signals:
- sustained utilization above a defined threshold;
- a stable hardware and software requirement;
- cost becoming material at the portfolio level;
- operating competence becoming available;
- provider constraints affecting the product;
- a data or latency boundary becoming durable.
The threshold values will be organization-specific. Their purpose is not to predict the future perfectly. It is to prevent inertia from turning an experiment into permanent architecture without a decision.
Elasticity must be designed and used
Possible elasticity and exercised elasticity are different. A service may run on elastic infrastructure while remaining provisioned for peak demand at all times. Another may scale in theory but take longer to add capacity than the workload can tolerate.
Measure the actual shape:
- time to add and remove capacity;
- minimum efficient unit;
- state that prevents horizontal scaling;
- reserved baseline versus variable peak;
- queuing tolerance;
- provider and regional capacity risk.
This also prevents false precision in cost models. A model based on perfect scale-to-zero behavior is not useful when the workload carries a continuous baseline or operational policy requires warm capacity.
Strategic control and concentration
Provider concentration is not automatically a reason to build everything twice. A nominally portable architecture can impose real cost while remaining untested. Multi-cloud can become two incomplete operating models rather than one resilient system.
Define the control that matters:
- the ability to restore service elsewhere within an acceptable time;
- access to data in a usable format;
- control of critical interfaces and identity;
- the ability to operate locally during upstream failure;
- commercial leverage for a material dependency;
- compliance with customer or jurisdictional commitments.
Then design and test for that control. Portability matters only when it changes the consequence of a plausible event. Source compatibility, deployability, data recovery and operating readiness are different claims; do not collapse them into one architecture label.
Keep a workload decision record
Architecture decisions decay when their assumptions disappear. A short decision record keeps the boundary legible.
| Field | What to capture | Review trigger |
|---|---|---|
| Shape and criticality | Demand, latency, queueability, recovery objective and physical coupling. | Material change in demand or consequence. |
| Capability | Commodity, provider-native, specialized or location-bound requirements. | A capability becomes stable, scarce or strategically differentiating. |
| Economics | Current assumptions for capacity, people, resilience, transfer and lifecycle. | Assumption expiry or a material price change. |
| Operating owner | Team, competence, support model and failure response. | Ownership or staffing changes. |
| Reversibility | Data exit, interface boundaries, recovery path and migration effort. | A dependency becomes materially harder to leave. |
Strategy is the ability to move the boundary
Infrastructure decisions rarely stay correct forever. Experiments become stable workloads. Commodity capacity becomes scarce. Internal competence grows or erodes. Provider economics, regulation and product architecture change.
A company-wide ideology makes those changes feel like reversals. A workload decision record makes them normal operating reviews.
This gives a CTO something better than a company-wide cloud policy: a portfolio that can change when the facts change. The ability to move the boundary is the strategy.
Sources and further reading
- NIST, SP 800-145: The NIST Definition of Cloud Computing, defines five essential cloud characteristics, including on-demand self-service, rapid elasticity and measured service.
- AWS, Cost Optimization Pillar—Well-Architected Framework, organizes cost decisions around business outcomes, workload review and the management of demand and supply.
- AWS, Manage demand and supply resources, discusses scaling as well as throttling, buffering and queueing as workload-shaping mechanisms.