TL;DR
- Hybrid pays off when latency, data residency or legacy integration block full public cloud — not when teams want to avoid a migration decision.
- Define a landing zone first: identity, networking, logging and backup — then place workloads.
- Measure total cost: egress, dual operations teams and incident paths across two estates.
- Exit criteria matter: every hybrid design should state what triggers consolidation or repatriation.
“We need hybrid” is one of the most common opening lines in cloud workshops — and one of the least precise. Hybrid is not a product SKU; it is an operating model where some workloads stay on premises or in a sovereign region while others run on public cloud. That can be the right answer for a mid-market manufacturer with shop-floor PLCs that cannot tolerate a 40 ms round-trip to Frankfurt. It can also be a way to postpone hard choices until two parallel platforms are burning budget and nobody owns the integration layer between them.
At QData we see both patterns every quarter. The difference is rarely technology. It is whether leadership has named the constraint that hybrid is meant to solve, set a review date, and funded operations for two estates without pretending one team can run both as a hobby.
When hybrid is a rational choice
Three signals consistently justify a split estate in our architecture reviews. The first is latency and locality. Manufacturing lines, trading floors, clinical devices or warehouse automation often need compute within metres or milliseconds of the equipment. Moving batch analytics to cloud while keeping the control loop on-prem is a classic, defensible split — provided you document which direction data flows and who approves changes to that path.
The second signal is data gravity and law. A Polish insurer we advised could not move raw policyholder records to a US-region SaaS, but could run actuarial models on anonymised aggregates in EU cloud. The boundary was explicit: pseudonymised extracts only, nightly, with a DPO sign-off on the transformation job. Hybrid here is not indecision; it is a legal partition with technical enforcement.
The third is legacy coupling. An ERP that cannot be containerised in the next 12–18 months but must exchange orders and inventory with a new customer portal is a legitimate hybrid candidate — if you write the integration contract, name the sunset date for the old interface, and refuse to add a fourth ad-hoc VPN for “just one more” satellite office.
In these cases hybrid is a transition or permanent partition with explicit boundaries. The mistake is treating it as “best of both worlds” without naming who operates each side, who owns incidents that cross the boundary, and how identity flows between them.
When hybrid is a trap
Hybrid fails when chosen for comfort. We regularly audit estates where old VMs run “just in case”, two monitoring stacks disagree on whether a service is up, and patch Tuesday happens twice a month because no one owns consolidation. Teams pay for cloud flexibility while still running a full on-prem ops team — without the automation that makes either side efficient.
Another trap is network spaghetti. Site-to-site links, hairpinned traffic and manual firewall tickets for every new microservice turn hybrid into friction. If adding a service requires a change window across two network teams and a CAB that meets fortnightly, you have not built resilience — you have built a queue. One client spent more on MPLS and cross-connect hours than they would have spent refactoring a single monolith into a managed API in cloud.
A subtler trap is identity drift. Users with duplicate accounts, different MFA policies on-prem vs cloud, and break-glass passwords stored in three places. Incidents that start in SaaS admin consoles end in Active Directory with no correlated logs. Hybrid without a single IdP strategy is two attack surfaces with a VPN between them.
A decision framework you can use in one workshop
Score each candidate workload on four axes from 1 to 5: refactorability, regulatory constraint, operational maturity and cost sensitivity to egress. Workloads with high regulatory constraint and low refactorability stay on-prem or in a managed private zone. Workloads with high refactorability and elastic demand move to public cloud. Everything in the middle goes to a pilot landing zone with a 90-day review and a named executive sponsor.
In a half-day workshop with a 200-person logistics firm, this exercise surfaced an obvious winner for cloud (customer-facing track-and-trace API) and an obvious stay (warehouse label printers on a flat L2 segment). The contentious middle — nightly route optimisation — went to pilot with a cap on data egress and a requirement to reproduce runs from Git. Politics quietened because the criteria were on the wall, not in someone’s inbox.
Document exit triggers alongside placement: for example, “if monthly cross-estate incident hours exceed 40, we consolidate networking” or “when the ERP REST facade ships, batch jobs move to cloud”. Without triggers, hybrid drifts forever and every new project defaults to “straddle both”.
BaseCloud as a landing zone — not a slogan
At QData we use BaseCloud as a structured starting point: identity integrated with your IdP, centralised logging, encrypted backup, baseline network segmentation and infrastructure-as-code for repeatable environments. Whether the landing zone sits on public cloud, a partner DC or a mix, the point is the same: one operational language — runbooks, monitoring, access reviews — before you scale workload count.
BaseCloud is deliberately boring. Few integration points. Standard tags. A map of which subnets may reach which data classes. Backup policies that match tier, not heroics. When an auditor or insurer asks “how do you know this VM is in scope?”, you open a dashboard, not a spreadsheet from 2019.
Hybrid architectures that survive audits and incidents are boring on purpose: strong observability, written data-flow diagrams, and a single on-call rotation that knows both sides — or a managed partner who does. If your current diagram needs a legend to explain arrow directions, simplify before you scale.
Scenario: manufacturer vs SaaS scale-up
Compare two clients with the same headline — “we are hybrid”. The manufacturer keeps SCADA and historians on-prem, sends anonymised telemetry to cloud for predictive maintenance, and uses SaaS for HR and CRM with SSO from Entra ID. Incidents have a runbook per tier; cross-boundary flows are three, not thirty. Total cost is known; exit trigger is “new production line gets edge gateway standard in Q3”.
The scale-up, meanwhile, runs production Kubernetes in cloud but keeps “temporary” PostgreSQL on a office NAS, dev databases on laptops, and a forgotten Jenkins on a VM “because pipelines are sensitive”. That is not hybrid; that is debt with a slide that says multi-cloud. The fix was consolidation into one landing zone and a moratorium on new on-prem except by architecture board approval.
Your organisation likely resembles one of these more than the other. Naming which story you are telling is the first step toward a hybrid design that will still make sense in eighteen months.
What to do next
Inventory ten workloads — not fifty in a wiki no one updates. Run the four-axis score in a room with infrastructure, security and a business owner. Pick one pilot that is allowed to fail safely, with rollback and a 90-day checkpoint. If you want an external sanity check on landing-zone design, migration sequencing or total cost of dual estates, our Cloud Services practice starts with architecture and operations — not shelfware licenses.
Want to discuss a project?
Contact us
