SR. CLOUD PLATFORM & SRC LEAD
kawan lama group- Posted 10 hours ago
- Be among the first 10 applicants
Job Description
We're hiring a Senior Cloud Platform & SRC Lead to set standards, architect platform-wide systems, and turn fuzzy organizational pain into funded, measurable platform programs.
This is an architect-and-influence role. You'll own the design and rollout of platform initiatives that span teams, regions, and accounts — multi-region HA, the Zero-Trust program, the supply-chain baseline, the FinOps strategy, the AI/ML infrastructure bet. You'll write architecture decision records that stick, mentor Mids toward Staff, and be the person the org trusts to de-risk big bets.
If you've architected multi-region platforms, driven org-level reliability or security programs, and want a team that treats the platform as a product with users and a roadmap — this is the role.
Responsibitly:
- Has architected multi-region or multi-account platforms — can defend active-active vs active-passive, cell-based vs replicated, and the consistency/cost/complexity trade-offs of each.
- Reliability program ownership — has run an org-wide SLO program, used error budgets as a governance tool (not just a metric), led game days / DR exercises, and owned post-incident reviews that changed how the org works.
- Security program ownership — has driven Zero Trust, supply-chain, or compliance-automation as a program, not a feature.
- FinOps leadership — has owned unit economics across teams, set commitment strategy, run chargeback/showback, and made platform cost a first-class engineering concern.
- Deep cloud-native fluency — Kubernetes (multi-cluster, operators, CRDs, policy), service mesh (Istio/Linkerd/Cilium), GitOps (ArgoCD/Flux + progressive delivery), IaC (Terraform + CDK/Pulumi/Crossplane), OpenTelemetry, eBPF-based observability. Can name the trade-off of every tool they propose.
- Architecture communication — writes ADRs and strategy docs that survive review and drive decisions. Can frame a problem before proposing a solution.
- Influence and ambiguity — turns fuzzy org pain into a funded initiative, aligns cross-functional stakeholders, sequences work, says what to cut.
- Mentorship — has grown at least one engineer to the next level; can calibrate the hiring bar; sponsors others growth.
Qualification (must-have):
- 5+ years in cloud infrastructure/platform/SRE, with 2+ years operating Kubernetes at scale (multiple clusters, production traffic, real on-call).
- Has architected multi-region or multi-account platforms — can defend active-active vs active-passive, cell-based vs replicated, and the consistency/cost/complexity trade-offs of each.
- Reliability program ownership — has run an org-wide SLO program, used error budgets as a governance tool (not just a metric), led game days / DR exercises, and owned post-incident reviews that changed how the org works.
- Security program ownership — has driven Zero Trust, supply-chain, or compliance-automation as a program, not a feature. Knows SPIFFE/SPIRE, workload identity, admission control, and breach-readiness.
- FinOps leadership — has owned unit economics across teams, set commitment strategy, run chargeback/showback, and made platform cost a first-class engineering concern.
- Deep cloud-native fluency — Kubernetes (multi-cluster, operators, CRDs, policy), service mesh (Istio/Linkerd/Cilium), GitOps (ArgoCD/Flux + progressive delivery), IaC (Terraform + CDK/Pulumi/Crossplane), OpenTelemetry, eBPF-based observability. Can name the trade-off of every tool they propose.
- Architecture communication — writes ADRs and strategy docs that survive review and drive decisions. Can frame a problem before proposing a solution.
- Influence and ambiguity — turns fuzzy org pain into a funded initiative, aligns cross-functional stakeholders, sequences work, says what to cut.
- Mentorship — has grown at least one engineer to the next level; can calibrate the hiring bar; sponsors others growth.
Nice to have (differentiators)
- AI/ML infrastructure — GPU scheduling, model serving, LLMOps, vector DB ops, inference cost optimization. Strong 2026 differentiator.
- Internal Developer Platform leadership — Backstage/Humanitec-style portals, golden-path programs, DX metrics, platform-team org model.
- Multi-cloud honesty — can articulate when multi-cloud is the right answer and when it's cargo cult.Regulator-grade DR — RTO/RPO discipline, audited game days, compliance automation that survives an audit.
- Open-source leadership — maintainer, frequent contributor, or recognized voice in cloud-native (talks, sigs, the CNCF).
- Hiring committee / calibration experience — has shaped a hiring bar at a previous org.
About Kawan Lama Group
Established in 1955, Kawan Lama Group is a multi-sector group of companies who are constantly innovating for improving the quality of lives. Manages 28 brand portfolios operating in six different sectors: Commercial & Industrial, Consumer Retail, Food & Beverages, Property & Hospitality, Manufacturing & Engineering, and Commercial Technology. Aiming to be more than family business - but beyond that, we are business for families, we carry the mission to bring values for betterment of lives through business development and continuous growth.
More Info
Key Skills
CRDs
unit economics
GitOps
compliance automation
IaC
multi-cluster operators
Cilium
OpenTelemetry
error budgets
DR exercises
eBPF-based observability
Crossplane
showback
policy service mesh
Linkerd
ArgoCD
Zero Trust
SLO program
