Cloud & Infrastructure Engineer
On-site | Full-timeAbout Us
PayMedia has evolved into a fully-fledged FinTech solutions provider with strong capabilities across digital payments, software engineering, and financial technology innovation. With a team of over 160+ skilled professionals, we deliver end-to-end solutions that meet the highest standards of security, compliance, and performance.
Job Summary
We are seeking a skilled Cloud & Infrastructure Engineer to design, deploy and manage a Kubernetes-based cloud platform powering our FinTech product suite. The role is responsible for infrastructure reliability, autoscaling, CI/CD pipelines and observability.
Key Responsibilities
- Design, implement, and maintain Kubernetes clusters (GKE / EKS / AKS) for a multi-tier FinTech microservices platform.
- Architect and tune Horizontal Pod Autoscaler (HPA), Vertical Pod Autoscaler (VPA), and KEDA-based scaling policies across Heavy, Large, Medium, and Small service tiers.
- Define and enforce pod resource requests and limits (CPU/RAM) in alignment with load-test benchmarks (e.g., BFF: 2000m/4Gi, Auth: 1500m/3Gi, Notification: 500m/1Gi).
- Conduct capacity planning and RPS-based pod replica calculations to sustain peak TPS across all services (BFF, Auth, Payment, Integration, Biller, Account, Admin, Notification, and more).
- Build and maintain CI/CD pipelines (GitLab CI / GitHub Actions / ArgoCD) with Helm chart management and GitOps delivery patterns.
- Provision and manage cloud infrastructure using Terraform / Terragrunt across GCP, AWS, and/or Azure environments.
- Implement and maintain observability stacks (Prometheus, Grafana, Alertmanager) with autoscaling trigger metrics and SLO-based alerting.
- Manage Kubernetes node pools, cluster upgrades, storage class migrations, and stateful workload operations (Kafka, Redis, PostgreSQL) with zero-downtime.
- Configure and maintain API traffic routing via Nginx, APISIX, or Ingress controllers including rate-limiting, TLS termination, and load balancing.
- Embed DevSecOps practices: image scanning (Trivy), IaC scanning (tfsec/Checkov), secret scanning, and SAST gate integration.
- Lead incident response (P1–P4), root cause analysis, and post-incident reviews for platform availability.
- Collaborate with development, security, and QA teams on architecture decisions, environment strategy (Dev/UAT/Production), and release readiness.
Required Qualifications & Skills
Special Skills Required — Kubernetes Pod Scaling
This role requires proven hands-on expertise in designing and operating pod scaling strategies for high throughput microservices platforms. The candidate must be able to:
- Traffic-based capacity modelling: Translate peak TPS targets (e.g., 400 user req/sec split 60/25/15 across Login, Payment, and Biller flows) into per-service RPS loads and minimum/maximum pod replica counts.
- Service-tier resource sizing: Correctly assign CPU and memory requests/limits per service tier — Heavy (2000m CPU / 4Gi RAM), Large (1500m / 3Gi), Medium (1000m / 2Gi), Small (500m / 1Gi) — based on workload classification (encryption overhead, API fan-out, notification throughput, etc.).
- HPA & KEDA configuration: Configure HPA rules driven by CPU utilisation, custom Prometheus metrics (RPS), or event-driven queues using KEDA; set appropriate cooldown periods and stabilisation windows to prevent thrashing.
- API fan-out awareness: Understand how a single user action maps to multiple downstream API calls across BFF, Auth, Account, Integration, Payment, Biller, and Notification services; factor this multiplier into per service pod sizing.
- Load test integration: Design and interpret load tests (k6, JMeter, Gatling) to validate pod-level RPS ceilings and inform HPA thresholds before production rollout.
- Stateful & persistent workload scaling: Manage scaling considerations for stateful services (Kafka, Redis, PostgreSQL) including persistent volume provisioning, storage class selection, and PodDisruptionBudget configuration.
- Zero-downtime deployments: Configure rolling update strategies, pod disruption budgets, and readiness/liveness probes to ensure uninterrupted service during peak traffic and deployment windows.
- Cluster autoscaling: Configure node pool autoscaling (GKE Cluster Autoscaler / Karpenter on EKS) to provision/deprovision nodes dynamically in response to pending pod scheduling pressure.
Required Technical Skills
- Kubernetes & Containers: Kubernetes (GKE / EKS / AKS), Docker, Helm, HPA, VPA, KEDA, pod security, node pools, network policies, persistent volumes.
- Cloud Platforms: Google Cloud Platform (GKE, GCS, IAM, Cloud DNS, Load Balancing) — primary; AWS (EKS, EC2, S3, RDS, IAM) or Azure (AKS, VNet, Key Vault) as additional platforms.
- CI/CD & GitOps: GitLab CI, GitHub Actions, ArgoCD, Helm chart management, environment promotion strategies (Dev → UAT → Prod).
- Infrastructure as Code: Terraform and/or Terragrunt; reusable module design; state management; IaC security scanning (tfsec / Checkov).
- Observability: Prometheus, Grafana, Alertmanager; custom metrics for autoscaling; SLO/SLA dashboards; log aggregation (ELK / Cloud Logging).
- Networking & Security: Nginx / APISIX / Ingress controllers; TLS/SSL lifecycle management (Cert-Manager); WAF; VPC/VNet segmentation; least-privilege IAM/RBAC; secrets management (Vault / AWS KMS / GCP Secret Manager).
- Scripting & Automation: Bash, Python; operational automation scripts for health checks, alert triage, compliance reporting.
- Databases & Messaging: PostgreSQL, MongoDB, Redis, Kafka — operational awareness for stateful workload scaling and high availability.
Nice To Have
- Experience in FinTech, banking, or payment processing platforms.
- Hands-on with Keycloak or similar identity/SSO platforms.
- Knowledge of PCI DSS infrastructure control requirements.
- Experience with Apache Kafka, EMQX, or Zookeeper cluster operations.
- Familiarity with KubeEdge or edge computing workloads.
- DevSecOps tooling: SonarQube (SAST), Trivy, Gitleaks, OWASP Dependency-Check.
- Certifications: CKA / CKAD, AWS Solutions Architect, GCP Professional Cloud Architect, or Azure Administrator (AZ-104).
Qualifications & Experience
- Bachelor's degree in Computer Science, Information Technology, Software Engineering, or equivalent.
- 3 – 6 years of hands-on experience in cloud infrastructure and/or DevOps engineering.
- At least 2 years of production-level Kubernetes experience including pod autoscaling.
- Demonstrated experience managing CI/CD pipelines for multi-environment (Dev/UAT/Prod) deployments.
- Strong written and verbal communication skills for cross-functional collaboration and incident documentation.