Home / Case Studies / Platform engineering
Platform engineering

One blueprint for every environment of an AppSec SaaS.

We moved an application-security vendor from hand-maintained Terraform to Terragrunt with one AWS account per environment, Argo CD delivery, and autoscaling built for bursty scanning workloads.

Application security SaaS · Client name withheld under NDA

85

Services: delivered by Argo CD from charts and values in Git.

94

Scanner tools: scaled from queue depth by KEDA.

11

Karpenter node pools: one per workload role, spot where safe.

Amazon EKSTerragruntArgo CDHelmKarpenterKEDAAmazon FSx for OpenZFSExternal SecretsDatadogKyverno

The challenge

The platform scans customers' code and cloud for security issues. Each environment was a separate plain-Terraform deployment, development and staging shared an AWS account, and releases ran from CI scripts with promotions done by automated merge requests.

The workload is unusual: 94 scanner tools, about 70 of them language-specific variants, that sit idle and then all start at once when customers connect repositories.

The constraints

  • Customers depend on the SaaS around the clock, so changes had to land without downtime.
  • The same platform also had to install inside customers' own environments.
  • Persistent storage ran on self-managed Rook-Ceph, which needed manual repair runbooks.

Decisions & tradeoffs

  • Terragrunt with a strict hierarchy: account, environment, region, unit. Every unit has its own state and lock, and each environment lives in its own AWS account to contain the blast radius.
  • Terraform installs Argo CD and hands it the cluster's facts, so add-ons are configured from infrastructure outputs instead of copy-pasted values.
  • Charts and values live in separate repositories and meet in multi-source Argo CD applications.
  • Karpenter with 11 node pools, one per workload role (scanners, workers, control, add-ons), tainted so each workload lands on the right capacity; spot for workers, on-demand only for add-ons.
  • KEDA scales scan jobs from Redis queue depth instead of keeping workers warm.
  • Shared storage moved from Rook-Ceph to Amazon FSx for OpenZFS through its CSI driver, removing a storage cluster the team had to operate.

The implementation

Infrastructure units cover the VPC, an optional NAT instance to cut data-transfer cost, Secrets Manager, and EKS with access entries; a second layer installs metrics-server, ingress, the AWS Load Balancer Controller, Argo CD, the FSx CSI driver, and Karpenter.

Secrets are created per environment in AWS Secrets Manager and synced into the cluster by External Secrets. A Kyverno policy injects corporate CA bundles into the product's pods, for customers whose networks inspect TLS. Developers can also spin up an ephemeral environment per pull request.

GitLab CI runs Terragrunt to build each AWS account's VPC, EKS, and add-ons; Argo CD syncs 85 services from Git; Karpenter and KEDA scale nodes and scan jobs; Secrets Manager feeds External Secrets; FSx for OpenZFS provides shared storage; Datadog monitors.
AppSec SaaS platform. Highlighted: provisioning and delivery paths.

Outcomes

Each environment is isolated in its own account, a new one is a set of files and an apply, and scanning capacity follows the work instead of being provisioned for the peak. The team no longer runs its own storage cluster for the SaaS.

Handover & ongoing ownership

Runbooks for the Terragrunt layout, node pools, and storage, with the platform maintained by the vendor's DevOps team.

Planning something similar? Talk to an AWS partner in Armenia that has built it before.

What’s next for
your business?

Let’s talk