Home / Case Studies / Platform and software engineering
Platform and software engineering

An online gaming platform on AWS, architected and built end to end.

We architected the system and the software and built the engineering processes for an online gaming operator: from an on-premises monolith to a secure, cost-optimized, event-driven platform on Amazon EKS with GitOps delivery from commit to production.

Online gaming · Client name withheld under NDA

Amazon EKSAmazon AuroraAmazon ElastiCacheAmazon SQSTerragruntKarpenterKEDAArgo CDKargoCloudflare

The challenge

When we started, the platform was one monolith on self-managed servers on the operator's own premises, with one database underneath every feature. Deploys and scaling were manual and nothing was defined as code.

The product was growing fast, every feature moved real money, and traffic arrived in sharp peaks. The monolith could not scale the parts under load independently, and every release risked the whole platform.

Before: one monolith with every feature in one codebase on self-managed on-premises servers behind a load balancer, with one shared database; deploys and scaling were manual.
Before: one monolith, one database, manual operations.

The constraints

  • Real money: every balance change must be traceable, and a payout must never go out by mistake.
  • Many external integrations, each with its own failure modes.
  • Real-time features and sharp traffic peaks.
  • A public target for bots and floods.
  • A small team that needed to release safely every day.

Decisions & tradeoffs

Event-driven services. We split the monolith along domain boundaries. Services that answer users scale on CPU and requests; work that can wait goes onto queues and is processed by workers that KEDA scales from zero with the queue depth. Events leave each service through a transactional outbox, so nothing is lost and nothing is processed twice.

Money safety by design. Every movement of money is a posting in a double-entry ledger, every money API is idempotent, and automatic payouts run through a chain of deterministic checks that fails closed: anything unexpected goes to a person. An AI reviewer can advise, but never move money.

Secure by design. Edge protection with a WAF, rate limits and bot rules; workloads only in private subnets; engineers through SSO, Zero Trust and VPN; organization-wide CloudTrail, GuardDuty and Security Hub; signed and scanned images; secrets in AWS Secrets Manager with rotation and one least-privilege role per service.

Cost-optimized by design. Workers scale to zero, Karpenter right-sizes and consolidates nodes, stateless work runs on Spot and Graviton, Savings Plans cover the steady baseline, read replicas take read traffic, and bots are stopped at the edge instead of being paid for.

The implementation

Migration. We moved the platform from its on-premises servers to Amazon EKS and its data to Amazon Aurora, switching production over in a planned window.

Accounts and infrastructure as code. Terragrunt defines an AWS Organization with separate management, development and production accounts, each production network spread across three Availability Zones. Services get their cloud resources and permissions from their own deployment definitions.

Delivery. Every change goes from Git through CI, tests and image scanning to a registry. Kargo promotes it from development to production by committing to the GitOps repository, and Argo CD applies it to each cluster.

Observability and AI operations. Monitoring runs centrally in the management account: metrics, logs and traces from every account, with alerts routed to the team that owns them. AI agents run beside it to triage alerts and summarize incidents, and they can investigate but never change production.

GitHubRepositoriesCI · test · scanGitOps repositoryAWS Cloud · OrganizationManagementContainer registrybuild · pushEKS · platformMonitoringAI agentsKargopromoteSecurity & governanceIdentity CenterCloudTrailGuardDutySecurity HubAWS BackupDevelopmentVPCEKSArgo CDsyncWorkloadsDatabasesProductionVPC · 3 AZsPublic subnetsLoad balancerPrivate subnetsEKSKarpenterArgo CDServicessyncData subnetsAuroraElastiCacheUsersEdge protectionHTTPSThird partiesDocument DBUser trafficDelivery (GitOps)Platform calls
Bird’s-eye view of the platform. Green: player traffic. Dashed: delivery from commit to production.
Private subnets · Amazon EKSAPI servicesUsersEdge protectionHTTPSPods scale on CPU and requestspublishQueue and event busAmazon SQS + SNS · dead-letter queuesconsumeKEDAdepthscale 0…NWorkersScale from zero with the queuepending podsNodesKarpenterprovisionOn-demandSpotGravitonconsolidatedwhen idleManaged AWS data servicesAurorawriter + read replicasElastiCachecache · sessionsS3object storagereads and writescachefilesGitOps deliveryGitCI · test · scanRegistryKargo · promoteArgo CD · syncdev → staging → prodSecure by designEdge WAF, rate limits and bot protectionWorkloads only in private subnets; no public admin endpointsSSO, Zero Trust, CloudTrail, GuardDuty and Security HubSigned, scanned images; secrets rotated; least-privilege IAMCost-optimized by designWorkers scale to zero when queues are emptyKarpenter right-sizes and consolidates nodesSpot and Graviton for stateless work; Savings Plans for the baselineRead replicas and edge filtering keep the bill flatUser trafficCalls and scalingQueues and eventsDelivery
How we build it: event-driven services with pod and node autoscaling, queues, GitOps, security and cost built in.

Outcomes

The operator moved from one on-premises monolith to a platform where each part scales on its own, idle work costs nothing, every release is an auditable promotion, and money can never move by mistake. The platform has absorbed flood attacks without taking players offline.

Handover & ongoing ownership

Runbooks, alert routing and every piece of infrastructure and configuration live in the operator's own repositories.

Planning something similar? Talk to an AWS partner in Armenia that has built it before.

What’s next for
your business?

Let’s talk