Cutting AI Infrastructure Costs by 83%: A Healthtech Migration Story

Replacing always-on Celery workers with Kubernetes Jobs on EKS

VirtueCloud helped a healthtech startup cut production AI infrastructure compute costs by ~83% (saving ~$2,500/month) and double peak processing capacity by replacing always-on Celery instances with on-demand Kubernetes Jobs and Karpenter autoscaling on Amazon EKS.

6 min read
Cutting AI Infrastructure Costs by 83%: A Healthtech Migration Story

Challenge

A healthtech startup came to us with an AI-powered product that ran heavy processing jobs in the background. Each job took about 30 minutes and was handled by Celery workers written in Python. To keep those workers available, the team ran 4 always-on instances in their production AWS account, costing roughly $3,000 per month. Most of the day those machines sat idle, waiting for jobs that arrived in bursts. For an early-stage startup, that was a significant share of the cloud budget spent on unused capacity. Scaling was the second problem. More load meant more workers, which meant more always-on machines and a higher baseline bill. Why Celery was the wrong fit: Celery is excellent for short, frequent tasks. This workload was the opposite: infrequent, heavy, 30-minute jobs. - Paying for idle time: Workers must be running to accept work, so the client was billed even with an empty queue. - Coupled scaling: Capacity was tied to the instances hosting the workers, so growth meant provisioning more machines. - Shared resources: Long jobs competed for CPU and memory on the same worker.

Solution

1

Moved the background AI workloads to Kubernetes Jobs on the client's existing Amazon EKS cluster, with Karpenter handling dynamic node autoscaling so compute exists only while a job is running.

2

Designed the application's pod to create and monitor Jobs directly through the Kubernetes API using the official Kubernetes Python SDK, avoiding separate queues or external schedulers.

3

Hardened the architecture with a scoped ServiceAccount and namespaced Role granting least-privilege permissions solely to manage batch Jobs, protecting the rest of the cluster.

4

Replaced Celery worker mechanisms with native Kubernetes Job settings including backoff_limit for retries, active_deadline_seconds for timeouts, and ttl_seconds_after_finished for automated resource cleanup.

5

Replaced fixed node groups with Karpenter just-in-time autoscaling, allowing cluster capacity to flex from zero up to 8 instances on demand and deprovisioning empty nodes immediately upon completion.

Core Architecture & Components

Container Orchestration (Amazon EKS)

  • •Amazon EKS cluster hosting core application pods and batch AI Jobs
  • •Workloads isolated within secure private subnets
  • •Dedicated CPU (2 vCPU) and memory (8Gi RAM) quotas per worker Job
  • •Native Kubernetes Batch API handling job scheduling and lifecycle

Dynamic Node Autoscaling (Karpenter)

  • •High-performance, just-in-time EC2 node provisioner
  • •Rapid node provisioning (~60s) upon detection of pending pods
  • •Elastic scaling from 0 to 8 instances based on burst demand
  • •Automatic node termination when jobs finish and pods exit

In-Cluster Application Job Launcher

  • •Application pod dispatches Jobs directly via the Kubernetes Python SDK
  • •In-cluster authentication using mounted ServiceAccount tokens
  • •Automated status tracking via batch_v1.read_namespaced_job_status
  • •Zero static credentials or external kubeconfig files to manage

Kubernetes RBAC & Least-Privilege Security

  • •Dedicated ServiceAccount (app-job-launcher) assigned to application pod
  • •Scoped Role (job-manager) restricting permissions exclusively to batch/jobs
  • •Namespace isolation ensuring no access to other cluster resources
  • •Aligns with stringent healthcare data protection standards

Application Stack

LayerTechnology
Container OrchestrationAmazon EKS
Node AutoscalingKarpenter
Workload ExecutionKubernetes Batch Jobs (v1)
Application SDKKubernetes Python Client (In-Cluster Config)
Security & Access ControlKubernetes RBAC (ServiceAccount, Role, RoleBinding)
Cloud InfrastructureAWS (EC2 On-Demand Compute)

Smart Workflow Automation

1. User Action Triggers AI Workload

Application Request → Job Handler

A user action within the healthtech application requires an intensive 30-minute AI batch processing task.

2. Application Creates Kubernetes Job

Python SDK → In-Cluster K8s API

The application pod uses the Kubernetes Python client with mounted ServiceAccount credentials to submit an on-demand batch Job specification.

3. Karpenter Just-In-Time Provisioning

Unschedulable Pod → Karpenter → EC2 Node (~60s)

If the existing cluster capacity cannot accommodate the job, Karpenter immediately detects the unschedulable pod and provisions a right-sized EC2 instance.

4. Isolated Execution with Dedicated Resources

Container Execution (2 vCPU, 8Gi RAM)

The AI worker container executes with dedicated resources, guaranteeing zero resource contention or noisy-neighbor interference.

5. Automated Scale Down to Zero

Job Success → TTL Cleanup (10m) → Node Deprovisioning

Once the job finishes, the pod terminates. The Kubernetes TTL controller cleans up the job after 10 minutes, and Karpenter terminates the idle node. No jobs means zero nodes.

Implementation Details

rbac-job-manager.yamlyaml
Scoped ServiceAccount and Role configuration enforcing least-privilege access for the application's job-launcher pod.
apiVersion: v1
kind: ServiceAccount
metadata:
  name: app-job-launcher
  namespace: app
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: job-manager
  namespace: app
rules:
  - apiGroups: ["batch"]
    resources: ["jobs"]
    verbs: ["create", "get", "list", "watch", "delete"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: app-job-launcher-binding
  namespace: app
subjects:
  - kind: ServiceAccount
    name: app-job-launcher
    namespace: app
roleRef:
  kind: Role
  name: job-manager
  apiGroup: rbac.authorization.k8s.io
trigger_ai_job.pypython
Python application handler using in-cluster configuration to create Jobs with native retries, timeouts, and automated TTL cleanup.
from kubernetes import client, config

# Loads the ServiceAccount token mounted into the pod
config.load_incluster_config()
batch_v1 = client.BatchV1Api()


def trigger_ai_job(job_input: str, namespace: str = "app") -> str:
    container = client.V1Container(
        name="ai-worker",
        image="your-registry/ai-worker:latest",
        env=[client.V1EnvVar(name="JOB_INPUT", value=job_input)],
        resources=client.V1ResourceRequirements(
            requests={"cpu": "2", "memory": "8Gi"},
            limits={"memory": "8Gi"},
        ),
    )

    job = client.V1Job(
        metadata=client.V1ObjectMeta(generate_name="ai-job-"),
        spec=client.V1JobSpec(
            backoff_limit=2,                  # retry failed jobs up to 2 times
            active_deadline_seconds=3600,     # kill jobs that run longer than 1 hour
            ttl_seconds_after_finished=600,   # clean up finished jobs after 10 minutes
            template=client.V1PodTemplateSpec(
                spec=client.V1PodSpec(
                    restart_policy="Never",
                    containers=[container],
                )
            ),
        ),
    )

    created = batch_v1.create_namespaced_job(namespace=namespace, body=job)
    return created.metadata.name

Results & Impact Comparison

Migrating from always-on Celery instances to on-demand Kubernetes Jobs cut compute expenditure by ~83% while simultaneously doubling peak system throughput.

MetricBefore (Celery)After (K8s Jobs + Karpenter)Impact
Compute model4 always-on instances0 to 8 instances, on demandTrue scale-to-zero compute
Production compute cost~$3,000 / month~$500 / month~$2,500 / month saved (~83%)
Peak capacity4 instances (fixed ceiling)8 instances (elastic)2x concurrency capacity
Startup overheadNone (always warm)15s to ~75s per job< 4% of 30-minute run time

The Cold Start Trade-Off Analysis

In an on-demand compute model, nodes only launch when needed. For 30-minute AI workloads, the maximum ~75 second startup represents negligible overhead.

ScenarioMeasured Startup TimeOverhead on 30m Job
Space available on an existing node~15 seconds< 1% of total job duration
New node required (Karpenter provisioned)~1 minute (node) + ~15 seconds (pod)~4% of total job duration

Infrastructure Optimization

Cost Optimization

  • Cut monthly production compute bills by ~83% (~$2,500/month saved)
  • Dropped cloud compute costs from ~$3,000 to ~$500 monthly
  • True zero compute spend when no background AI jobs are in progress
  • Replaced over-provisioned always-on instances with just-in-time compute

Performance & Scalability

  • Elastic autoscaling from 0 to 8 instances on demand
  • Doubled peak capacity compared to the previous 4-instance ceiling
  • Negligible ~4% cold-start overhead for 30-minute batch tasks
  • Dedicated CPU and memory quotas preventing worker resource contention

Security & Governance

  • Scoped ServiceAccount with namespaced Role granting batch/jobs permissions only
  • In-cluster authentication eliminating static secrets or long-lived keys
  • Aligns with healthcare data protection and least-privilege standards
  • Automated TTL cleanup preventing orphaned resource sprawl

Key Solutions

Celery to Kubernetes Jobs Migration:

Replaced continuously running Celery worker instances with native Kubernetes Jobs that execute on demand and terminate upon completion, eliminating idle costs.

Karpenter Dynamic Node Autoscaling:

Configured Karpenter to watch for unschedulable pods, provision right-sized nodes in ~60 seconds, and decommission idle nodes automatically when jobs finish.

In-Cluster Scoped Job Dispatch:

Enabled the application pod to directly dispatch Jobs via the Kubernetes Python SDK using a least-privilege ServiceAccount and Role without external queue brokers.

Built-in Job Lifecycle Controls:

Leveraged native Kubernetes Job parameters for automated retries (backoff_limit), execution timeouts (active_deadline_seconds), and post-run cleanup (ttl_seconds_after_finished).

Objectives & Key Results

Objective 1: Eliminate idle cloud compute costs for background AI processing

01

Slashed production compute costs by ~83% (~$2,500/month saved)

02

Reduced monthly baseline bill from ~$3,000 to ~$500

03

Achieved true scale-to-zero during off-peak periods

Objective 2: Modernize scaling to handle bursty client workloads

01

Doubled peak concurrency capacity from 4 to 8 instances

02

Startup overhead maintained below 4% for 30-minute jobs

03

Zero interference between concurrent batch processing jobs

Objective 3: Maintain least-privilege security and operational simplicity

01

Enforced scoped RBAC ensuring the launcher pod can only manage batch Jobs

02

Eliminated external Celery broker and worker daemon management overhead

Business Impact

~83%
Production compute cost reduction
~$2,500/mo
Direct cloud cost savings
2x
Peak job concurrency capacity
0
Idle instances running during lulls
< 4%
Startup overhead on 30-min jobs

Key Takeaways & Considerations

Key Takeaways

  • •For long-running, bursty AI workloads, always-on workers mean paying for idle time.
  • •Kubernetes Jobs plus Karpenter provide compute that scales with demand, down to zero.
  • •Letting the application pod launch Jobs through a scoped ServiceAccount and Role keeps the architecture simple and secure.
  • •Cold starts matter far less when jobs run for tens of minutes (~4% overhead on 30-minute workloads).

When This Approach Isn't the Right Fit

  • •Short, high-frequency tasks: If tasks take seconds, a 15-second pod startup dominates. Celery or persistent workers remain better suited.
  • •Real-time, latency-sensitive work: Cold starts may be unacceptable unless warm capacity is maintained.
  • •Added operational surface: Job lifecycle, cleanup, and monitoring now reside in Kubernetes rather than Celery tooling.

Project Outcome

VirtueCloud successfully migrated the healthtech startup from an expensive, 4-instance always-on Celery infrastructure to an on-demand Kubernetes Jobs model powered by Karpenter on Amazon EKS. The migration slashed monthly production compute costs by ~83% (saving ~$2,500 every month) while doubling peak concurrent capacity from 4 to 8 instances. By replacing always-on worker instances with on-demand Jobs and Karpenter just-in-time autoscaling, compute now exists only while workloads are executing. The streamlined architecture eliminated external queue broker maintenance and enforced strict least-privilege RBAC security, establishing a cost-effective, scalable foundation for the startup's continued AI product innovation.

Future Roadmap

Planned future optimizations include configuring Karpenter to incorporate AWS Spot instances for fault-tolerant jobs to reduce costs by an additional 30–50%, creating specialized Karpenter node pools tailored to varying AI model memory footprints, and integrating automated Prometheus and Grafana dashboards for granular per-job cost attribution.