Cutting AI Infrastructure Costs by 83%: A Healthtech Migration Story
Replacing always-on Celery workers with Kubernetes Jobs on EKS
VirtueCloud helped a healthtech startup cut production AI infrastructure compute costs by ~83% (saving ~$2,500/month) and double peak processing capacity by replacing always-on Celery instances with on-demand Kubernetes Jobs and Karpenter autoscaling on Amazon EKS.

Challenge
Solution
Moved the background AI workloads to Kubernetes Jobs on the client's existing Amazon EKS cluster, with Karpenter handling dynamic node autoscaling so compute exists only while a job is running.
Designed the application's pod to create and monitor Jobs directly through the Kubernetes API using the official Kubernetes Python SDK, avoiding separate queues or external schedulers.
Hardened the architecture with a scoped ServiceAccount and namespaced Role granting least-privilege permissions solely to manage batch Jobs, protecting the rest of the cluster.
Replaced Celery worker mechanisms with native Kubernetes Job settings including backoff_limit for retries, active_deadline_seconds for timeouts, and ttl_seconds_after_finished for automated resource cleanup.
Replaced fixed node groups with Karpenter just-in-time autoscaling, allowing cluster capacity to flex from zero up to 8 instances on demand and deprovisioning empty nodes immediately upon completion.
Core Architecture & Components
Container Orchestration (Amazon EKS)
- •Amazon EKS cluster hosting core application pods and batch AI Jobs
- •Workloads isolated within secure private subnets
- •Dedicated CPU (2 vCPU) and memory (8Gi RAM) quotas per worker Job
- •Native Kubernetes Batch API handling job scheduling and lifecycle
Dynamic Node Autoscaling (Karpenter)
- •High-performance, just-in-time EC2 node provisioner
- •Rapid node provisioning (~60s) upon detection of pending pods
- •Elastic scaling from 0 to 8 instances based on burst demand
- •Automatic node termination when jobs finish and pods exit
In-Cluster Application Job Launcher
- •Application pod dispatches Jobs directly via the Kubernetes Python SDK
- •In-cluster authentication using mounted ServiceAccount tokens
- •Automated status tracking via batch_v1.read_namespaced_job_status
- •Zero static credentials or external kubeconfig files to manage
Kubernetes RBAC & Least-Privilege Security
- •Dedicated ServiceAccount (app-job-launcher) assigned to application pod
- •Scoped Role (job-manager) restricting permissions exclusively to batch/jobs
- •Namespace isolation ensuring no access to other cluster resources
- •Aligns with stringent healthcare data protection standards
Application Stack
| Layer | Technology |
|---|---|
| Container Orchestration | Amazon EKS |
| Node Autoscaling | Karpenter |
| Workload Execution | Kubernetes Batch Jobs (v1) |
| Application SDK | Kubernetes Python Client (In-Cluster Config) |
| Security & Access Control | Kubernetes RBAC (ServiceAccount, Role, RoleBinding) |
| Cloud Infrastructure | AWS (EC2 On-Demand Compute) |
Smart Workflow Automation
1. User Action Triggers AI Workload
A user action within the healthtech application requires an intensive 30-minute AI batch processing task.
2. Application Creates Kubernetes Job
The application pod uses the Kubernetes Python client with mounted ServiceAccount credentials to submit an on-demand batch Job specification.
3. Karpenter Just-In-Time Provisioning
If the existing cluster capacity cannot accommodate the job, Karpenter immediately detects the unschedulable pod and provisions a right-sized EC2 instance.
4. Isolated Execution with Dedicated Resources
The AI worker container executes with dedicated resources, guaranteeing zero resource contention or noisy-neighbor interference.
5. Automated Scale Down to Zero
Once the job finishes, the pod terminates. The Kubernetes TTL controller cleans up the job after 10 minutes, and Karpenter terminates the idle node. No jobs means zero nodes.
Implementation Details
apiVersion: v1
kind: ServiceAccount
metadata:
name: app-job-launcher
namespace: app
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: job-manager
namespace: app
rules:
- apiGroups: ["batch"]
resources: ["jobs"]
verbs: ["create", "get", "list", "watch", "delete"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: app-job-launcher-binding
namespace: app
subjects:
- kind: ServiceAccount
name: app-job-launcher
namespace: app
roleRef:
kind: Role
name: job-manager
apiGroup: rbac.authorization.k8s.iofrom kubernetes import client, config
# Loads the ServiceAccount token mounted into the pod
config.load_incluster_config()
batch_v1 = client.BatchV1Api()
def trigger_ai_job(job_input: str, namespace: str = "app") -> str:
container = client.V1Container(
name="ai-worker",
image="your-registry/ai-worker:latest",
env=[client.V1EnvVar(name="JOB_INPUT", value=job_input)],
resources=client.V1ResourceRequirements(
requests={"cpu": "2", "memory": "8Gi"},
limits={"memory": "8Gi"},
),
)
job = client.V1Job(
metadata=client.V1ObjectMeta(generate_name="ai-job-"),
spec=client.V1JobSpec(
backoff_limit=2, # retry failed jobs up to 2 times
active_deadline_seconds=3600, # kill jobs that run longer than 1 hour
ttl_seconds_after_finished=600, # clean up finished jobs after 10 minutes
template=client.V1PodTemplateSpec(
spec=client.V1PodSpec(
restart_policy="Never",
containers=[container],
)
),
),
)
created = batch_v1.create_namespaced_job(namespace=namespace, body=job)
return created.metadata.nameResults & Impact Comparison
Migrating from always-on Celery instances to on-demand Kubernetes Jobs cut compute expenditure by ~83% while simultaneously doubling peak system throughput.
| Metric | Before (Celery) | After (K8s Jobs + Karpenter) | Impact |
|---|---|---|---|
| Compute model | 4 always-on instances | 0 to 8 instances, on demand | True scale-to-zero compute |
| Production compute cost | ~$3,000 / month | ~$500 / month | ~$2,500 / month saved (~83%) |
| Peak capacity | 4 instances (fixed ceiling) | 8 instances (elastic) | 2x concurrency capacity |
| Startup overhead | None (always warm) | 15s to ~75s per job | < 4% of 30-minute run time |
The Cold Start Trade-Off Analysis
In an on-demand compute model, nodes only launch when needed. For 30-minute AI workloads, the maximum ~75 second startup represents negligible overhead.
| Scenario | Measured Startup Time | Overhead on 30m Job |
|---|---|---|
| Space available on an existing node | ~15 seconds | < 1% of total job duration |
| New node required (Karpenter provisioned) | ~1 minute (node) + ~15 seconds (pod) | ~4% of total job duration |
Infrastructure Optimization
Cost Optimization
- Cut monthly production compute bills by ~83% (~$2,500/month saved)
- Dropped cloud compute costs from ~$3,000 to ~$500 monthly
- True zero compute spend when no background AI jobs are in progress
- Replaced over-provisioned always-on instances with just-in-time compute
Performance & Scalability
- Elastic autoscaling from 0 to 8 instances on demand
- Doubled peak capacity compared to the previous 4-instance ceiling
- Negligible ~4% cold-start overhead for 30-minute batch tasks
- Dedicated CPU and memory quotas preventing worker resource contention
Security & Governance
- Scoped ServiceAccount with namespaced Role granting batch/jobs permissions only
- In-cluster authentication eliminating static secrets or long-lived keys
- Aligns with healthcare data protection and least-privilege standards
- Automated TTL cleanup preventing orphaned resource sprawl
Key Solutions
Celery to Kubernetes Jobs Migration:
Replaced continuously running Celery worker instances with native Kubernetes Jobs that execute on demand and terminate upon completion, eliminating idle costs.
Karpenter Dynamic Node Autoscaling:
Configured Karpenter to watch for unschedulable pods, provision right-sized nodes in ~60 seconds, and decommission idle nodes automatically when jobs finish.
In-Cluster Scoped Job Dispatch:
Enabled the application pod to directly dispatch Jobs via the Kubernetes Python SDK using a least-privilege ServiceAccount and Role without external queue brokers.
Built-in Job Lifecycle Controls:
Leveraged native Kubernetes Job parameters for automated retries (backoff_limit), execution timeouts (active_deadline_seconds), and post-run cleanup (ttl_seconds_after_finished).
Objectives & Key Results
Objective 1: Eliminate idle cloud compute costs for background AI processing
Slashed production compute costs by ~83% (~$2,500/month saved)
Reduced monthly baseline bill from ~$3,000 to ~$500
Achieved true scale-to-zero during off-peak periods
Objective 2: Modernize scaling to handle bursty client workloads
Doubled peak concurrency capacity from 4 to 8 instances
Startup overhead maintained below 4% for 30-minute jobs
Zero interference between concurrent batch processing jobs
Objective 3: Maintain least-privilege security and operational simplicity
Enforced scoped RBAC ensuring the launcher pod can only manage batch Jobs
Eliminated external Celery broker and worker daemon management overhead
Business Impact
Key Takeaways & Considerations
Key Takeaways
- •For long-running, bursty AI workloads, always-on workers mean paying for idle time.
- •Kubernetes Jobs plus Karpenter provide compute that scales with demand, down to zero.
- •Letting the application pod launch Jobs through a scoped ServiceAccount and Role keeps the architecture simple and secure.
- •Cold starts matter far less when jobs run for tens of minutes (~4% overhead on 30-minute workloads).
When This Approach Isn't the Right Fit
- •Short, high-frequency tasks: If tasks take seconds, a 15-second pod startup dominates. Celery or persistent workers remain better suited.
- •Real-time, latency-sensitive work: Cold starts may be unacceptable unless warm capacity is maintained.
- •Added operational surface: Job lifecycle, cleanup, and monitoring now reside in Kubernetes rather than Celery tooling.
Project Outcome
VirtueCloud successfully migrated the healthtech startup from an expensive, 4-instance always-on Celery infrastructure to an on-demand Kubernetes Jobs model powered by Karpenter on Amazon EKS. The migration slashed monthly production compute costs by ~83% (saving ~$2,500 every month) while doubling peak concurrent capacity from 4 to 8 instances. By replacing always-on worker instances with on-demand Jobs and Karpenter just-in-time autoscaling, compute now exists only while workloads are executing. The streamlined architecture eliminated external queue broker maintenance and enforced strict least-privilege RBAC security, establishing a cost-effective, scalable foundation for the startup's continued AI product innovation.
Future Roadmap
Planned future optimizations include configuring Karpenter to incorporate AWS Spot instances for fault-tolerant jobs to reduce costs by an additional 30–50%, creating specialized Karpenter node pools tailored to varying AI model memory footprints, and integrating automated Prometheus and Grafana dashboards for granular per-job cost attribution.