Armada iconArmada text
Armada iconArmada text

Multi-cluster batch scheduler for Kubernetes

Latest releaseGitHub starsArtifact HubLFX Health Score
CNCF SandboxOpenSSF Best PracticesLicense

One API.
Any number of clusters.
Millions of jobs.

Armada is the open-source batch job meta-scheduler that makes Kubernetes handle massive-scale workloads, with fair queuing, gang scheduling, and multi-cluster orchestration built in.


As a CNCF Sandbox project, Armada is actively maintained and used in production environments, including at G-Research where it processes millions of jobs daily.


What is Armada?

The batch scheduler Kubernetes was missing.

Armada sits above your Kubernetes clusters as a control plane. It does not replace Kubernetes, but allows K8 to handle millions of jobs a day across tens of thousands of nodes.

Multi-cluster native

Run jobs across many clusters through one API, and add or remove capacity without disrupting what's already running.

Fair-share scheduling

Every team gets a fair share of resources over time, so heavy users can't crowd everyone else out.

Gang scheduling

All the workers in a job start together or not at all, which is what frameworks like MPI, PyTorch, and Spark need.

Intelligent preemption

Urgent work can preempt lower-priority jobs to run in time, and you decide how that works per queue.

High throughput

Handle millions of queued jobs by moving queueing onto PostgreSQL and Redis instead of leaning on etcd.

Built for production

Prometheus metrics, Lookout web UI, secure auth, and automatic handling of failed nodes, all come built in.


Use Cases

Is Armada right for you?

01
Machine learning training at scale

Your workers need to start together or not at all. Armada's gang scheduling makes sure they do, across however many clusters have the GPU capacity you need.

02
Quantitative research and financial modelling

Millions of short-lived jobs, every day. Armada keeps them moving fairly and fast, with priority controls for the calculations that can't wait.

03
High-performance computing

MPI workloads in containers, scheduled across clusters, with hardware-aware placement. Cloud-native tooling without giving up the reproducibility HPC teams depend on.

04
Multi-tenant compute environments

Multiple teams, one infrastructure. Fair-share scheduling means no single team can quietly consume everything while others wait.

05
CI/CD build and test

Critical merges go first. Large test suites don't block urgent builds. Priority and fairness built in, no manual queue management.

Armada's niche is multi-cluster. If you're not there yet, another project may be a better fit.

See how Armada compares →


Ready to run batch at scale?

Get Armada running locally in minutes. Join the community of organisations running batch workloads on Kubernetes.


Edit on GitHub

Last updated on