This program is for idle capacity in Kubernetes clusters. For a direct relationship around current or planned bare-metal capacity, see Bare-metal partnerships.
How it works
1
Keep your workloads first
Your existing jobs retain priority. Lilac uses only the capacity made eligible through your GPU pool configuration.
2
Run inference in idle windows
The operator detects eligible GPUs, starts inference pods, and serves traffic from Lilac’s inference network.
3
Earn per token
You receive 70% of the gross inference revenue processed on your hardware. Lilac retains 30%.
4
Reclaim capacity
When your workloads need the GPUs, the operator drains Lilac inference pods according to the configured preemption policy.
Revenue model
Requirements
- A Kubernetes cluster with supported GPUs
kubectlaccess to the cluster- A Lilac supplier account
Start operator onboarding
Create an account, share your cluster details, and schedule onboarding.

