Skip to main content
Lilac’s Kubernetes operator lets suppliers make idle GPUs available to the shared inference network without displacing their own workloads. The operator starts inference workloads when eligible capacity is idle and steps aside when the cluster needs those GPUs back.
This program is for idle capacity in Kubernetes clusters. For a direct relationship around current or planned bare-metal capacity, see Bare-metal partnerships.

How it works

1

Keep your workloads first

Your existing jobs retain priority. Lilac uses only the capacity made eligible through your GPU pool configuration.
2

Run inference in idle windows

The operator detects eligible GPUs, starts inference pods, and serves traffic from Lilac’s inference network.
3

Earn per token

You receive 70% of the gross inference revenue processed on your hardware. Lilac retains 30%.
4

Reclaim capacity

When your workloads need the GPUs, the operator drains Lilac inference pods according to the configured preemption policy.

Revenue model

Requirements

  • A Kubernetes cluster with supported GPUs
  • kubectl access to the cluster
  • A Lilac supplier account

Start operator onboarding

Create an account, share your cluster details, and schedule onboarding.