# Lilac > Documentation for Lilac inference, subscriptions, dedicated GPUs, and GPU supply partnerships. ## Docs - [Inference Quickstart](https://docs.getlilac.com/inference/quickstart.md): Get started with the Lilac inference API in under five minutes. Create an account, add credits, generate an API key, and send your first request. - [Supported Models](https://docs.getlilac.com/inference/models.md): Browse all models available on Lilac including Kimi K2.6, GLM 5.2, Gemma 4, and MiniMax M3 with context lengths, capabilities, and per-token pricing. - [API Keys](https://docs.getlilac.com/inference/api-keys.md): Create, rotate, and manage API keys that authenticate your inference requests to the Lilac API. Each key is scoped to an organization. - [Connect to local coding tools](https://docs.getlilac.com/inference/local-tools.md): Connect Lilac inference to OpenCode, Continue, Cursor, and other OpenAI-compatible AI coding assistants running in your local development environment. - [Organizations & Invites](https://docs.getlilac.com/inference/organizations.md): Manage your Lilac organization, invite team members by email, assign owner or member roles, and control API key and billing access. - [Chat Completions](https://docs.getlilac.com/inference/chat-completions.md): Use the OpenAI-compatible chat completions endpoint to generate model responses from conversation history, with streaming and tool calling. - [Completions (Legacy)](https://docs.getlilac.com/inference/completions.md): Use the legacy completions endpoint to generate text from a raw prompt string. For new integrations, use the chat completions endpoint instead. - [Responses API](https://docs.getlilac.com/inference/responses.md): Use the responses endpoint, OpenAI's newer API format, to generate structured output and call tools with built-in support for JSON schemas. - [OpenAI Compatibility](https://docs.getlilac.com/inference/openai-compatibility.md): Learn which OpenAI API features Lilac supports and how to migrate from OpenAI by changing just the base URL and API key in your existing code. - [API Status and Model Performance](https://docs.getlilac.com/inference/status.md): Query the public Lilac status endpoint for live model uptime, throughput, and time-to-first-token metrics across configurable aggregation windows. - [API Rate Limits and 429 Handling](https://docs.getlilac.com/inference/rate-limits.md): Default per-organization rate limits for the Lilac inference API, how 429 Too Many Requests responses work, and recommended retry and backoff behavior. - [Inference Pricing](https://docs.getlilac.com/inference/pricing.md): Lilac offers pay-per-token inference pricing with no minimums or contracts. See per-model rates for input and output tokens powered by idle GPUs. - [Usage & Billing](https://docs.getlilac.com/inference/usage.md): Monitor your token consumption, view per-model cost breakdowns, and manage prepaid credit billing from the Lilac dashboard in real time. - [Personal subscriptions](https://docs.getlilac.com/billing/subscription-rates.md): Lilac personal subscriptions bundle monthly included model usage with live per-model discounts. - [Dedicated GPUs](https://docs.getlilac.com/dedicated-gpus/overview.md): Source dedicated GPU virtual machines, bare-metal nodes, and private multi-node clusters through Lilac's network. - [Bare-metal partnerships](https://docs.getlilac.com/suppliers/bare-metal.md): Work with Lilac to connect current or planned bare-metal GPU capacity with suitable customer demand. - [Kubernetes operator program](https://docs.getlilac.com/suppliers/operator/overview.md): Contribute idle GPU capacity to Lilac's shared inference network from an existing Kubernetes cluster. - [Kubernetes operator onboarding](https://docs.getlilac.com/suppliers/getting-started.md): Create a supplier account, share your cluster details, complete onboarding, and install the Kubernetes operator. - [How the Operator Works](https://docs.getlilac.com/suppliers/operator/how-it-works.md): Understand the Lilac GPU operator architecture, its 30-second sync loop with the control plane, and how it manages inference pods on idle GPUs. - [Operator Installation](https://docs.getlilac.com/suppliers/operator/installation.md): Install the Lilac GPU operator in your Kubernetes cluster using Helm. Covers prerequisites, namespace setup, chart configuration, and verification. - [GPU Pool Configuration](https://docs.getlilac.com/suppliers/operator/gpu-pools.md): Define GPU pool custom resources to control which nodes, how many GPUs, availability schedules, and preemption rules Lilac uses in your cluster. - [GPU Preemption](https://docs.getlilac.com/suppliers/operator/preemption.md): Learn how the Lilac operator gracefully reclaims GPUs when your workloads need them back, using LIFO eviction and configurable grace periods. - [Operator revenue and payouts](https://docs.getlilac.com/suppliers/revenue.md): Understand how earnings, reporting, and payouts work for suppliers in the Kubernetes operator program. - [Operator cluster monitoring](https://docs.getlilac.com/suppliers/monitoring.md): Monitor Kubernetes operator connectivity, workload activity, and cluster health through the Lilac dashboard and Kubernetes tools. ## Optional - [Talk to Us](https://calendly.com/d/ctxy-jd8-585/lilac-support) - [Contact Us](mailto:contact@getlilac.com)