https://github.com/defilantech/llmkube
GitHub - defilantech/LLMKube: Kubernetes operator for local LLM inference with llama.cpp, vLLM, and...
Kubernetes operator for local LLM inference with llama.cpp, vLLM, and TGI - multi-GPU, autoscaling, air-gapped, production-ready - defilantech/LLMKube