https://github.com/vllm-project/vllm
GitHub - vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for...
A high-throughput and memory-efficient inference and serving engine for LLMs - vllm-project/vllm
https://vllm.ai/
vLLM
vLLM is a high-throughput and memory-efficient inference and serving engine for Large Language Models (LLMs). Deploy AI models faster with state-of-the-art...
vllm
https://docs.vllm.ai/en/latest/
vLLM
vllm
https://docs.vllm.ai/en/stable/getting_started/installation/index.html
Installation - vLLM
installationvllm
https://recipes.vllm.ai/baidu/Unlimited-OCR
baidu/Unlimited-OCR | vLLM Recipes
Baidu's state-of-the-art document-parsing model with Reference Sliding Window Attention (R-SWA), optimized for full-page OCR and markdown generation.
baiduunlimitedocrvllmrecipes
https://docs.vllm.ai/projects/recipes/en/latest/Google/Gemma4.html
Gemma 4 Usage Guide - vLLM Recipes
A collection of recipes and guides for using vLLM with a variety of models.
usage guidegemmavllmrecipes
https://github.com/defilantech/llmkube
GitHub - defilantech/LLMKube: Kubernetes operator for local LLM inference with llama.cpp, vLLM, and...
Kubernetes operator for local LLM inference with llama.cpp, vLLM, and TGI - multi-GPU, autoscaling, air-gapped, production-ready - defilantech/LLMKube
https://recipes.vllm.ai/moonshotai/Kimi-K3
moonshotai/Kimi-K3 | vLLM Recipes
Pre-release 2.8T-parameter native multimodal MoE with Kimi Delta Attention, Gated MLA, Attention Residuals, and a 1M-token context window
moonshotaikimivllmrecipes
https://www.docker.com/blog/docker-model-runner-vllm-metal-macos/
Docker Model Runner Adds vLLM Support on macOS | Docker
Mar 16, 2026 - Run vLLM on your Mac with Docker Model Runner. The vllm-metal backend enables high-performance LLM inference on Apple Silicon with Metal GPU acceleration.
docker model runnersupport onaddsvllmmacos
https://docs.vllm.ai/en/stable/features/quantization/torchao/
TorchAO - vLLM
torchaovllm
https://docs.vllm.ai/en/stable/serving/parallelism_scaling/
Parallelism and Scaling - vLLM
parallelism and scalingvllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/glm4_moe_mtp/
glm4_moe_mtp - vLLM
moemtpvllm
https://communityinviter.com/apps/vllm-dev/join-vllm-developers-slack
Join vLLM on Slack - Community Inviter
Join vLLM on Slack. Powered by Community Inviter. You will get an invitation soon. Check your inbox.
on slackjoinvllmcommunityinviter
https://docs.cloud.google.com/ai-hypercomputer/docs/tutorials/tpu/serve-qwen2-7b-instruct
Serve Qwen2-7B-Instruct with vLLM on TPUs | AI Hypercomputer | Google Cloud Documentation
Serve the Qwen2-7B-Instruct on Cloud TPU Trillium using the vLLM serving framework.
https://developers.llamaindex.ai/python/framework/integrations/llm/vllm/
vLLM | Developer Documentation
vllmdeveloperdocumentation
https://docs.vllm.ai/en/stable/api/vllm/models/deepseek_v32/nvidia/attention/
attention - vLLM
attentionvllm
https://docs.vllm.ai/en/stable/api/vllm/v1/attention/backends/mla/flashinfer_mla_sparse_sm120/
flashinfer_mla_sparse_sm120 - vLLM
flashinfermlasparsevllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/quantization/compressed_tensors/compressed_tensors_moe/compressed_tensors_moe_w4a8_int8/
compressed_tensors_moe_w4a8_int8 - vLLM
compressedtensorsmoevllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/fused_moe/config/
config - vLLM
configvllm
https://recipes.vllm.ai/fishaudio
Fish Audio on vLLM — 1 recipe
vLLM serve recipes for Fish Audio models — 1 recipe with hardware-tuned commands.
fish audiovllmrecipe
https://docs.vllm.ai/en/stable/serving/online_serving/derenderer/
Derenderer APIs - vLLM
apisvllm
https://docs.vllm.ai/en/stable/getting_started/installation/gpu/
GPU - vLLM
gpuvllm
https://recipes.vllm.ai/moonshotai/Kimi-K2.6
moonshotai/Kimi-K2.6 | vLLM Recipes
Open-source native multimodal agentic MoE model with vision-language understanding, tool calling, and thinking modes
moonshotaikimivllmrecipes
https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/quantization/quark/quark/
quark - vLLM
quarkvllm
https://docs.nvidia.com/nemo/gym/v0.2.1/model-server/vllm/
vLLM | NeMo Gym
vllmnemogym
https://docs.vllm.ai/en/stable/models/extensions/instanttensor/
Loading Model Weights with InstantTensor - vLLM
model weightsloadingvllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/mistral_large_3_eagle/
mistral_large_3_eagle - vLLM
mistral largeeaglevllm
https://docs.vllm.ai/en/stable/api/vllm/v1/worker/gpu/pool/pooling_runner/
pooling_runner - vLLM
poolingrunnervllm
https://docs.vllm.ai/en/stable/api/vllm/tool_parsers/xlam_tool_parser/
xlam_tool_parser - vLLM
xlamtoolparservllm
https://recipes.vllm.ai/
vLLM Recipes — Deploy any model on any hardware with vLLM
How do I run model X on hardware Y? Pick a model, get a working vllm serve command.
any modelvllmrecipesdeployhardware
https://forums.developer.nvidia.com/t/new-bleeding-edge-vllm-docker-image-avarok-vllm-nvfp4-gb10-sm120/354231
New bleeding-edge vLLM Docker Image: avarok/vllm-nvfp4-gb10-sm120 - DGX Spark / GB10 Projects -...
Dec 11, 2025 - Running NVFP4 MoE Models Copy and paste this code to get a snappy SOTA qwen3-next model running on your DGX Spark at an NVFP4 quant: # Pull the pre-built...
https://docs.vllm.ai/en/stable/api/vllm/entrypoints/scale_out/token_in_token_out/
token_in_token_out - vLLM
in outtokenvllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/llama/
llama - vLLM
llamavllm
https://docs.vllm.ai/en/stable/examples/tool_calling/openai_responses_client_with_tools/
OpenAI Responses Client With Tools - vLLM
openai responses clienttoolsvllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/gemma3n/
gemma3n - vLLM
vllm
https://vllm.ai/blog
Blog | vLLM
Technical articles, release announcements, model guides, and community updates from the vLLM project.
blogvllm
https://advisories.gitlab.com/pkg/pypi/vllm/GHSA-j828-28rj-hfhp/
https://advisories.gitlab.com/pypi/vllm/GHSA-j828-28rj-hfhp/
httpsadvisoriesgitlabpypivllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/quantization/compressed_tensors/schemes/compressed_tensors_w8a8_fp8/
compressed_tensors_w8a8_fp8 - vLLM
compressedtensorsvllm
https://docs.vllm.ai/en/stable/api/vllm/parser/parser_manager/
parser_manager - vLLM
parsermanagervllm
https://forums.developer.nvidia.com/t/problem-running-qwen-models-via-vllm-on-jetson-orin/368978
Problem running Qwen Models via vllm on Jetson Orin - Jetson AGX Orin - NVIDIA Developer Forums
May 5, 2026 - I have been trying to run qwen models e.g. qwen 3.6 35b a3b and qwen 3.5 35b a3b, qwen 3.5 9b on my jetson orin but i have been getting this error...
https://docs.vllm.ai/en/stable/api/vllm/entrypoints/cli/main/
main - vLLM
mainvllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/mamba/linear/minimax_linear_attn/
minimax_linear_attn - vLLM
minimaxlinearattnvllm
https://docs.vllm.ai/projects/tpu/en/latest/recommended_models_features/
Recommended Models and Features - vLLM TPU
recommendedmodelsfeaturesvllmtpu
https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/glm4_moe_lite/
glm4_moe_lite - vLLM
moelitevllm
https://docs.vllm.ai/en/stable/features/sleep_mode/
Sleep Mode - vLLM
sleep modevllm
https://recipes.vllm.ai/JetBrains
JetBrains on vLLM — 2 recipes
vLLM serve recipes for JetBrains models — 2 recipes with hardware-tuned commands.
jetbrainsvllmrecipes
https://docs.vllm.ai/en/stable/api/vllm/logits_process/
logits_process - vLLM
logitsprocessvllm
https://recipes.vllm.ai/bosonai
Boson AI on vLLM — 1 recipe
vLLM serve recipes for Boson AI models — 1 recipe with hardware-tuned commands.
boson aivllmrecipe
https://advisories.gitlab.com/pkg/pypi/vllm/CVE-2025-62372/
https://advisories.gitlab.com/pypi/vllm/CVE-2025-62372/
httpsadvisoriesgitlabpypivllm
https://docs.vllm.ai/en/stable/api/vllm/v1/sample/ops/topk_topp_triton/
topk_topp_triton - vLLM
topktopptritonvllm
https://docs.vllm.ai/en/stable/usage/
Using vLLM - vLLM
usingvllm
https://docs.vllm.ai/en/stable/api/vllm/reasoning/abs_reasoning_parsers/
abs_reasoning_parsers - vLLM
absreasoningparsersvllm
https://docs.vllm.ai/en/stable/api/vllm/triton_utils/allocation/
allocation - vLLM
allocationvllm
https://grafana.com/grafana/dashboards/24756-vllm-monitoring-v2/
vLLM Monitoring - V2 | Grafana Labs
The vLLM Monitoring - V2 dashboard uses the prometheus data source to create a Grafana dashboard with the bargauge, gauge, heatmap, piechart, stat and...
vllmmonitoringgrafanalabs
https://run-ai-docs.nvidia.com/saas/workloads-in-nvidia-run-ai/workload-templates/inference-templates/hugging-face-inference-templates
vLLM Inference Templates | SaaS | Run:ai Documentation
run aivllminferencetemplatessaas
https://docs.vllm.ai/en/stable/api/vllm/reasoning/basic_parsers/
basic_parsers - vLLM
basicparsersvllm
https://docs.vllm.ai/en/stable/api/vllm/models/deepseek_v4/attention/
attention - vLLM
attentionvllm
https://advisories.gitlab.com/pypi/vllm/CVE-2026-34756/
vLLM: Unauthenticated OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server |...
CVE-2026-34756 vLLM: Unauthenticated OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server: A Denial of Service vulnerability exists in the...
https://recipes.vllm.ai/moonshotai/Kimi-K2.5
moonshotai/Kimi-K2.5 | vLLM Recipes
Open-source native multimodal agentic MoE model with vision-language understanding, tool calling, and thinking modes
moonshotaikimivllmrecipes
https://docs.vllm.ai/en/stable/api/vllm/model_executor/kernels/linear/scaled_mm/aiter/
aiter - vLLM
aitervllm
https://docs.vllm.ai/en/stable/api/vllm/models/deepseek_v4/amd/rocm/
rocm - vLLM
rocmvllm
https://recipes.vllm.ai/Google
Google on vLLM — 7 recipes
vLLM serve recipes for Google models — 7 recipes with hardware-tuned commands.
google onvllmrecipes
https://aws.amazon.com/blogs/machine-learning/boost-cold-start-recommendations-with-vllm-on-aws-trainium/
Boost cold-start recommendations with vLLM on AWS Trainium | Artificial Intelligence
Jul 24, 2025 - In this post, we demonstrate how to use vLLM for scalable inference and use AWS Deep Learning Containers (DLC) to streamline model packaging and deployment....
cold starton awsboostrecommendations
https://docs.cloud.google.com/kubernetes-engine/docs/tutorials/serve-vllm-tpu
Serve an LLM using TPU Trillium on GKE with vLLM | GKE AI/ML | Google Cloud Documentation
https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/deepseek_eagle/
deepseek_eagle - vLLM
deepseekeaglevllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/rotary_embedding/gemma4_rope/
gemma4_rope - vLLM
ropevllm
https://play.vllm-semantic-router.com/
vLLM Semantic Router
vLLM Semantic Router - Manage your AI-powered Intelligent Router
vllmsemanticrouter
https://www.globenewswire.com/news-release/2025/03/12/3041490/0/en/Pliops-Announces-Collaboration-with-vLLM-Production-Stack-to-Enhance-LLM-Inference-Performance.html
Pliops Announces Collaboration with vLLM Production Stack
Pliops has announced a strategic collaboration with the vLLM Production Stack to revolutionize large language model (LLM) inference performance. ...
collaboration withpliopsannouncesvllmproduction
https://docs.vllm.ai/en/stable/api/vllm/model_executor/kernels/
kernels - vLLM
kernelsvllm
https://docs.vllm.ai/en/stable/api/vllm/v1/simple_kv_offload/metadata/
metadata - vLLM
metadatavllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/openpangu/
openpangu - vLLM
openpanguvllm
https://research.ibm.com/publications/vllm-in-confidential-cpu-gpu-enclaves-does-it-perform
vLLM in Confidential CPU-GPU Enclaves: Does it Perform? for AICS 2024 - IBM Research
vLLM in Confidential CPU-GPU Enclaves: Does it Perform? for AICS 2024 by Mengmei Ye et al.
https://advisories.gitlab.com/pypi/vllm/CVE-2026-34753/
vLLM: Server-Side Request Forgery (SSRF) in `download_bytes_from_url ` | GitLab Advisory Database...
CVE-2026-34753 vLLM: Server-Side Request Forgery (SSRF) in `download_bytes_from_url `: A Server Side Request Forgery (SSRF) vulnerability in...
server side request forgery
https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/fused_moe/experts/lora_context/
lora_context - vLLM
loracontextvllm
https://docs.vllm.ai/en/stable/api/vllm/models/minimax_m3/common/ops/
ops - vLLM
opsvllm
https://docs.vllm.ai/en/stable/features/interleaved_thinking/
Interleaved Thinking - vLLM
interleaved thinkingvllm
https://qwen.readthedocs.io/en/latest/deployment/vllm.html
vLLM - Qwen
vllmqwen
https://www.vllm.ch/
vLLM Experts Switzerland – LLM Inference Consulting | VSHN
vLLM consulting and operations in Switzerland. VSHN deploys and manages high-throughput LLM inference on Kubernetes with full Swiss data residency. ISO 27001.
llm inferencevllmexpertsswitzerlandconsulting
https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/fused_moe/experts/nvfp4_emulation_moe/
nvfp4_emulation_moe - vLLM
emulationmoevllm
https://docs.vllm.ai/en/stable/api/vllm/cute_utils/cvt/
cvt - vLLM
cvtvllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/telechat2/
telechat2 - vLLM
vllm
https://docs.vllm.ai/en/stable/api/vllm/entrypoints/pooling/embed/serving/
serving - vLLM
servingvllm
https://docs.vllm.ai/projects/recipes/en/latest/StepFun/Step-3.5-Flash.html
Step-3.5-Flash Guide - vLLM Recipes
A collection of recipes and guides for using vLLM with a variety of models.
flash guidestepvllmrecipes
https://docs.vllm.ai/en/stable/api/vllm/multimodal/encoder_budget/
encoder_budget - vLLM
encoderbudgetvllm
https://cloud.google.com/blog/products/compute/in-q3-2025-ai-hypercomputer-adds-vllm-tpu-and-more
In Q3 2025, AI Hypercomputer adds vLLM TPU and more | Google Cloud Blog
Amongst the latest developments to AI Hypercomputer is vLLM TPU.
https://docs.vllm.ai/en/stable/models/extensions/runai_model_streamer/
Loading models with Run:ai Model Streamer - vLLM
run ailoadingmodelsstreamervllm
https://docs.vllm.ai/en/stable/features/batch_invariance/
Batch Invariance - vLLM
batch invariancevllm
https://docs.vllm.ai/en/stable/api/vllm/entrypoints/pooling/classify/
classify - vLLM
classifyvllm
https://docs.vllm.ai/en/stable/training/weight_transfer/nccl/
NCCL Engine - vLLM
nccl enginevllm
https://recipes.vllm.ai/stabilityai
Stability AI on vLLM — 2 recipes
vLLM serve recipes for Stability AI models — 2 recipes with hardware-tuned commands.
stability aivllmrecipes
https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/fused_moe/fused_moe_modular_method/
fused_moe_modular_method - vLLM
fusedmoemodularmethodvllm
https://docs.vllm.ai/en/stable/api/vllm/entrypoints/cli/benchmark/latency/
latency - vLLM
latencyvllm
https://docs.vllm.ai/en/stable/api/vllm/tool_parsers/gemma4_engine_tool_parser/
gemma4_engine_tool_parser - vLLM
engine toolparservllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/fused_moe/prepare_finalize/deepep_v2/
deepep_v2 - vLLM
deepepvllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/siglip/
siglip - vLLM
siglipvllm
https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/clip/
clip - vLLM
clipvllm
https://www.redhat.com/en/blog/unleash-full-potential-llms-optimize-performance-vllm
Unleash the full potential of LLMs: Optimize for performance with vLLM
What if you could harness LLM power without breaking the bank? Model compression and efficient inference with vLLM offer a game-changing answer, helping reduce...
full potentialfor performanceunleash
https://advisories.gitlab.com/pypi/vllm/CVE-2025-48943/
vLLM allows clients to crash the openai server with invalid regex | GitLab Advisory Database (GLAD)
CVE-2025-48943 vLLM allows clients to crash the openai server with invalid regex: A denial of service bug caused the vLLM server to crash if an invalid regex...
https://docs.redhat.com/en/documentation/red_hat_ai_inference_server/3.1/html/llm_compressor/integration-with-rhaiis-and-vllm_llm-compressor
Chapter 3. Integration with Red Hat AI Inference Server and vLLM | LLM Compressor | Red Hat AI...
Chapter 3. Integration with Red Hat AI Inference Server and vLLM | LLM Compressor | Red Hat AI Inference Server | 3.1 | Red Hat Documentation
red hat ai inference
https://docs.vllm.ai/en/stable/api/vllm/v1/worker/cpu_model_runner/
cpu_model_runner - vLLM
model runnercpuvllm