Robuta

https://github.com/vllm-project/vllm GitHub - vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for... A high-throughput and memory-efficient inference and serving engine for LLMs - vllm-project/vllm https://vllm.ai/ vLLM vLLM is a high-throughput and memory-efficient inference and serving engine for Large Language Models (LLMs). Deploy AI models faster with state-of-the-art... vllm https://docs.vllm.ai/en/latest/ vLLM vllm https://docs.vllm.ai/en/stable/getting_started/installation/index.html Installation - vLLM installationvllm https://recipes.vllm.ai/baidu/Unlimited-OCR baidu/Unlimited-OCR | vLLM Recipes Baidu's state-of-the-art document-parsing model with Reference Sliding Window Attention (R-SWA), optimized for full-page OCR and markdown generation. baiduunlimitedocrvllmrecipes https://docs.vllm.ai/projects/recipes/en/latest/Google/Gemma4.html Gemma 4 Usage Guide - vLLM Recipes A collection of recipes and guides for using vLLM with a variety of models. usage guidegemmavllmrecipes https://github.com/defilantech/llmkube GitHub - defilantech/LLMKube: Kubernetes operator for local LLM inference with llama.cpp, vLLM, and... Kubernetes operator for local LLM inference with llama.cpp, vLLM, and TGI - multi-GPU, autoscaling, air-gapped, production-ready - defilantech/LLMKube https://recipes.vllm.ai/moonshotai/Kimi-K3 moonshotai/Kimi-K3 | vLLM Recipes Pre-release 2.8T-parameter native multimodal MoE with Kimi Delta Attention, Gated MLA, Attention Residuals, and a 1M-token context window moonshotaikimivllmrecipes https://www.docker.com/blog/docker-model-runner-vllm-metal-macos/ Docker Model Runner Adds vLLM Support on macOS | Docker Mar 16, 2026 - Run vLLM on your Mac with Docker Model Runner. The vllm-metal backend enables high-performance LLM inference on Apple Silicon with Metal GPU acceleration. docker model runnersupport onaddsvllmmacos https://docs.vllm.ai/en/stable/features/quantization/torchao/ TorchAO - vLLM torchaovllm https://docs.vllm.ai/en/stable/serving/parallelism_scaling/ Parallelism and Scaling - vLLM parallelism and scalingvllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/glm4_moe_mtp/ glm4_moe_mtp - vLLM moemtpvllm https://communityinviter.com/apps/vllm-dev/join-vllm-developers-slack Join vLLM on Slack - Community Inviter Join vLLM on Slack. Powered by Community Inviter. You will get an invitation soon. Check your inbox. on slackjoinvllmcommunityinviter https://docs.cloud.google.com/ai-hypercomputer/docs/tutorials/tpu/serve-qwen2-7b-instruct Serve Qwen2-7B-Instruct with vLLM on TPUs | AI Hypercomputer | Google Cloud Documentation Serve the Qwen2-7B-Instruct on Cloud TPU Trillium using the vLLM serving framework. https://developers.llamaindex.ai/python/framework/integrations/llm/vllm/ vLLM | Developer Documentation vllmdeveloperdocumentation https://docs.vllm.ai/en/stable/api/vllm/models/deepseek_v32/nvidia/attention/ attention - vLLM attentionvllm https://docs.vllm.ai/en/stable/api/vllm/v1/attention/backends/mla/flashinfer_mla_sparse_sm120/ flashinfer_mla_sparse_sm120 - vLLM flashinfermlasparsevllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/quantization/compressed_tensors/compressed_tensors_moe/compressed_tensors_moe_w4a8_int8/ compressed_tensors_moe_w4a8_int8 - vLLM compressedtensorsmoevllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/fused_moe/config/ config - vLLM configvllm https://recipes.vllm.ai/fishaudio Fish Audio on vLLM — 1 recipe vLLM serve recipes for Fish Audio models — 1 recipe with hardware-tuned commands. fish audiovllmrecipe https://docs.vllm.ai/en/stable/serving/online_serving/derenderer/ Derenderer APIs - vLLM apisvllm https://docs.vllm.ai/en/stable/getting_started/installation/gpu/ GPU - vLLM gpuvllm https://recipes.vllm.ai/moonshotai/Kimi-K2.6 moonshotai/Kimi-K2.6 | vLLM Recipes Open-source native multimodal agentic MoE model with vision-language understanding, tool calling, and thinking modes moonshotaikimivllmrecipes https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/quantization/quark/quark/ quark - vLLM quarkvllm https://docs.nvidia.com/nemo/gym/v0.2.1/model-server/vllm/ vLLM | NeMo Gym vllmnemogym https://docs.vllm.ai/en/stable/models/extensions/instanttensor/ Loading Model Weights with InstantTensor - vLLM model weightsloadingvllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/mistral_large_3_eagle/ mistral_large_3_eagle - vLLM mistral largeeaglevllm https://docs.vllm.ai/en/stable/api/vllm/v1/worker/gpu/pool/pooling_runner/ pooling_runner - vLLM poolingrunnervllm https://docs.vllm.ai/en/stable/api/vllm/tool_parsers/xlam_tool_parser/ xlam_tool_parser - vLLM xlamtoolparservllm https://recipes.vllm.ai/ vLLM Recipes — Deploy any model on any hardware with vLLM How do I run model X on hardware Y? Pick a model, get a working vllm serve command. any modelvllmrecipesdeployhardware https://forums.developer.nvidia.com/t/new-bleeding-edge-vllm-docker-image-avarok-vllm-nvfp4-gb10-sm120/354231 New bleeding-edge vLLM Docker Image: avarok/vllm-nvfp4-gb10-sm120 - DGX Spark / GB10 Projects -... Dec 11, 2025 - Running NVFP4 MoE Models Copy and paste this code to get a snappy SOTA qwen3-next model running on your DGX Spark at an NVFP4 quant: # Pull the pre-built... https://docs.vllm.ai/en/stable/api/vllm/entrypoints/scale_out/token_in_token_out/ token_in_token_out - vLLM in outtokenvllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/llama/ llama - vLLM llamavllm https://docs.vllm.ai/en/stable/examples/tool_calling/openai_responses_client_with_tools/ OpenAI Responses Client With Tools - vLLM openai responses clienttoolsvllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/gemma3n/ gemma3n - vLLM vllm https://vllm.ai/blog Blog | vLLM Technical articles, release announcements, model guides, and community updates from the vLLM project. blogvllm https://advisories.gitlab.com/pkg/pypi/vllm/GHSA-j828-28rj-hfhp/ https://advisories.gitlab.com/pypi/vllm/GHSA-j828-28rj-hfhp/ httpsadvisoriesgitlabpypivllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/quantization/compressed_tensors/schemes/compressed_tensors_w8a8_fp8/ compressed_tensors_w8a8_fp8 - vLLM compressedtensorsvllm https://docs.vllm.ai/en/stable/api/vllm/parser/parser_manager/ parser_manager - vLLM parsermanagervllm https://forums.developer.nvidia.com/t/problem-running-qwen-models-via-vllm-on-jetson-orin/368978 Problem running Qwen Models via vllm on Jetson Orin - Jetson AGX Orin - NVIDIA Developer Forums May 5, 2026 - I have been trying to run qwen models e.g. qwen 3.6 35b a3b and qwen 3.5 35b a3b, qwen 3.5 9b on my jetson orin but i have been getting this error... https://docs.vllm.ai/en/stable/api/vllm/entrypoints/cli/main/ main - vLLM mainvllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/mamba/linear/minimax_linear_attn/ minimax_linear_attn - vLLM minimaxlinearattnvllm https://docs.vllm.ai/projects/tpu/en/latest/recommended_models_features/ Recommended Models and Features - vLLM TPU recommendedmodelsfeaturesvllmtpu https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/glm4_moe_lite/ glm4_moe_lite - vLLM moelitevllm https://docs.vllm.ai/en/stable/features/sleep_mode/ Sleep Mode - vLLM sleep modevllm https://recipes.vllm.ai/JetBrains JetBrains on vLLM — 2 recipes vLLM serve recipes for JetBrains models — 2 recipes with hardware-tuned commands. jetbrainsvllmrecipes https://docs.vllm.ai/en/stable/api/vllm/logits_process/ logits_process - vLLM logitsprocessvllm https://recipes.vllm.ai/bosonai Boson AI on vLLM — 1 recipe vLLM serve recipes for Boson AI models — 1 recipe with hardware-tuned commands. boson aivllmrecipe https://advisories.gitlab.com/pkg/pypi/vllm/CVE-2025-62372/ https://advisories.gitlab.com/pypi/vllm/CVE-2025-62372/ httpsadvisoriesgitlabpypivllm https://docs.vllm.ai/en/stable/api/vllm/v1/sample/ops/topk_topp_triton/ topk_topp_triton - vLLM topktopptritonvllm https://docs.vllm.ai/en/stable/usage/ Using vLLM - vLLM usingvllm https://docs.vllm.ai/en/stable/api/vllm/reasoning/abs_reasoning_parsers/ abs_reasoning_parsers - vLLM absreasoningparsersvllm https://docs.vllm.ai/en/stable/api/vllm/triton_utils/allocation/ allocation - vLLM allocationvllm https://grafana.com/grafana/dashboards/24756-vllm-monitoring-v2/ vLLM Monitoring - V2 | Grafana Labs The vLLM Monitoring - V2 dashboard uses the prometheus data source to create a Grafana dashboard with the bargauge, gauge, heatmap, piechart, stat and... vllmmonitoringgrafanalabs https://run-ai-docs.nvidia.com/saas/workloads-in-nvidia-run-ai/workload-templates/inference-templates/hugging-face-inference-templates vLLM Inference Templates | SaaS | Run:ai Documentation run aivllminferencetemplatessaas https://docs.vllm.ai/en/stable/api/vllm/reasoning/basic_parsers/ basic_parsers - vLLM basicparsersvllm https://docs.vllm.ai/en/stable/api/vllm/models/deepseek_v4/attention/ attention - vLLM attentionvllm https://advisories.gitlab.com/pypi/vllm/CVE-2026-34756/ vLLM: Unauthenticated OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server |... CVE-2026-34756 vLLM: Unauthenticated OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server: A Denial of Service vulnerability exists in the... https://recipes.vllm.ai/moonshotai/Kimi-K2.5 moonshotai/Kimi-K2.5 | vLLM Recipes Open-source native multimodal agentic MoE model with vision-language understanding, tool calling, and thinking modes moonshotaikimivllmrecipes https://docs.vllm.ai/en/stable/api/vllm/model_executor/kernels/linear/scaled_mm/aiter/ aiter - vLLM aitervllm https://docs.vllm.ai/en/stable/api/vllm/models/deepseek_v4/amd/rocm/ rocm - vLLM rocmvllm https://recipes.vllm.ai/Google Google on vLLM — 7 recipes vLLM serve recipes for Google models — 7 recipes with hardware-tuned commands. google onvllmrecipes https://aws.amazon.com/blogs/machine-learning/boost-cold-start-recommendations-with-vllm-on-aws-trainium/ Boost cold-start recommendations with vLLM on AWS Trainium | Artificial Intelligence Jul 24, 2025 - In this post, we demonstrate how to use vLLM for scalable inference and use AWS Deep Learning Containers (DLC) to streamline model packaging and deployment.... cold starton awsboostrecommendations https://docs.cloud.google.com/kubernetes-engine/docs/tutorials/serve-vllm-tpu Serve an LLM using TPU Trillium on GKE with vLLM | GKE AI/ML | Google Cloud Documentation https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/deepseek_eagle/ deepseek_eagle - vLLM deepseekeaglevllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/rotary_embedding/gemma4_rope/ gemma4_rope - vLLM ropevllm https://play.vllm-semantic-router.com/ vLLM Semantic Router vLLM Semantic Router - Manage your AI-powered Intelligent Router vllmsemanticrouter https://www.globenewswire.com/news-release/2025/03/12/3041490/0/en/Pliops-Announces-Collaboration-with-vLLM-Production-Stack-to-Enhance-LLM-Inference-Performance.html Pliops Announces Collaboration with vLLM Production Stack Pliops has announced a strategic collaboration with the vLLM Production Stack to revolutionize large language model (LLM) inference performance. ... collaboration withpliopsannouncesvllmproduction https://docs.vllm.ai/en/stable/api/vllm/model_executor/kernels/ kernels - vLLM kernelsvllm https://docs.vllm.ai/en/stable/api/vllm/v1/simple_kv_offload/metadata/ metadata - vLLM metadatavllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/openpangu/ openpangu - vLLM openpanguvllm https://research.ibm.com/publications/vllm-in-confidential-cpu-gpu-enclaves-does-it-perform vLLM in Confidential CPU-GPU Enclaves: Does it Perform? for AICS 2024 - IBM Research vLLM in Confidential CPU-GPU Enclaves: Does it Perform? for AICS 2024 by Mengmei Ye et al. https://advisories.gitlab.com/pypi/vllm/CVE-2026-34753/ vLLM: Server-Side Request Forgery (SSRF) in `download_bytes_from_url ` | GitLab Advisory Database... CVE-2026-34753 vLLM: Server-Side Request Forgery (SSRF) in `download_bytes_from_url `: A Server Side Request Forgery (SSRF) vulnerability in... server side request forgery https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/fused_moe/experts/lora_context/ lora_context - vLLM loracontextvllm https://docs.vllm.ai/en/stable/api/vllm/models/minimax_m3/common/ops/ ops - vLLM opsvllm https://docs.vllm.ai/en/stable/features/interleaved_thinking/ Interleaved Thinking - vLLM interleaved thinkingvllm https://qwen.readthedocs.io/en/latest/deployment/vllm.html vLLM - Qwen vllmqwen https://www.vllm.ch/ vLLM Experts Switzerland – LLM Inference Consulting | VSHN vLLM consulting and operations in Switzerland. VSHN deploys and manages high-throughput LLM inference on Kubernetes with full Swiss data residency. ISO 27001. llm inferencevllmexpertsswitzerlandconsulting https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/fused_moe/experts/nvfp4_emulation_moe/ nvfp4_emulation_moe - vLLM emulationmoevllm https://docs.vllm.ai/en/stable/api/vllm/cute_utils/cvt/ cvt - vLLM cvtvllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/telechat2/ telechat2 - vLLM vllm https://docs.vllm.ai/en/stable/api/vllm/entrypoints/pooling/embed/serving/ serving - vLLM servingvllm https://docs.vllm.ai/projects/recipes/en/latest/StepFun/Step-3.5-Flash.html Step-3.5-Flash Guide - vLLM Recipes A collection of recipes and guides for using vLLM with a variety of models. flash guidestepvllmrecipes https://docs.vllm.ai/en/stable/api/vllm/multimodal/encoder_budget/ encoder_budget - vLLM encoderbudgetvllm https://cloud.google.com/blog/products/compute/in-q3-2025-ai-hypercomputer-adds-vllm-tpu-and-more In Q3 2025, AI Hypercomputer adds vLLM TPU and more | Google Cloud Blog Amongst the latest developments to AI Hypercomputer is vLLM TPU. https://docs.vllm.ai/en/stable/models/extensions/runai_model_streamer/ Loading models with Run:ai Model Streamer - vLLM run ailoadingmodelsstreamervllm https://docs.vllm.ai/en/stable/features/batch_invariance/ Batch Invariance - vLLM batch invariancevllm https://docs.vllm.ai/en/stable/api/vllm/entrypoints/pooling/classify/ classify - vLLM classifyvllm https://docs.vllm.ai/en/stable/training/weight_transfer/nccl/ NCCL Engine - vLLM nccl enginevllm https://recipes.vllm.ai/stabilityai Stability AI on vLLM — 2 recipes vLLM serve recipes for Stability AI models — 2 recipes with hardware-tuned commands. stability aivllmrecipes https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/fused_moe/fused_moe_modular_method/ fused_moe_modular_method - vLLM fusedmoemodularmethodvllm https://docs.vllm.ai/en/stable/api/vllm/entrypoints/cli/benchmark/latency/ latency - vLLM latencyvllm https://docs.vllm.ai/en/stable/api/vllm/tool_parsers/gemma4_engine_tool_parser/ gemma4_engine_tool_parser - vLLM engine toolparservllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/layers/fused_moe/prepare_finalize/deepep_v2/ deepep_v2 - vLLM deepepvllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/siglip/ siglip - vLLM siglipvllm https://docs.vllm.ai/en/stable/api/vllm/model_executor/models/clip/ clip - vLLM clipvllm https://www.redhat.com/en/blog/unleash-full-potential-llms-optimize-performance-vllm Unleash the full potential of LLMs: Optimize for performance with vLLM What if you could harness LLM power without breaking the bank? Model compression and efficient inference with vLLM offer a game-changing answer, helping reduce... full potentialfor performanceunleash https://advisories.gitlab.com/pypi/vllm/CVE-2025-48943/ vLLM allows clients to crash the openai server with invalid regex | GitLab Advisory Database (GLAD) CVE-2025-48943 vLLM allows clients to crash the openai server with invalid regex: A denial of service bug caused the vLLM server to crash if an invalid regex... https://docs.redhat.com/en/documentation/red_hat_ai_inference_server/3.1/html/llm_compressor/integration-with-rhaiis-and-vllm_llm-compressor Chapter 3. Integration with Red Hat AI Inference Server and vLLM | LLM Compressor | Red Hat AI... Chapter 3. Integration with Red Hat AI Inference Server and vLLM | LLM Compressor | Red Hat AI Inference Server | 3.1 | Red Hat Documentation red hat ai inference https://docs.vllm.ai/en/stable/api/vllm/v1/worker/cpu_model_runner/ cpu_model_runner - vLLM model runnercpuvllm