Robuta

https://github.com/sgl-project/sglang GitHub - sgl-project/sglang: SGLang is a high-performance serving framework for large language... SGLang is a high-performance serving framework for large language models and multimodal models. - sgl-project/sglang https://lmsysorg.mintlify.app/cookbook/autoregressive/Google/Gemma4 Gemma 4 - SGLang Documentation gemmasglangdocumentation https://docs.sglang.io/cookbook/autoregressive/Google/Gemma4 Gemma 4 - SGLang Documentation gemmasglangdocumentation https://docs.sglang.io/cookbook/autoregressive/Moonshotai/Kimi-K3 Kimi-K3 - SGLang Documentation Jul 29, 2026 - Deploy Moonshot AI's Kimi-K3 with SGLang — a 2.8T-parameter hybrid Mixture-of-Experts vision-language model (Kimi Delta Attention + MLA, 16/896 active experts)... kimisglangdocumentation https://fergusfinn.com/blog/fast-sglang-starts/ 70x faster cold(ish) starts for SGLang Checkpoint/restore with CRIU and cuda-checkpoint, from 12 minutes to 10 seconds on a B200. fastercoldishstartssglang https://docs.sglang.io/docs/get-started/install Installation - SGLang Documentation Jul 25, 2026 - Install SGLang with pip/uv, source, Docker, Kubernetes, and cloud deployment options. installationsglangdocumentation https://docs.sglang.io/cookbook/autoregressive/Moonshotai/Kimi-K2 Kimi-K2 - SGLang Documentation kimisglangdocumentation https://lmsysorg.mintlify.app/ Welcome to SGLang - SGLang Documentation High-performance serving framework for large language and multimodal models. welcome tosglangdocumentation https://docs.sglang.io/docs/hardware-platforms/nvidia_jetson NVIDIA Jetson Orin - SGLang Documentation Jul 13, 2026 - Guide for installing and running SGLang on NVIDIA Jetson Orin devices. nvidia jetsonorinsglangdocumentation https://docs.sglang.io/docs/hardware-platforms/tpu TPU - SGLang Documentation Apr 21, 2026 - SGLang supports high-performance TPU inference through the SGLang-JAX backend, which is specifically optimized for Google Cloud TPUs. The JAX-based... tpusglangdocumentation https://docs.sglang.io/docs/hardware-platforms/ascend-npus/ascend_npu SGLang installation with NPUs support - SGLang Documentation Jul 20, 2026 - Complete installation guide for SGLang on Ascend NPUs, including component version mapping, environment setup, and launching inference services. sglanginstallationsupportdocumentation https://aur.archlinux.org/packages/sglang AUR (en) - sglang aurensglang https://docs.sglang.io/cookbook/intro SGLang Cookbook - SGLang Documentation sglangcookbookdocumentation https://build.nvidia.com/spark/sglang SGLang for Inference | DGX Spark Install and use SGLang on DGX Spark sglanginferencedgxspark https://docs.sglang.io/docs/basic_usage/send_request Tutorial: Sending a request - SGLang Documentation a requesttutorialsendingsglangdocumentation https://docs.sglang.io/docs/hardware-platforms/amd_gpu AMD GPUs - SGLang Documentation amd gpussglangdocumentation https://docs.sglang.io/docs/get-started/quickstart Quickstart - SGLang Documentation Jun 29, 2026 - Get up and running with SGLang in minutes: install, launch a server, and send your first request. quickstartsglangdocumentation https://docs.nvidia.com/deeplearning/frameworks/sglang-release-notes/rel-26-03.html SGLang Release 26.03 - NVIDIA Docs NVIDIA Optimized Frameworks such as Kaldi, NVIDIA Optimized Deep Learning Framework (powered by Apache MXNet), NVCaffe, PyTorch, and TensorFlow (which includes... sglangreleasenvidiadocs https://docs.sglang.io/docs/hardware-platforms/apple_metal Apple Silicon with Metal - SGLang Documentation apple siliconmetalsglangdocumentation https://www.sglang.io/ Welcome to SGLang - SGLang Homepage SGLang powers fast, scalable inference for large language and multimodal models. Open-source serving framework with state-of-the-art performance. welcome tosglanghomepage https://www.frontierswe.com/inference-system-optimization FrontierSWE - SGLang Inference System Optimization Make SGLang serving for Qwen3.5-4B faster on a B200 GPU. frontierswesglanginferencesystemoptimization https://docs.sglang.io/docs/hardware-platforms/overview Hardware Platforms - SGLang Documentation Jul 31, 2026 - Platform-specific guides for running SGLang on GPUs, TPUs, NPUs, CPUs, and more. hardware platformssglangdocumentation https://docs.sglang.io/cookbook/autoregressive/Moonshotai/Kimi-K2.6 Kimi-K2.6 - SGLang Documentation kimisglangdocumentation https://github.com/sgl-project/sglang/discussions/16031 SGLang Public Community Events · sgl-project/sglang · Discussion #16031 · GitHub SGLang Public Community Events public community eventsproject discussionsglanggithub https://docs.sglang.io/docs/hardware-platforms/xpu XPU - SGLang Documentation xpusglangdocumentation https://docs.sglang.io/docs/basic_usage/aws_sagemaker Amazon SageMaker AI - SGLang Documentation Jun 16, 2026 - Deploy SGLang on Amazon SageMaker AI endpoints using the AWS Deep Learning Container. amazon sagemaker aisglangdocumentation https://docs.nvidia.com/dynamo/v-0-9-1/integrations/sg-lang-hi-cache SGLang HiCache | NVIDIA Dynamo Documentation nvidia dynamosglangdocumentation https://dev.to/zkaria_gamal_3cddbbff21c8/concurrent-llm-serving-benchmarking-vllm-vs-sglang-vs-ollama-1cpn Concurrent LLM Serving: Benchmarking vLLM vs SGLang vs Ollama - DEV Community I wanted to know exactly how the three most popular open-source LLM serving engines perform when real... Tagged with llm, ai, cloudnative. concurrentllmservingbenchmarkingvs https://docs.nvidia.com/dynamo/dev/backends/sg-lang SGLang | NVIDIA Dynamo Documentation nvidia dynamosglangdocumentation https://docs.nvidia.com/dynamo/dev/backends/sg-lang/chat-processor SGLang Chat Processor | NVIDIA Dynamo Documentation SGLang-native preprocessing and postprocessing for chat completions nvidia dynamosglangchatprocessordocumentation