https://arxiv.org/abs/2501.08090
[2501.08090] Hierarchical Autoscaling for Large Language Model Serving with Chiron
Abstract page for arXiv paper 2501.08090: Hierarchical Autoscaling for Large Language Model Serving with Chiron
large language modelhierarchicalautoscaling
https://www.alibabacloud.com/help/en/asm/sidecar/integrate-the-kserve-on-asm-feature-with-fluid-to-implement-ai-serving-that-accelerates-data-access
Accelerate AI Model Serving with KServe and Fluid - Service Mesh - Alibaba Cloud
Integrate KServe on ASM with Fluid to cache AI models locally from OSS, enabling faster inference pod startup and eliminating remote download latency.
accelerate aimodel serving
https://www.analyticsvidhya.com/blog/2025/06/ml-model-serving/
ML Model Serving with FastAPI and Redis for faster predictions
ml modelservingfastapiredisfaster
https://www.pluralsight.com/courses/ai-ops-model-serving-architecture
AIOps: Model Serving Architecture
model servingaiopsarchitecture
https://arxiv.org/abs/2504.15856
[2504.15856] FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
Abstract page for arXiv paper 2504.15856: FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
model serving
https://ubiops.com/
UbiOps - AI Model Serving & Orchestration
ai modelubiopsservingorchestration
https://kthena.volcano.sh/
Kthena - AI Model Serving Platform for Kubernetes | Kthena
Kubernetes-native AI serving platform for scalable model serving.
ai modelkthenaservingplatformkubernetes
https://www.infoq.com/presentations/llm-meta/
Scaling Large Language Model Serving Infrastructure at Meta - InfoQ
large language modelscalingservinginfrastructuremeta
https://speakerdeck.com/kahnwong/pycon-thailand-2025-ml-model-serving-optimization-with-onnx
Pycon Thailand 2025 - ML Model Serving Optimization with ONNX - Speaker Deck
thailand 2025ml modelpycon
https://emrrc.com/
The Elkhart Model Railroad Club – Serving Northern Indiana and Southwestern Michigan Since 1950
Serving Northern Indiana and Southwestern Michigan Since 1950
https://deepai.org/publication/sensix-bringing-mlops-and-multi-tenant-model-serving-to-sensory-edge-devices
SensiX++: Bringing MLOPs and Multi-tenant Model Serving to Sensory Edge Devices | DeepAI
Sep 8, 2021 - 09/08/21 - We present SensiX++ - a multi-tenant runtime for adaptive model execution with integrated MLOps on edge devices, e.g., a camera, a...