https://epoch.ai/data-insights/llm-inference-price-trends
LLM inference prices have fallen rapidly but unequally across tasks | Epoch AI
Epoch AI is a research institute investigating key trends and questions that will shape the trajectory and governance of Artificial Intelligence.
llm inference
https://omlx.ai/
oMLX — LLM inference, optimized for your Mac
llm inferencefor youromlxoptimizedmac
https://llama-cpp.com/
Llama.cpp - Run LLM Inference in C/C++
Apr 25, 2026 - Llama.cpp (LLaMA C++) allows you to run efficient Large Language Model Inference in pure C/C++. Download llama.cpp for Windows, Linux and Mac.
llm inferencellamacpprun
https://redis.io/blog/get-faster-llm-inference-and-cheaper-responses-with-lmcache-and-redis/
Get faster LLM inference and cheaper responses with LMCache and Redis | Redis
Developers love Redis. Unlock the full potential of the Redis database with Redis Enterprise and start building blazing fast apps.
get fasterllm inferencecheaperresponseslmcache
https://developer.microsoft.com/it-it/reactor/events/23187/
Deploying and Monitoring LLM Inference Endpoints | Microsoft Reactor
Acquisisci nuove competenze, incontra nuovi colleghi e trova un tutor. Gli eventi virtuali si tengono 24 ore su 24. Unisciti a noi ovunque e in qualsiasi...
llm inferencedeployingmonitoringendpointsmicrosoft
https://www.digitalocean.com/community/conceptual-articles/bottlenecks-llm-inference-optimization
The Hidden Bottlenecks in LLM Inference and How to Fix Them | DigitalOcean
Discover LLM inference bottlenecks like GPU underuse, memory limits, and latency, plus practical strategies to optimize performance and scalability.
how to fixthe hiddenllm inference
https://www.vllm.ch/
vLLM Experts Switzerland – LLM Inference Consulting | VSHN
vLLM consulting and operations in Switzerland. VSHN deploys and manages high-throughput LLM inference on Kubernetes with full Swiss data residency. ISO 27001.
llm inferencevllmexpertsswitzerlandconsulting
https://blogs.oracle.com/cloud-infrastructure/benchmarking-oci-compute-shapes-llm-serving
Effectively benchmarking OCI Compute Shapes for LLM inference serving | cloud-infrastructure
The blog post explores the rapidly evolving landscape of generative AI models and the corresponding maturation of the AI software ecosystem, emphasizing the...
for llmeffectivelybenchmarkingocicompute
https://www.kronkai.com/
Kronk — Hardware Accelerated LLM Inference for Go
Kronk is a Go library for hardware accelerated local LLM inference with llama.cpp. OpenAI-compatible API.
llm inferencekronkhardwareacceleratedgo
https://dinference.com/
DInference - OpenSource LLM Inference API
Open Source LLM Inference. US Hosted. OpenRouter Compatible. Decentralized inference with OpenAI API compatibility.
llm inferenceopensourceapi
https://juicefactory.ai/en
GDPR-Compliant AI API — EU-Hosted LLM Inference | Juice Factory
Stateless, zero-retention AI inference hosted in Sweden. OpenAI-compatible API, GDPR compliant by design. Your data never leaves the EU.
gdpr compliantai apillm inference
https://www.redhat.com/ja/blog/efficient-and-reproducible-llm-inference-red-hat-mlperf-inference-v51-results
Efficient and reproducible LLM inference with Red Hat: MLPerf Inference v5.1 results
As generative AI (gen AI) workloads become central to enterprise applications, benchmarking their inference performance has never been more critical for...
llm inference
https://dmatora.github.io/LLM-inference-speed-benchmarks/
LLM Inference Speeds
llm inferencespeeds
https://resources.nvidia.com/121231541151516515612310231231231/en-us-retail-resource-library/watch-7
Understanding LLM Inference: How AI Generates Words
In the last eighteen months, large language models (LLMs) have become commonplace. For many people, simply being able to use AI chat tools is enough, but for...
llm inferenceunderstandingaiwords
https://developer.nvidia.com/blog/llm-performance-benchmarking-measuring-nvidia-nim-performance-with-genai-perf/
LLM Inference Benchmarking Guide: NVIDIA GenAI-Perf and NIM | NVIDIA Technical Blog
May 29, 2025 - This is the second post in the LLM Benchmarking series, which shows how to use GenAI-Perf to benchmark the Meta Llama 3 model when deployed with NVIDIA NIM.
llm inferencebenchmarkingguidenvidia
https://blogs.vmware.com/cloud-foundation/2024/09/25/llm-inference-sizing-and-performance-guidance/
LLM Inference Sizing and Performance Guidance - VMware Cloud Foundation (VCF) Blog
Feb 20, 2026 - Editorial Note: This blog post was updated on August 4, 2025, to include considerations for GQA in memory consumption and to incorporate additional GPUs and...
vmware cloud foundationllm inferencesizingperformance
https://www.baseten.co/solutions/llms/
LLM Inference for Performance and Scale | Baseten
Ship LLM-powered apps that scale in any cloud. Performant, compliant, and reliable inference for every LLM.
llm inferencefor performancescalebaseten
https://github-danielrosehill.vercel.app/category/llm-inference
llm-inference - Daniel Rosehill GitHub Index
A periodically updating index of GitHub repositories and open source development projects by Daniel Rosehill. Browse by category, date, and topic tags.
llm inferencedaniel rosehillgithubindex
https://redis.io/blog/prefill-vs-decode/
Prefill vs Decode: LLM Inference Phases Explained
Learn how prefill and decode phases affect LLM app speed, what drives TTFT and inter-token latency, and which optimizations fix each bottleneck.
llm inferenceprefillvsdecodephases
https://forums.developer.nvidia.com/t/llm-inference-results/347999
LLM inference results? - Jetson AGX Xavier - NVIDIA Developer Forums
Oct 17, 2025 - Hey, I am considering acquiring a Xavier agx second hand and use it for LLM inference. Does anyone have any benchmarks? The only thing I found so far was in...
jetson agx xavierllm inferencenvidia developerresultsforums
https://ollama.linkworksinc.com/
LLM Monitor · live GPU-cluster inference monitoring
Hourly automated monitoring of a homelab GPU inference cluster — tokens/second, uptime, and incidents across every endpoint. Open methodology, no marketing...
llm monitorgpu clusterliveinferencemonitoring
https://github.com/defilantech/llmkube
GitHub - defilantech/LLMKube: Kubernetes operator for local LLM inference with llama.cpp, vLLM, and...
Kubernetes operator for local LLM inference with llama.cpp, vLLM, and TGI - multi-GPU, autoscaling, air-gapped, production-ready - defilantech/LLMKube
https://github.com/winfunc/deepreasoning
GitHub - winfunc/deepreasoning: A high-performance LLM inference API and Chat UI that integrates...
A high-performance LLM inference API and Chat UI that integrates DeepSeek R1's CoT reasoning traces with Anthropic Claude models. - winfunc/deepreasoning
https://aws.amazon.com/de/about-aws/whats-new/2023/11/amazon-sagemaker-large-model-inference-dlc-tensorrt-llm-support/
Amazon SageMaker launches a new version of Large Model Inference DLC with TensorRT-LLM support
https://arxiv.org/abs/2504.06261
[2504.06261] Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Abstract page for arXiv paper 2504.06261: Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
inferenceparallelllmgenerationvia
https://arxiv.org/html/2604.09718v2
Agentic Compilation: Mitigating the LLM Rerun Crisis for Minimized-Inference-Cost Web Automation
https://www.redhat.com/fr/technically-speaking/distributed-inference-llm-d
Technically Speaking | Inside distributed inference with llm-d
technically speakingdistributed inferenceinsidellm
https://arxiv.org/html/2604.18529v1
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
efficientllmgenerativeinferencevia
https://bitnet.xin/
BitNet.XIN - 1-Bit LLM Tutorials, CPU Inference & Edge AI Guides
Master BitNet and 1-bit large language models with expert tutorials. Learn CPU inference, edge deployment, model architecture, and performance tuning to run...
edge aibitnetxinllm
https://docs.redhat.com/en/documentation/red_hat_ai_inference_server/3.1/html/llm_compressor/integration-with-rhaiis-and-vllm_llm-compressor
Chapter 3. Integration with Red Hat AI Inference Server and vLLM | LLM Compressor | Red Hat AI...
Chapter 3. Integration with Red Hat AI Inference Server and vLLM | LLM Compressor | Red Hat AI Inference Server | 3.1 | Red Hat Documentation
red hat ai inference
https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/kleidiai-on-android-with-mediapipe-and-xnnpack/
Run LLM inference on Android with KleidiAI, MediaPipe, and XNNPACK | Arm Learning Paths
Learn how to run LLM inference on Android devices using MediaPipe with KleidiAI-enhanced Arm i8mm features to benchmark the Gemma 2B model.
https://blog.aks.azure.com/2025/09/12/pair-llmd-and-rag-on-aks
Pair llm-d Inference with KAITO RAG Advanced Search to Enhance your AI Workflows | AKS Engineering...
Sep 12, 2025 - Accelerate AI-driven discovery on Kubernetes with faster insights, greater accuracy, and scalable performance.
https://research.ibm.com/publications/cloud-native-sustainable-llm-inference-in-action
Cloud Native Sustainable LLM Inference in Action for KubeCon EU 2024 - IBM Research
Cloud Native Sustainable LLM Inference in Action for KubeCon EU 2024 by Chen Wang et al.
https://developers.googleblog.com/en/supercharging-llm-inference-on-google-tpus-achieving-3x-speedups-with-diffusion-style-speculative-decoding/
Supercharging LLM inference on Google TPUs: Achieving 3X speedups with diffusion-style speculative...
Researchers at UCSD have achieved a breakthrough in AI serving efficiency by integrating DFlash, a block-diffusion speculative decoding framework, into the...
https://llm-doc.iheaan.com/
Encrypted LLM Experience for Private Inference | HEaaN Encrypted LLM Experience
for privateencryptedllmexperienceinference
https://docs.oracle.com/en/learn/llm-infebench-ocicom/index.html
Set up a Simple LLM Inference Benchmarking System with vLLM on Oracle Cloud Infrastructure Compute
Learn how to set up and run LLM inference benchmarks on Oracle Cloud Infrastructure Compute shapes using virtual large language model (vLLM).
https://developer.nvidia.com/blog/nvidia-tensorrt-llm-supercharges-large-language-model-inference-on-nvidia-h100-gpus/
NVIDIA TensorRT-LLM Supercharges Large Language Model Inference on NVIDIA H100 GPUs | NVIDIA...
Nov 7, 2023 - Large language models (LLMs) offer incredible new capabilities, expanding the frontier of what is possible with AI. However, their large size and unique…
large language modelnvidia tensorrtllm
https://huggingface.co/papers/2502.04416
Paper page - CMoE: Fast Carving of Mixture-of-Experts for Efficient LLM Inference
Join the discussion on this paper page
paper page
https://build.nvidia.com/spark/trt-llm
TRT LLM for Inference | DGX Spark
Install and use TensorRT-LLM on DGX Spark
trtllminferencedgxspark
https://arxiv.org/abs/2404.15420
[2404.15420] XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
Abstract page for arXiv paper 2404.15420: XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
https://arxiv.org/abs/2604.09613
[2604.09613] Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
Abstract page for arXiv paper 2604.09613: Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
https://www.nvidia.com/en-us/on-demand/session/gtc24-dlit61523/
Unlocking Local LLM Inference with Jetson AGX Orin: A Hands-On Lab DLIT61523 | GTC San Jose 2024 |...
Join us for this hands-on lab and learn how to harness the power of large language models (LLMs) and Jetson AGX Orin to build next-generation AI applicatio
https://engineering.fb.com/2026/03/31/ml-applications/meta-adaptive-ranking-model-bending-the-inference-scaling-curve-to-serve-llm-scale-models-for-ads/
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads...
Apr 21, 2026 - Meta continues to lead the industry in utilizing groundbreaking AI Recommendation Systems (RecSys) to deliver better experiences for people, and better results...
https://inference.roboflow.com/workflows/blocks/open_ai_compatible_llm/
OpenAI-Compatible LLM - Roboflow Inference
Open-source computer vision inference server for object detection, segmentation, classification, and foundation models. Deploy on-device or in the cloud.
openaicompatiblellmroboflowinference
https://developer.nvidia.com/blog/achieving-top-inference-performance-with-the-nvidia-h100-tensor-core-gpu-and-nvidia-tensorrt-llm/
Achieving Top Inference Performance with the NVIDIA H100 Tensor Core GPU and NVIDIA TensorRT-LLM |...
Dec 14, 2023 - Best-in-class AI performance requires an efficient parallel computing architecture, a productive tool stack, and deeply optimized algorithms.
https://resources.nvidia.com/en-us-ai-inference-content/nvidia-tensorrt-llm-supercharges-large-language-model-inference-on-nvidia-h100-gpus
NVIDIA TensorRT-LLM Supercharges Large Language Model Inference on NVIDIA H100 GPUs
Read on how TensorRT-LLM significantly enhances the convenience of utilization and expandability through an open-source modular python API that allows for the...
large language modelnvidia tensorrtllm
https://docs.cloud.google.com/kubernetes-engine/docs/best-practices/machine-learning/inference/autoscaling?ref=blog.premai.io
Best practices for autoscaling large language model (LLM) inference workloads with GPUs on Google...
Learn best practices for autoscaling inference large language model (LLM) workloads with GPUs on Google Kubernetes Engine (GKE) with the Horizonal Pod...
https://arxiv.org/html/2404.15420v3
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
xccachecrossattending
https://research.google/blog/generating-synthetic-data-with-differentially-private-llm-inference/?ref=kenpriore.com
Generating synthetic data with differentially private LLM inference
synthetic dataprivate llmgeneratinginference
https://resources.nvidia.com/en-us-ai-inference-cloud-deployments/other2024-aiinference2
Harness the Power of Cloud-Ready AI Inference Solutions and Experience a Step-By-Step Demo of LLM...
Watch a hands-on demonstration of the effortless process of optimizing, deploying, and managing your AI-inferencing solutions within the public cloud...
https://arxiv.org/abs/2404.15420v3
[2404.15420v3] XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
Abstract page for arXiv paper 2404.15420v3: XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
https://arxiv.org/abs/2601.08511
[2601.08511] STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition...
Abstract page for arXiv paper 2601.08511: STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/embed/
Mastering LLM Techniques: Inference Optimization | NVIDIA Technical Blog
inference optimizationmasteringllmtechniquesnvidia