https://engineering.fb.com/2026/03/31/ml-applications/meta-adaptive-ranking-model-bending-the-inference-scaling-curve-to-serve-llm-scale-models-for-ads/
Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads...
Apr 21, 2026 - Meta continues to lead the industry in utilizing groundbreaking AI Recommendation Systems (RecSys) to deliver better experiences for people, and better results...
inference scalingscale modelsmetaadaptiveranking
https://www.eyalo.com/259212/science-news-4/inference-scaling-test-time-compute-why-reasoning-models-raise-your-compute-bill/
Inference Scaling (Test-Time Compute): Why Reasoning Models Raise Your Compute Bill - Drops of...
inference scalingreasoning modelsyour billtesttime
https://arize.com/blog/sleep-time-compute-beyond-inference-scaling-at-test-time/
Sleep-time Compute: Beyond Inference Scaling at Test-time - Arize AI
May 9, 2025 - We summarize a new concept called Sleep-time Compute, a new way to scale AI capabilities: letting models "think" during downtime.
sleep timeinference scalingarize aicomputebeyond
https://www.nec-labs.com/blog/sfs-smarter-code-space-search-improves-llm-inference-scaling/
SFS: Smarter Code Space Search improves LLM Inference Scaling
space searchllm inferencesfssmartercode
https://metavert.io/inference-scaling
Inference Scaling
Inference scaling is the shift from training-dominated AI compute to inference-dominated compute, driven by chain-of-thought reasoning, agentic loops, and...
inference scaling
https://www.redhat.com/de/technically-speaking/scaling-AI-inference
Technically Speaking | Scaling AI inference with open source
Explore the critical role of production-quality AI inference, the power of open source projects like vLLM, and the future of the enterprise AI stack.
technically speakingscaling aiopen sourceinference
https://huggingface.co/papers/2507.02559
Paper page - Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to...
Join the discussion on this paper page
inference time scalingpapertransformersneedremoval
https://tldr.takara.ai/p/2410.05318
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification |...
Despite significant advancements in the general capability of large language models (LLMs), they continue to struggle with consistent and accurate reasoning,...
llm reasoningimprovingscalinginferencecomputation
https://research.ibm.com/publications/rollout-roulette-a-probabilistic-inference-approach-to-inference-time-scaling-of-llms-using-particle-based-monte-carlo-methods
Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using...
Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods for NeurIPS 2025 by Isha Puri...
rolloutrouletteprobabilisticinferenceapproach
https://aidisruption.ai/p/deepseek-unveils-new-paper-on-inference
DeepSeek Unveils New Paper on Inference-Time Scaling, Is R2 Coming?
DeepSeek's new Self-Principled Critique Tuning (SPCT) boosts AI reward models. Is R2 coming? Read the arXiv paper now!
inference time scalingnew paperdeepseekcoming
https://www.zenml.io/llmops-database/scaling-llm-training-and-inference-with-fp8-precision
DeepL: Scaling LLM Training and Inference with FP8 Precision - ZenML LLMOps Database
DeepL needed to scale their Language AI capabilities while maintaining low latency for production inference and handling increasing request volumes. The...
llm trainingllmops databasedeeplscalinginference
https://openreview.net/forum?id=bl884pjXhN
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria...
Modern Large Language Model (LLM) systems typically rely on Retrieval Augmented Generation (RAG) which aims to gather context that is useful for response...
all you needrag systemsrelevancescalinginference
https://cloud.google.com/blog/products/compute/scaling-moe-inference-with-nvidia-dynamo-on-google-cloud-a4x?e=48754805
Scaling MoE inference with NVIDIA Dynamo on Google Cloud A4X | Google Cloud Blog
A new reference architecture for mixture-of-experts (MoE) workloads uses AI Hypercomputer with A4X machines, NVIDIA GB200 NVL72 and NVIDIA Dynamo.
moe inferenceon googlescalingnvidiadynamo