https://cloud.google.com/blog/products/compute/scaling-moe-inference-with-nvidia-dynamo-on-google-cloud-a4x?e=48754805
Scaling MoE inference with NVIDIA Dynamo on Google Cloud A4X | Google Cloud Blog
A new reference architecture for mixture-of-experts (MoE) workloads uses AI Hypercomputer with A4X machines, NVIDIA GB200 NVL72 and NVIDIA Dynamo.
moe inferenceon googlescalingnvidiadynamo
https://scholar.hit.edu.cn/en/publications/secmoe-communication-efficient-secure-moe-inference-viselect-then/
SecMoE: Communication-Efficient Secure MoE Inference viSelect-Then-Compute - Harbin Institute of...
moe inferencecommunicationefficientsecurecompute
https://gist.github.com/DocShotgun/a02a4c0c0a57e43ff4f038b46ca66ae0
Guide to optimizing inference performance of large MoE models across CPU+GPU using llama.cpp and...
Guide to optimizing inference performance of large MoE models across CPU+GPU using llama.cpp and its derivatives - llamacpp-moe-offload-guide.md
guide tollama cppoptimizinginferenceperformance
https://friendli.ai/models/Flink-ddd/MoE-Pilot-Align-2.7B
Flink-ddd/MoE-Pilot-Align-2.7B - Fast, Reliable, and Scalable Inference on FriendliAI
Run Flink-ddd/MoE-Pilot-Align-2.7B with fast, reliable, and scalable inference on FriendliAI. Get low-latency performance with advanced quantization (FP4, FP8,...
flinkdddmoepilotalign