https://cloud.google.com/tpu
Tensor Processing Units (TPUs) | Google Cloud
Google Cloud's Tensor Processing Units (TPUs) are custom-built to help speed up machine learning workloads. Contact Google Cloud today to learn more.
tensor processing unitstpusgooglecloud
https://docs.ray.io/en/latest/cluster/kubernetes/user-guides/tpu.html
Use TPUs with KubeRay โ Ray 2.56.0
This document provides tips on TPU usage with KubeRay. TPUs are available on Google Kubernetes Engine (GKE). To use TPUs with Kubernetes, configure both the...
usetpuskuberay
https://docs.cloud.google.com/ai-hypercomputer/docs/tutorials/tpu/serve-qwen2-7b-instruct
Serve Qwen2-7B-Instruct with vLLM on TPUs | AI Hypercomputer | Google Cloud Documentation
Serve the Qwen2-7B-Instruct on Cloud TPU Trillium using the vLLM serving framework.
https://cloud.google.com/blog/products/ai-machine-learning/bfloat16-the-secret-to-high-performance-on-cloud-tpus?hl=en
BFloat16: The secret to high performance on Cloud TPUs | Google Cloud Blog
How the high performance of Google Cloud TPUs is driven by Brain Floating Point Format, or bfloat16
the secrethigh performanceon cloud
https://docs.cloud.google.com/kubernetes-engine/docs/tutorials/serve-open-models-tpu-terraform
Serve open LLMs on GKE using TPUs with a pre-configured architecture | GKE AI/ML | Google Cloud...
Deploy and serve popular open large language models (LLMs) on GKE with TPUs for inference using a pre-configured, production-ready reference architecture with...
https://tpumlir.org/
TPU-MLIR ยท An Open-Source MLIR-Based Compiler for TPUs
TPU-MLIR is an open-source MLIR-based compiler that converts PyTorch, ONNX, TFLite, Caffe and HuggingFace LLM models into bmodel files running efficiently on...
open sourcetpumlirbasedcompiler
https://docs.cloud.google.com/vertex-ai/docs/predictions/serve-llama3-with-saxml-tpu
Serve Llama 3 open models using multi-host Cloud TPUs on Vertex AI with Saxml | Google Cloud...
Learn about serving Llama 3 open models using multi-host Cloud TPUs with Saxml.
https://docs.cloud.google.com/sdk/gcloud/reference/alpha/compute/tpus/tpu-vm/versions
gcloud alpha compute tpus tpu-vm versions | Google Cloud SDK | Google Cloud Documentation
google cloud sdkgcloudalphacomputetpus
https://www.prnewswire.com/news-releases/anthropic-expands-use-of-google-cloud-and-tpus-302735047.html
Anthropic Expands Use of Google Cloud and TPUs
/PRNewswire/ -- Anthropic today announced an expansion of its use of TPU chips and cloud services, as it scales its development of foundation models, agents,...
google cloudanthropicexpandsusetpus
https://docs.cloud.google.com/sdk/gcloud/reference/compute/tpus/queued-resources/ssh
gcloud compute tpus queued-resources ssh | Google Cloud SDK | Google Cloud Documentation
google cloud sdkgcloudcomputetpusresources
https://www.business-standard.com/technology/tech-news/google-ai-race-strategy-bard-gemini-evolution-tpu-8i-8t-explained-126050500903_1.html
Bard to Gemini and TPUs: Decoding Google's multi-pronged AI strategy | Tech News - Business Standard
Google's latest TPU innovations and the transition to more advanced Gemini models suggest the company may be moving beyond catch-up mode in the fast-evolving...
https://developers.googleblog.com/en/maxtext-expands-post-training-capabilities-introducing-sft-and-rl-on-single-host-tpus/
MaxText Expands Post-Training Capabilities: Introducing SFT and RL on Single-Host TPUs - Google...
MaxText now supports Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) on single-host TPUs. Leverage JAX-based efficiency and advanced algorithms...
https://cloud.google.com/tpu?authuser=0&hl=th
Tensor Processing Units (TPUs) | Google Cloud
Google Cloud's Tensor Processing Units (TPUs) are custom-built to help speed up machine learning workloads. Contact Google Cloud today to learn more.
tensor processing unitstpusgooglecloud
https://cloud.google.com/tpu?ref=maxoxo.me
Tensor Processing Units (TPUs) | Google Cloud
Google Cloud's Tensor Processing Units (TPUs) are custom-built to help speed up machine learning workloads. Contact Google Cloud today to learn more.
tensor processing unitstpusgooglecloud
https://docs.cloud.google.com/sdk/gcloud/reference/alpha/compute/tpus/locations
gcloud alpha compute tpus locations | Google Cloud SDK | Google Cloud Documentation
google cloud sdkgcloudalphacomputetpus
https://developers.googleblog.com/en/supercharging-llm-inference-on-google-tpus-achieving-3x-speedups-with-diffusion-style-speculative-decoding/
Supercharging LLM inference on Google TPUs: Achieving 3X speedups with diffusion-style speculative...
Researchers at UCSD have achieved a breakthrough in AI serving efficiency by integrating DFlash, a block-diffusion speculative decoding framework, into the...
https://cloud.google.com/tpu?ref=mrdbourke.com
Tensor Processing Units (TPUs) | Google Cloud
Google Cloud's Tensor Processing Units (TPUs) are custom-built to help speed up machine learning workloads. Contact Google Cloud today to learn more.
tensor processing unitstpusgooglecloud
https://blog.google/feed/trillium-tpus/
Google Cloud announces Trillium TPUs now available
google cloudannouncestrilliumtpusavailable
https://cloud.google.com/tpu?authuser=00
Tensor Processing Units (TPUs) | Google Cloud
Google Cloud's Tensor Processing Units (TPUs) are custom-built to help speed up machine learning workloads. Contact Google Cloud today to learn more.
tensor processing unitstpusgooglecloud
https://opensource.googleblog.com/2025/12/grl-turning-verifiable-games-into-a-post-training-suite-for-llm-agents-with-tunix-on-tpus.html?m=0
GRL: Turning verifiable games into a post-training suite for LLM agents with Tunix on TPUs | Google...
https://inferencesystemsauthority.com/inference-hardware-accelerators/
Inference Hardware Accelerators: GPUs, TPUs, and Custom Chips
The selection and deployment of hardware accelerators โ graphics processing units GPUs, tensor processing units TPUs, and application-specific custom silicon โ...
hardware acceleratorsinferencegpustpuscustom
https://www.tensorflow.org/guide/tpu?authuser=5
Use TPUs | TensorFlow Core
usetpustensorflowcore
https://cloud.google.com/blog/topics/sustainability/tpus-improved-carbon-efficiency-of-ai-workloads-by-3x?hl=en
TPUs improved carbon-efficiency of AI workloads by 3x | Google Cloud Blog
A new study finds that TPU hardware has seen a 3x improvement in the carbon-efficiency of AI workloads from TPU v4 to Trillium.
carbon efficiencyai workloads
https://docs.cloud.google.com/kubernetes-engine/docs/tutorials/serve-gemma-tpu-jetstream
Serve Gemma using TPUs on GKE with JetStream | Kubernetes Engine | Google Cloud Documentation
For efficient inference serving, deploy and serve Gemma large language models (LLMs) on GKE using TPUs with JetStream and MaxText.
https://docs.cloud.google.com/kubernetes-engine/docs/tutorials/ray-kueue-dws
Optimize AI training on TPUs with DWS and Kueue | GKE AI/ML | Google Cloud Documentation
Learn how to optimize AI training workloads on TPUs using Dynamic Workload Scheduler (DWS) and Kueue in Google Kubernetes Engine (GKE).
https://docs.cloud.google.com/sdk/gcloud/reference/beta/compute/tpus/versions/describe
gcloud beta compute tpus versions describe | Google Cloud SDK | Google Cloud Documentation
google cloud sdkgcloudbetacomputetpus