Geek Out Time: Simulating High-Bandwidth and Cache-Like Memory on Google Colab’s T4 GPU
Nov 2024·~95 words in full
This week, we’re delving into the GPU memory hierarchies and exploring how different types of memory impact the performance of transformer models like GPT, BERT, and T5. While Google Colab’s T4 GPU doesn’t provide direct access to High Bandwidth Memory (HBM) or detailed cache controls, we can simulate aspects of these memory behaviors to understand their effects on computational tasks.
Before jumping into our experiments, let’s explore how CPUs, GPUs, and various types of memory contribute to the efficiency of transformer models.
This is an excerpt — the full article continues on Medium.
Read the full article on Medium →Related Posts
- Geek Out Time: Simulating LLM Short and Long Memory with FAISS, LangChain, and Google ColabJun 2025
- Geek Out Time: Simulating Distributed Training on TPU & GPU in Google ColabFeb 2025
- Geek Out Time: From Ontology to AI Agents -Determinism Matters in High-Compliance IndustriesJan 2026
- Geek Out Time: Testing TPU vs GPU on Google Colab After Meta’s Reported Shift Toward TPUsNov 2025