{"componentChunkName":"component---src-templates-blog-post-js","path":"/efficiently_serving_llms/","result":{"data":{"site":{"siteMetadata":{"title":"Aparna's Personal Space","author":"Aparna Ravindra"}},"markdownRemark":{"id":"251a83a4-5632-5b91-9618-14ada1842d3b","excerpt":"Course: Efficiently Serving LLMs\nPlatform: DeepLearning.io Deep Learning Text Generation Input text is tokenized into numbers (called tensors). The tensor is…","html":"<p>Course: Efficiently Serving LLMs\nPlatform: DeepLearning.io</p>\n<h2>Deep Learning</h2>\n<h3>Text Generation</h3>\n<ol>\n<li>\n<p><strong>Input text is tokenized</strong> into numbers (called tensors).</p>\n</li>\n<li>\n<p>The tensor is fed to a model that gives the <strong>top-k predictions for the next token</strong>.</p>\n</li>\n<li>\n<p>Take the top prediction and append it to the input, either until the maximum number of tokens are created or a stop token is reached.</p>\n</li>\n<li>\n<p>For a sequence without padding, the attention mask can contain 1 for every token.</p>\n</li>\n<li>\n<p>During attention, the model computes Key (K) and Value (V) tensors. These can be stored in a KV cache for each transformer layer and reused during subsequent token generation.</p>\n</li>\n<li>\n<p>Cache the K-V matrix and pass that to the model, so that only the incremental value can be computed instead of starting computation from the first token while calculating attention.</p>\n<p><em>The K-V cache contains the values that have already been computed and can be reused for subsequent tokens.</em></p>\n</li>\n</ol>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">                    Attention computation\n                         +-----------+\n                         |     A     |\n                         +-----+-----+\n                               |\n              +----------------+----------------+\n              |                                 |\n              v                                 v\n        New input token                    Cached K-V\n              |                         +-----------------+\n              v                         | K1 K2 K3 ...   |\n        +-----------+                   | V1 V2 V3 ...   |\n        |  Q, K, V  |                   +-----------------+\n        | for token |                          |\n        +-----------+                          |\n              |                                 |\n              +----------------+----------------+\n                               |\n                               v\n                         Attention uses:\n                         Q(new) with\n                         K1...Kn + K(new)\n                         V1...Vn + V(new)\n                               |\n                               v\n                       Next-token logits\n                               |\n                               v\n                         Next token\n                               |\n                               |\n                               +------------------+\n                                                  |\n                           K(new), V(new)         |\n                                                  v\n                 +---------------------+      +-------+\n                 |     Updated K-V     | &lt;----| Append|\n                 |                     |      +-------+\n                 | K1 K2 K3 ... Kn+1  |\n                 | V1 V2 V3 ... Vn+1  |\n                 +---------------------+\n                              |\n                              +----&gt; cache for next step</code></pre></div>\n<h3>Prefill and Decode</h3>\n<ol start=\"7\">\n<li>\n<p><strong>Prefill phase</strong> — process the entire input prompt and build the initial KV cache. The output is used to generate the first token. (slowest)</p>\n<p><strong>Decode phase</strong> — generate subsequent tokens one at a time. (can use K-V cache)</p>\n</li>\n</ol>\n<hr>\n<h2>Batching</h2>\n<ol start=\"8\">\n<li>\n<p><strong>Batching:</strong> Parallelly process N sequences of tokens.</p>\n<p>If the number of tokens in each sequence is different, use padding tokens on the left and set the attention mask value to `0`, so that the model does not pay attention to it.</p>\n<p>For example:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">[The] [cat] [sat]\n[The] [home] [on] [the]\n[A]   [boy]</code></pre></div>\n<p>After padding:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">[PAD] [PAD] [The] [cat] [sat]\n[PAD] [The] [home] [on] [the]\n[PAD] [PAD] [PAD] [A] [boy]</code></pre></div>\n<p>The model can then generate the next token for each sequence.</p>\n</li>\n<li>\n<p>Here, in the output from the model, keep all the elements in the batch.</p>\n<p>Next, instead of taking `argmax` of the entire output, take the <strong>max of each row</strong> to choose the next token for that sequence.</p>\n</li>\n<li>\n<p>If you do this efficiently for every loop to see if a sequence is ready to be removed from the batch or a new one can be added, that's continuous batching. This can improve throughput.</p>\n</li>\n</ol>\n<img width=\"1024\" height=\"1536\" alt=\"image\" src=\"https://github.com/user-attachments/assets/a9c784df-f406-43f0-87c8-089dfdc9d151\">","fields":{"slug":"/efficiently_serving_llms/"},"frontmatter":{"title":"Notes from Efficiently Serving LLMs","date":"2026, Sep 26","tags":["tech"],"memorydata":null,"practicedata":null,"img":{"childImageSharp":{"resize":{"src":"/static/e64fde1008cd2041cf38b0bde7482e0e/f3583/efficiently_serving_llms.png","height":675,"width":1200},"gatsbyImageData":{"layout":"fullWidth","placeholder":{"fallback":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsTAAALEwEAmpwYAAACGUlEQVQoz33Q20sbQRTH8f3n24eCTyLVahURLK31UktMxIpNGgUTjFrTmIRoNrubvbrXue3szJwppu+FD+ft+3s4xlMKFtX2gkW0WWiPaSaUEkxJCpL9h+EU7Lbf6w+6j4/X5vQmDf9EdheH91DlUlDBsRIUJAW1uItmMUq1pEbJk9OjlffLb3c2lzZW340fjo+/rl23drUulJJag5JMCaokA0GUIP/KijPKuUFwcXLSWF3fWd/cbbbazdZl/fT8unPz0B92uncftw9939eAFQ0UDYD5UMYpEy4CD4PBGd7YOqg1WnuHZ1s7R5/3G1/2G73b32fnl9/rzfppKw49zSOFbE0cIE6GkV3oOQIHgSErurG9v7S89a32Y+XDp7XNvePahQSYI2UhhUDLKgMykWWMSBET4SJt5zDLFrESbDK2OleDp6E7enzu3T7U6hdOUTYn3q/noGvnk2nqjGauj9yUjCw3yMkk5h07tQtlLP4mgOkKSwCptVSgzRd8NRz1ZmZnbN/99NoHpnk/60ynb65u2rP5wHHb4+EsRwbJBUW8KikoUpbsJRFhAp4rhveZPeSuk7nPgdW3C9+Ok2BgRtM4m4f4qZ8HiBhZCJkPWQBFIsNIhRG8JJAkkCYqT2SJS8Wx5oGm9qvSliKJGfKK1GOpkQWvcepDMAfLBMd6NZ9DFCicC06ZrKiqsCoTxSLBUlVREJQJhHjxF/E+UO0pRXc4AAAAAElFTkSuQmCC"},"images":{"fallback":{"src":"/blog/static/e64fde1008cd2041cf38b0bde7482e0e/56785/efficiently_serving_llms.png","srcSet":"/blog/static/e64fde1008cd2041cf38b0bde7482e0e/0dee1/efficiently_serving_llms.png 750w,\n/blog/static/e64fde1008cd2041cf38b0bde7482e0e/8beaa/efficiently_serving_llms.png 1080w,\n/blog/static/e64fde1008cd2041cf38b0bde7482e0e/a4198/efficiently_serving_llms.png 1366w,\n/blog/static/e64fde1008cd2041cf38b0bde7482e0e/56785/efficiently_serving_llms.png 1672w","sizes":"100vw"},"sources":[{"srcSet":"/blog/static/e64fde1008cd2041cf38b0bde7482e0e/ba386/efficiently_serving_llms.avif 750w,\n/blog/static/e64fde1008cd2041cf38b0bde7482e0e/8adb1/efficiently_serving_llms.avif 1080w,\n/blog/static/e64fde1008cd2041cf38b0bde7482e0e/a3131/efficiently_serving_llms.avif 1366w,\n/blog/static/e64fde1008cd2041cf38b0bde7482e0e/bc8d5/efficiently_serving_llms.avif 1672w","type":"image/avif","sizes":"100vw"},{"srcSet":"/blog/static/e64fde1008cd2041cf38b0bde7482e0e/a66aa/efficiently_serving_llms.webp 750w,\n/blog/static/e64fde1008cd2041cf38b0bde7482e0e/65dd5/efficiently_serving_llms.webp 1080w,\n/blog/static/e64fde1008cd2041cf38b0bde7482e0e/f9724/efficiently_serving_llms.webp 1366w,\n/blog/static/e64fde1008cd2041cf38b0bde7482e0e/52a32/efficiently_serving_llms.webp 1672w","type":"image/webp","sizes":"100vw"}]},"width":1,"height":0.562799043062201}}}}}},"pageContext":{"slug":"/efficiently_serving_llms/","previous":{"fields":{"slug":"/agentic_memory/"},"frontmatter":{"title":"Agentic Memory - Mem0, MemGPT, A-Mem","tags":["tech"],"img":{"childImageSharp":{"gatsbyImageData":{"layout":"fullWidth","placeholder":{"fallback":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsTAAALEwEAmpwYAAACK0lEQVQoz03SSU/bQAAFYP//ayV6qCq1UisOVF2oQkEkJVASQ0II4ITESRzv23gZj2f3klSUS6Xv8C7vnZ4SIR4UDFO2r1graVuxpmKNZLvqRSNZI+lO8layVrKCQrMI1tBf5c4auopb8CXg25RtM75JmB6VQrB9zSgXhPNasrbipUBU4qZ+6TccCRJBmptFoECEhpPZdL6ZaHp/eG97QVNLUAp1k97ZRUrqEAF1q02thZV5FQ6pPc3nfRkuaxwrpIQH7z69efvxvHt5eHT8+KQ1dWXmsjdZXWt2SHYODHv3ake91BO7Kmxs3c+Gp9ia7HGooCI/eH/YOb340emedq/Ozi9zCFO5Pxmv+5oTkV0q4O/l3dFNL6BJXfrEvLHHHWqNdzhSsgR8+f7r50n32/HZ0beTD5+/LvWNEZfXmnm79ELSWlkwfH4YrbR16pZg684GtjaAm1FbBkpZwIvBuPdn1B+Mr9RJrz9Y6puFmw7n1mBmbgE2E199frhdPq4Sp4gNY3oV6KNgocrcVSpO4ySNk9TxAi+IEgDqSlgZf3KLRUjWMXWLaB4aOrAMHMBIx+ZEBrPGnzaZqUhO9ZXh2I7vea7jBL5fCRoX1E5KPytBjhIEvDwIYWQnXpl5HBg43pJQlyhSahTloSsZ3jeyrXlb80ayWuAXHDeCNJJyghDMEcxjAECSlKhgFAtGlDYx9rm5g27D0Ouf/mH/Z8kJI0iwsuKkFqR6HRXkLwQWTuRnKt9JAAAAAElFTkSuQmCC"},"images":{"fallback":{"src":"/blog/static/c641efa396a6e4ec1e5e13a7595af860/56785/agentic_memory.png","srcSet":"/blog/static/c641efa396a6e4ec1e5e13a7595af860/0dee1/agentic_memory.png 750w,\n/blog/static/c641efa396a6e4ec1e5e13a7595af860/8beaa/agentic_memory.png 1080w,\n/blog/static/c641efa396a6e4ec1e5e13a7595af860/a4198/agentic_memory.png 1366w,\n/blog/static/c641efa396a6e4ec1e5e13a7595af860/56785/agentic_memory.png 1672w","sizes":"100vw"},"sources":[{"srcSet":"/blog/static/c641efa396a6e4ec1e5e13a7595af860/a66aa/agentic_memory.webp 750w,\n/blog/static/c641efa396a6e4ec1e5e13a7595af860/65dd5/agentic_memory.webp 1080w,\n/blog/static/c641efa396a6e4ec1e5e13a7595af860/f9724/agentic_memory.webp 1366w,\n/blog/static/c641efa396a6e4ec1e5e13a7595af860/52a32/agentic_memory.webp 1672w","type":"image/webp","sizes":"100vw"}]},"width":1,"height":0.562799043062201}}}}},"next":null}},"staticQueryHashes":["251720178","764694655"]}