Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
davmre
5 months ago
|
parent
|
context
|
favorite
| on:
Pro Max 5x quota exhausted in 1.5 hours despite mo...
The KV cache consists of activation vectors for every attention head at every layer of the model for every token, so it gets quite large. ChatGPT also estimates 60-100GB for full token context of an Opus-sized model:
https://chatgpt.com/share/69dc5030-268c-83e8-92c2-6cef962dc5...
CraigRood
5 months ago
|
next
[–]
That is actually nuts.... I'm trying to understand the true costs of AI, wonder how I plug this in!
visarga
5 months ago
|
prev
[–]
There are ways to quantize or compress KV cache down.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search:
https://chatgpt.com/share/69dc5030-268c-83e8-92c2-6cef962dc5...