Needs about 419 GB of accelerator memory at 4-bit, roughly 9 professional cards. Available: 200 GB. Mixture-of-experts models pay memory for TOTAL parameters because every expert must be resident, even though only a few activate per token.
| Memory needed (GB) | 419 |
| 48 GB cards needed | 9 |
| Weights alone (GB) | 336 |
| Gates passed | 3 |
| Self-host route | Not available |
| Hosted route | Available |
| What to do | Self-hosting is out of reach at this capacity. The hosted API remains available for this data class. |
At 4-bit, 671B parameters need about 336 GB for weights alone and roughly 419 GB once runtime overhead is included, which is about 9 professional cards. The first gate to fail is hardware feasible. Needs about 419 GB of accelerator memory at 4-bit, roughly 9 professional cards. Available: 200 GB. Mixture-of-experts models pay memory for TOTAL parameters because every expert must be resident, even though only a few activate per token. Self-host route: not available. Hosted route: available.