{"ok":true,"engine":"aicost.msp-open-weight-suitability","inputs":{"total_params_b":671,"precision_bits":4,"capacity":"gb_200","weights_shipped":"true","licence":"permissive","data_class":"general","jurisdiction":"CN","is_moe":"true"},"result":{"ok":true,"verdict":"HARDWARE SHORT","verdictReason":"Needs about 419 GB of accelerator memory at 4-bit, roughly 9 professional cards. Available: 200 GB. Mixture-of-experts models pay memory for TOTAL parameters because every expert must be resident, even though only a few activate per token.","decision":"Self-hosting is out of reach at this capacity. The hosted API remains available for this data class.","decisionDetail":"Memory covers WEIGHTS ONLY. KV cache, activations, batching and long context add materially on top. Verify against the model card and your inference framework before committing hardware.","selfHostAvailable":false,"hostedAvailable":true,"selfHostLabel":"Not available","hostedLabel":"Available","anyRouteAvailable":true,"weightMemoryGb":336,"requiredMemoryGb":419,"cards48gbNeeded":9,"precisionBits":4,"firstFailure":"hardware_feasible","gatesPassed":3,"gatesTotal":4,"gates":[{"id":"weights_shipped","passed":true,"detail":"Weights are published and downloadable."},{"id":"licence_permits","passed":true,"detail":"Licence permits commercial use."},{"id":"hardware_feasible","passed":false,"required_gb":419,"available_gb":200,"detail":"Needs about 419 GB of accelerator memory at 4-bit, roughly 9 professional cards. Available: 200 GB. Mixture-of-experts models pay memory for TOTAL parameters because every expert must be resident, even though only a few activate per token."},{"id":"hosted_route_data_class","passed":true,"detail":"The hosted API is acceptable for this data class."}],"memo":"At 4-bit, 671B parameters need about 336 GB for weights alone and roughly 419 GB once runtime overhead is included, which is about 9 professional cards. The first gate to fail is hardware feasible. Needs about 419 GB of accelerator memory at 4-bit, roughly 9 professional cards. Available: 200 GB. Mixture-of-experts models pay memory for TOTAL parameters because every expert must be resident, even though only a few activate per token. Self-host route: not available. Hosted route: available.","questions":["Is the client asking because of a headline, or because they have a workload the current model cannot serve?","Who patches, monitors and re-benchmarks this model once it is self-hosted, and is that billable?","Has anyone read the licence past the word \"open\"?","What data class will actually reach this model in month three, not in the pilot?","If the hosted route is used instead, which jurisdiction processes the request?"],"assumptions":["Memory covers WEIGHTS ONLY at 4-bit. KV cache, activations, batching and long context add materially on top.","Runtime overhead multiplier of 1.25 applied to weight memory *","Card count assumes 48 GB professional cards *","Mixture-of-experts sized on TOTAL parameters, because every expert must be resident."],"caveat":"Memory covers WEIGHTS ONLY. KV cache, activations, batching and long context add materially on top. Verify against the model card and your inference framework before committing hardware.","inputs":{"total_params_b":671,"precision_bits":4,"capacity":"gb_200","weights_shipped":"true","licence":"permissive","data_class":"general","jurisdiction":"CN","is_moe":"true"}}}