Ask what inference costs and you will be quoted a rate per GPU-hour because that is the number infrastructure providers publish. It is close to useless for a business case. Identical hardware serving an identical model can produce costs differing by more than an order of magnitude and the variable driving most of that spread […]