Author: Angelo Saraceno House rule: every claim in this post is sourced; if I can’t back something up I cut it rather than handwave. Here is the move most “best AI hosting” lists fumble in the first paragraph: they treat “host an LLM” and “host the app that calls an LLM” as the same problem. They are not. One is a GPU scheduling problem with weights, cold starts measured in tens of seconds, and per-token economics. The other is a long-running web app with a Postgres next to it, a worker pool beh