A developer reported running 4-bit GLM 5.2 with DwarfStar mixed RAM and VRAM inference on a DGX Station. The report cites 35 tokens per second generation in a single session.
Need help?
Contact usA developer reported running 4-bit GLM 5.2 with DwarfStar mixed RAM and VRAM inference on a DGX Station. The report cites 35 tokens per second generation in a single session.