Squeezing 2.7x More Tokens Per Second Out of Two Mismatched GPUs
Or: how I spent a weekend benchmarking local LLM inference, caught my optimization guide lying to me four times, and …
Or: how I spent a weekend benchmarking local LLM inference, caught my optimization guide lying to me four times, and …