Skip to main content
Loading feed…
I used speculative decoding to make my local LLM feel instant, and now I actually prefer it to cloud APIs · 8 Sync News