Liquid AI’s new 280M draft model speeds up its vision model up to 3.13x — but only decoding, and only free under a $10M revenue cap.
Tag: Hugging Face
Hugging Face’s Transformers library can now load and run llama.cpp’s GGUF quantized models directly, matching its speed on Apple Silicon.