Not on main yet I think (as of ~2:50PM UTC on 2026-08-26) – there’s a link to the PR for it in my other comment though. Unsloth’s fork has that integrated (they submitted the PR). I wouldn’t be surprised if something lands quickly in main, but this is a new architecture so may take a bit for people to figure out how to get the most out of it – bunch of discussion about e.g. SSD offloading for the ngrams and stuff like that in the github thread.
Not on main yet I think (as of ~2:50PM UTC on 2026-08-26) – there’s a link to the PR for it in my other comment though. Unsloth’s fork has that integrated (they submitted the PR). I wouldn’t be surprised if something lands quickly in main, but this is a new architecture so may take a bit for people to figure out how to get the most out of it – bunch of discussion about e.g. SSD offloading for the ngrams and stuff like that in the github thread.
llama.cpp support for it got merged a few hours ago 🏆
Nice! Hopefully I can figure out how to actually get it to load tomorrow… (It keeps getting OOM-killed when I try.)