so this is just a go binary that exec's `llama-server --model $MODEL_PATH ...`? there's room for llama wrappers, sure, but the readme only shows features that are already exposed directly from llama-server. it doesn't seem to do anything besides rename the CLI arguments: i don't get it.
I got excited about someone paying attention to intel. Oh well.
In the end i took the sligth slowdown of vulkan to have a more stable and higher development velocity backend While being able to use identical setups on both intel and and gpus.
https://github.com/ggml-org/llama.cpp/blob/master/docs/backe...
Also for the 3 people that ever read this and are curious about local models still, Qwen 27B 3.8 matched Sonnet 5 in the 17 DeepSWE tasks I have run so far, solving the exact same 7 it has. Caveat: datacurve combined low/medium/high Sonnet 5 data.