Artificial Intelligence
Maple-Preview, MiniCPM5-2B or Spark-X2.5, which CPU-only LLM to pick for an arm64 machine
Without a GPU, on an 18-core arm64 CPU with llama.cpp v0.4.1, Maple-Preview generated 172.84 tokens/s on 6 threads: twice MiniCPM5-2B and 3.6 times Spark-X2.5-4B. The trade-off is 5.79 GiB of RAM and reasoning it cannot switch off. MiniCPM5 fits in 3.27 GiB and Spark writes the best Spanish. All three made facts up.