Rare find

llama.cpp. Port of Facebook's LLaMA model in C/C++

github.com/ishandutta2007/llama.cpp

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

June 2026
  • jinja, chat: add --reasoning-preserve flag (#25105)
  • Revert "ui: fix accessibility for hover-gated interactive elements as…
  • ui: fix stop and reasoning skip in single-model mode (#25084)
  • dflash: refactor draft model conversion (#25110)
  • chat : implement minicpm5 parser (#24889)
  • jinja: add --dump-prog for debugging (#25086)
  • spec : add DFlash support (#22105)
  • common : allow --offline in llama download (#25091)
  • logs : reduce v2 (#25078)
  • opencl: flash attention improvement (#25069)
  • [CUDA] Added a cudaMemcpy2DAsync fast path to ggml_cuda_cpy (#25057)
  • sycl : fix failed ut cases of norm (#25044)
  • vulkan: fix step operator for 0 input (#25036)
  • binaries : Improve rpc-server and export-graph-ops names. (#25045)
  • ci : add windows-openvino to check-release (#25022)
  • tests : fix test-chat-template --no-common option (#25075)
  • app : allow --version, --licenses & --help (#25054)
  • sched : reintroduce less synchronizations during split compute (#20793)
  • devops : add llama in all docker images (#25035)
  • arg: fix handling --spec-draft-hf and --hf-repo-v (#25043)
October 2025
  • Release —b6795
  • Release —b6794
  • Release —b6792
  • Release —b6791
  • Release —b6788
  • Release —b6783
  • Release —b6782
  • Release —b6781
  • Release —b6779
  • Release —b6776