The ggml-org organization on Github develops and supports the ggml machine learning library and related projects.
- https://huggingface.co/ggml-org -
ggml-orgat Hugging Face - https://llama.app -
llama.cpp's official website
graph TD;
ggml --> whisper.cpp
ggml --> llama.cpp
llama.cpp --> coding
llama.cpp --> providers
subgraph coding[Coding]
llama.vim
llama.vscode
llama.qtcreator
end
subgraph providers[Providers]
llama
LlamaBarn
end
ggml[ggml
Machine learning library];
whisper.cpp[whisper.cpp
speech-to-text];
llama.cpp[llama.cpp
LLM inference];
llama.vim[llama.vim
Vim/Neovim plugin];
llama.vscode[llama.vscode
VSCode plugin];
llama.qtcreator[llama.qtcreator
Qt Creator plugin];
llama[llama
CLI app];
LlamaBarn[Llama-macOS
macOS app];
[2026 May 29]llama.app released[2026 May 29]whisper.cpp v1.8.5[2026 Mar 22]VMware Private AI Foundation with NVIDIA 9.0[2026 Mar 19]whisper.cpp v1.8.4[2026 Mar 16]ggml v0.9.8 released[2026 Feb 20]ggml.ai joins Hugging Face 🎉[2026 Jan 28]LlamaBarn v0.24.0 released[2026 Jan 15]whisper.cpp v1.8.3 released
News 2025
[2025 Dec 31]ggml v0.9.5 released[2025 Dec 12]Building zero trust generative AI applications in healthcare with AWS Nitro Enclaves[2025 Oct 28]ggml-org/llama.cpp featured in GitHub's Octoverse 2025 report as Top OSS by contributors[2025 Oct 21]NVIDIA RTX 5090 outperforms AMD and Apple running local OpenAI language models[2025 Sep 18]Latest Open-Source AMD Improvements Allowing For Better Llama.cpp AI Performance Against Windows 11[2025 Sep 09]Llama.cpp Meets Instinct: A New Era of Open-Source AI Acceleration[2025 Aug 19]Firefox 142 Allows Browser Extensions/Add-Ons To Use AI LLMs[2025 Aug 13]FFmpeg 8.0 Merges OpenAI Whisper Filter For Automatic Speech Recognition[2025 Jul 30]MLCommons Releases MLPerf Client v1.0: A New Standard for AI PC and Client LLM Benchmarking[2025 Jul 26]Shotcut 25.07 Video Editor Introduces Speech to Text Model Downloader[2025 Jul 10]Introducing LFM2: The Fastest On-Device Foundation Models on the Market[2025 Jun 26]Introducing Gemma 3n: The developer guide[2025 Jun 23]Running and optimizing small language models on-premises and at the edge[2025 Jun 10]Docker Model Runner adds Qualcomm support[2025 Jun 05]Run small language models cost-efficiently with AWS Graviton and Amazon SageMaker AI[2025 Jun 03]Try out Link Previews in Firefox Labs 138[2025 May 29]Llama.cpp and GGML are optimized for NVIDIA RTX GPUs and the fifth-generation Tensor Cores[2025 May 08]LM Studio Accelerates LLM Performance With NVIDIA GeForce RTX GPUs and CUDA 12.8[2025 Apr 18]Gemma 3 QAT Models: Bringing state-of-the-Art AI to consumer GPUs[2025 Apr 16]Llama 4 Runs on Arm[2025 Apr 04]Run LLMs Locally with Docker[2025 Mar 25]Deploy a Large Language Model (LLM) chatbot with llama.cpp using KleidiAI on Arm servers[2025 Feb 11]OLMoE, meet iOS[2024 Oct 02]Accelerating LLMs with llama.cpp on NVIDIA RTX Systems