GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
-
Updated
May 27, 2025 - C++
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Find secrets with Gitleaks 🔑
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
Official inference library for Mistral models
High-speed Large Language Model Serving for Local Deployment
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
A Datacenter Scale Distributed Inference Serving Framework
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Open-source implementation of AlphaEvolve
Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
FlashInfer: Kernel Library for LLM Serving
《深入理解 AI Infra:量化分析与系统设计》(李博杰 著)开源书稿:从硬件约束和模型架构出发,量化推导 LLM 推理与训练系统设计。含全书正文、PDF、配套计算工具与实验
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
To associate your repository with the llm-inference topic, visit your repo's landing page and select "manage topics."