Awesome AI Agent Stack
Local Models, Inference & Hardware
Run, fine-tune and quantize open models on your own hardware.
Serving
- ollama/ollama - Run models locally with one command; the default for development.
- ggml-org/llama.cpp - CPU/GPU inference in C++; the engine behind most local tooling.
- vllm-project/vllm - High-throughput production serving with paged attention.
- sgl-project/sglang - Fast serving with structured generation and prefix caching.
- lmstudio-ai/lms - CLI for LM Studio's local model server.
- bentoml/OpenLLM - Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
- mudler/LocalAI - LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
- mozilla-ai/llamafile - Distribute and run LLMs with a single file.
- mlc-ai/mlc-llm - Universal LLM Deployment Engine with ML Compilation.
- NVIDIA/TensorRT-LLM - NVIDIA's official optimized inference engine for LLMs on CUDA GPUs.
- kvcache-ai/Mooncake - KVCache-centric disaggregated LLM serving platform behind Moonshot AI's Kimi.
- microsoft/BitNet - Official inference framework for 1-bit LLMs.
- lyogavin/airllm - AirLLM 70B inference with single 4GB GPU.
- FlashML-org/FreeToken - Datacenter-scale model serving on your desktop for massive local models.
- vllm-project/vllm-omni - Framework for efficient inference with omni-modality models.
- Michael-A-Kuykendall/shimmy - Pure-Rust WebGPU inference engine, GGUF-native and OpenAI-API compatible.
- PaddlePaddle/FastDeploy - High-performance inference and deployment toolkit for LLMs and VLMs.
- ikawrakow/ik_llama.cpp - Fork of llama.cpp with state-of-the-art quantization formats and faster inference.
- containers/ramalama - Serve local AI models through familiar container workflows, from any model source.
- sammcj/gollama - Terminal manager for your Ollama models, written in Go.
- sybil-solutions/local-studio - Control panel for local vLLM, SGLang, llama.cpp, and exllamav3 servers.
- intentee/paddler - Open-source LLM and VLM load balancer for self-hosting at scale.
- nobodywho-ooo/nobodywho - Inference engine for running LLMs locally and efficiently on any device.
- turboderp-org/exllamav3 - Optimized quantization and inference library for LLMs on consumer GPUs.
- GradientHQ/parallax - Distributed model serving framework to build your own AI cluster anywhere.
- sgl-project/sglang-omni - High-performance serving framework for audio and unified multimodal models.
- openvinotoolkit/model_server - Scalable inference server for models optimized with OpenVINO.
- foldl/chatllm.cpp - Pure C++ implementation of several models for real-time chat on CPU and GPU.
- PaddlePaddle/Serving - Flexible, high-performance framework for serving machine learning models.
- mosecorg/mosec - High-performance ML model serving with dynamic batching and CPU/GPU pipelines.
- gavamedia/deltafin - Run the full Kimi K3 model on a single device with an OpenAI-compatible server.
- EricLBuehler/candle-vllm - Efficient platform for inference and serving of local LLMs with OpenAI-compatible API.
- ParisNeo/lollms_hub - Proxy server for managing multiple Ollama instances with key-based security.
- brontoguana/krasis - Hybrid LLM runtime for running larger models on VRAM-limited consumer hardware.
- spark-arena/sparkrun - Launch and manage LLM inference workloads on NVIDIA DGX Spark systems.
- toverainc/willow-inference-server - Self-hosted inference server for LLMs, speech-to-text, and text-to-speech.
- mudler/vllm.cpp - Community C++ engine mirroring vLLM with continuous batching and paged KV cache.
- thushan/olla - Lightweight proxy and load balancer for LLM infrastructure with failover.
- mohitsoni48/TurboLLM - Run any local LLM engine auto-tuned to your GPU with web UI and OpenAI-compatible API.
- onnx/turnkeyml - No-code CLI for accelerating ONNX model deployment workflows.
- timtoole02/Camelid - Rust-native local inference backend with evidence-gated model compatibility.
- wladimiravila/esp32s3-distributed-ai - Distributed LLM inference across ESP32-S3 boards, fully offline.
- b4rtaz/distributed-llama - Distributed LLM inference that clusters home devices to run larger models faster.
- xLLM-AI/xllm - High-performance inference engine for LLM, VLM and DiT models on diverse accelerators.
- Luce-Org/lucebox - Speculative LLM inference server tuned for heterogeneous hardware and consumer GPUs.
- syv-ai/HyperQwen - Serves large Qwen models quickly on consumer 24 GB GPUs.
- 0xShug0/audio.cpp - Pure C++ inference engine for audio models, built on ggml.
- shell-nlp/gpt_server - Open framework for production deployment of LLMs, embeddings, and rerankers.
- Tencent-Hunyuan/Hunyuan-A13B - Open fine-grained MoE model.
- kvcache-ai/ktransformers - CPU-GPU heterogeneous inference for giant MoE models on modest hardware.
- Tiiny-AI/PowerInfer - High-speed LLM serving for local deployment on consumer GPUs.
- EricLBuehler/mistral.rs - Fast, flexible LLM inference engine in Rust with an OpenAI-compatible API.
- huggingface/candle - Minimalist ML framework for Rust with CUDA and inference support.
- nomic-ai/gpt4all - Run local LLMs on any device with a desktop app and API.
- abetlen/llama-cpp-python - Python bindings for llama.cpp with an OpenAI-compatible API server.
- tracel-ai/burn - Next-generation deep learning framework for Rust: flexible and portable.
- mostlygeek/llama-swap - Reliably hot-swap models behind any OpenAI-compatible local server.
- exo-explore/exo - Run frontier AI models locally across a cluster of everyday devices.
- antirez/ds4 - Local inference engine for DeepSeek 4 on Metal, CUDA and ROCm.
- Niko1221/Strata - One-click local inference engine for Qwen models with OpenAI and Anthropic APIs.
- magnitudedev/magnitude - Agent inference engine that tunes its kernels to your hardware.
- Neroued/ninfer - High-performance single-GPU inference for selected models.
Apple Silicon & MLX
- ml-explore/mlx - Array framework for machine learning on Apple silicon.
- jundot/omlx - LLM inference server for Apple Silicon with continuous batching and SSD caching.
- ml-explore/mlx-lm - Run large language models on Apple silicon using MLX.
- Arthur-Ficial/apfel - Free on-device AI for Mac: CLI, OpenAI-compatible server, and chat.
- Blaizzy/mlx-vlm - Inference and fine-tuning of vision-language models on Mac with MLX.
- raullenchai/Rapid-MLX - Fast local AI engine for Apple Silicon with tool calling and prompt caching.
- vllm-project/vllm-metal - Community hardware plugin that enables vLLM on Apple Silicon.
- waybarrios/vllm-mlx - OpenAI-compatible LLM inference server for Apple Silicon built on MLX.
- ddalcu/mlx-serve - Native Apple Silicon LLM inference server, OpenAI and Anthropic API compatible.
- ml-explore/mlx-swift-lm - Run LLMs and VLMs on Apple silicon from Swift with MLX.
- SharpAI/SwiftLM - Native MLX Swift inference server for Apple Silicon with SSD streaming.
- ARahim3/mlx-dspark - Up to 4x faster lossless LLM decoding on Apple Silicon via speculative decoding.
- Trans-N-ai/swama - High-performance MLX-based LLM inference engine for macOS in native Swift.
- carloslfu/slotstream - Stream a 105GB mixture-of-experts model from SSD on Macs with 16 to 64GB RAM.
- Epistates/pmetal - High-performance Apple Silicon framework for local LLM inference and serving.
- youssofal/MTPLX - Fast local inference of Qwen models on Apple Silicon Macs using multi-token prediction.
- incoai/splash - Local inference engine for Apple silicon built around the specific model.
- drumih/turbo-fieldfare - Runs Gemma 4 26B-A4B in about 2 GB of RAM on M-series MacBooks.
On-device, mobile & browser
- mlc-ai/web-llm - High-performance In-browser LLM Inference Engine.
- huggingface/transformers.js - State-of-the-art Machine Learning for the web. Run Transformers directly in your browser, with no need for a server!.
- OpenBMB/MiniCPM - Small yet powerful on-device language models for phones and PCs.
- RunanywhereAI/runanywhere-sdks - Production-ready SDKs for running AI models locally on device.
- qualcomm/GenieX - Run frontier LLMs and VLMs locally across Qualcomm NPU, GPU, and CPU.
- cactus-compute/cactus - Quantization, kernels, and inference runtime for phones, wearables, and robots.
- lemonade-sdk/lemonade - Discover and run local AI apps serving optimized LLMs on your GPU or NPU.
- RightNow-AI/picolm - Run a 1-billion-parameter LLM on a $10 board with 256MB of RAM.
- withcatai/node-llama-cpp - Node.js bindings for llama.cpp to run AI models locally with JSON schema output.
- ROCm/FastFlowLM - Run LLMs on AMD Ryzen AI NPUs, purpose-built and deeply optimized.
- john-rocky/CoreML-Models - Core ML model zoo for iOS and macOS with conversion scripts and sample apps.
- callstackincubator/ai - On-device LLM execution in React Native apps with Vercel AI SDK compatibility.
- ngxson/wllama - WebAssembly bindings for llama.cpp enabling in-browser LLM inference.
- OpenSparX/MasterAgent - Build AI agents that run 100 percent on-device with sub-100ms Qualcomm NPU latency.
- kessler/gemma-gem - Run Google's Gemma model entirely on-device via WebGPU, no cloud needed.
- mybigday/llama.rn - React Native bindings for llama.cpp.
- shubham0204/SmolChat-Android - Run any GGUF small language model locally on Android devices.
- tetherto/qvac - Open-source local AI SDK for on-device inference across desktop and mobile.
- Helldez/BigMoeOnEdge - Run mixture-of-experts models bigger than RAM on a 12GB phone, CPU only.
- NotPunchnox/rkllama - Ollama alternative for Rockchip NPUs to run AI models on Rockchip devices.
- NVIDIA/TensorRT-Edge-LLM - Lightweight C++ LLM and VLM inference software for physical AI at the edge.
- zolotukhin/zinc - Zig inference engine for local LLM inference on AMD GPUs and Apple Silicon.
- NightMean/OlliteRT - Turn an Android phone into an OpenAI-compatible local LLM inference server.
- xybrid-ai/xybrid - Cross-platform on-device AI toolkit for phones, desktops, and edge devices.
- dineshsoudagar/local-llms-on-android - Guide to running Gemma, Qwen, and LLaMA locally on Android with LiteRT and ONNX.
- eleiton/ollama-intel-arc - Run Ollama, Stable Diffusion, and Whisper on Intel Arc GPUs.
- Picovoice/picollm - On-device LLM inference powered by X-bit quantization.
- google-ai-edge/LiteRT-LM - Google inference framework for running LLMs on edge and mobile devices.
- software-mansion/react-native-executorch - On-device AI inference library for React Native built on ExecuTorch.
- sauravpanda/BrowserAI - Run local LLMs and speech models directly in the browser via WebGPU.
- SciSharp/LLamaSharp - C# and .NET bindings for llama.cpp to run LLMs locally.
- off-grid-ai/OGAM - Offline mobile AI app for chat, vision, speech and image generation on-device.
- Mobile-Artificial-Intelligence/maid - Free open-source app for running llama.cpp models locally on mobile and desktop.
- AtomicBot-ai/Atomic-Chat - Local AI app and inference engine for running open-weight LLMs privately.
- cactus-compute/needle - Tiny 2-bit tool-calling model for running agents on small devices.
- alibaba/MNN - Blazing-fast lightweight inference engine for on-device and edge AI.
- Tencent/ncnn - High-performance neural network inference optimized for mobile.
- openvinotoolkit/openvino - Optimize and deploy AI inference across Intel hardware.
- google-ai-edge/mediapipe - Cross-platform ML solutions for live and streaming media.
- apple/coremltools - Tools for Core ML model conversion, editing, and validation.
- dusty-nv/jetson-containers - Machine learning containers for NVIDIA Jetson edge deployment.
- pytorch/executorch - On-device AI across mobile and embedded for PyTorch models.
- Edge0-AI/Edge0 - Streaming mixture-of-experts inference with SSD expert offload on any device.
KV cache, kernels & speculative decoding
Quantization & compression
Training & fine-tuning
Hardware fit & benchmarks