Awesome AI Agent Stack

Local Models, Inference & Hardware

Run, fine-tune and quantize open models on your own hardware.

Serving

Apple Silicon & MLX

On-device, mobile & browser

KV cache, kernels & speculative decoding

Quantization & compression

Training & fine-tuning

Hardware fit & benchmarks