Awesome AI Agent Stack
Deployment, Serving & MLOps
Ship and operate models and agents in production.
Model serving
- microsoft/onnxruntime - Cross-platform inference engine for serving ML models in production.
- onnx/onnx - Open standard format for portable, interoperable ML models.
- kserve/kserve - Standardized serverless inference platform on Kubernetes for predictive and generative AI.
- triton-inference-server/server - NVIDIA's multi-framework inference server with dynamic batching and ensembles.
- SeldonIO/seldon-core - MLOps framework to package, deploy, monitor, and manage production ML models.
- bentoml/BentoML - Build and serve model inference APIs, job queues, and multi-model pipelines.
- SeldonIO/MLServer - Multi-framework ML inference server with multi-model serving support.
- basetenlabs/truss - The simplest way to package and serve AI/ML models in production.
- tensorflow/serving - Flexible, high-performance serving system for machine learning models.
- replicate/cog - Containers for machine learning: package models reproducibly behind an API.
- vllm-project/production-stack - Reference system for Kubernetes-native, cluster-wide vLLM deployment.
- vllm-project/aibrix - Cost-efficient, pluggable infrastructure components for GenAI inference.
- llm-d/llm-d - Kubernetes-native distributed LLM inference with state-of-the-art performance.
- ogx-ai/ogx - OpenAI-compatible agentic API server; run any model on any infrastructure.
- ai-dynamo/dynamo - Datacenter-scale distributed inference serving framework from NVIDIA.
- xorbitsai/inference - Unified production API to serve LLMs, embeddings, and multimodal models.
- InternLM/lmdeploy - Toolkit for compressing, deploying, and serving LLMs at high throughput.
- lm-sys/FastChat - Open platform for training, serving, and evaluating large language models.
- dphnAI/sonar - Large-scale LLM inference engine, successor of the Aphrodite Engine fork.
- ServerlessLLM/ServerlessLLM - Serverless LLM serving with fast cold starts for everyone.
- predibase/lorax - Multi-LoRA inference server scaling to thousands of fine-tuned adapters.
- ModelTC/LightLLM - Lightweight, easy-to-scale, high-speed Python LLM serving framework.
- OpenNMT/CTranslate2 - Fast inference engine for Transformer models in C++ and Python.
- beam-cloud/beta9 - Ultrafast serverless GPU inference, sandboxes, and background jobs.
- vllm-project/vllm-ascend - Community hardware plugin running vLLM on Huawei Ascend NPUs.
- alibaba/rtp-llm - Alibaba's high-performance LLM inference engine.
MLOps platforms & pipelines
Experiment tracking
- mlflow/mlflow - The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of.
- clearml/clearml - ClearML - Auto-Magical CI/CD to streamline your AI workload. Experiment Management, Data.
- wandb/wandb - The AI developer platform. Use Weights & Biases to train and fine-tune models, and manage models from experimentation to production.
- marimo-team/marimo - A reactive notebook for Python — run reproducible experiments, query with SQL, execute as a.
- kubeflow/katib - Automated hyperparameter tuning and neural architecture search on K8s.
- optuna/optuna - Automatic hyperparameter optimization framework for ML.
- aimhubio/aim - Self-hostable experiment tracker built for high-volume run comparison.
- treeverse/dvclive - Log and track ML metrics, parameters, and models with Git/DVC.
- iterative/gto - Turn any Git repository into an artifact and model registry.
- kubeflow/hub - Central model registry UI for the Kubeflow MLOps lifecycle.
- DagsHub/client - Client libraries for the DagsHub ML experiment tracking platform.
- SwanHubX/SwanLab - Experiment tracking and visualization for AI training.
Kubernetes for AI workloads
Edge & lightweight Kubernetes
Serverless
- knative/serving - Scale-to-zero, request-driven serverless compute on Kubernetes.
- openfaas/faas - Serverless functions made simple, on Kubernetes or faasd.
- nuclio/nuclio - High-performance serverless event and data processing platform.
- fission/fission - Fast and simple serverless functions for Kubernetes.
GitOps, CI & dev environments