Awesome AI Agent Stack
Evaluation, Observability & Code Review
Trace, test and red-team agents, and review the code they write.
Tracing & observability
- langfuse/langfuse - Open-source LLM observability: traces, evals, prompt management; self-hostable.
- Arize-ai/phoenix - OpenTelemetry-based tracing and evaluation for LLM apps.
- traceloop/openllmetry - OpenTelemetry instrumentation for every major LLM library.
- langchain-ai/langsmith-sdk - SDK for LangSmith tracing and datasets.
- raga-ai-hub/RagaAI-Catalyst - Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features.
- evidentlyai/evidently - Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor.
- Helicone/helicone - Open-source LLM observability platform plus AI gateway with caching and cost tracking.
- openlit/openlit - OpenTelemetry-native observability and evaluation for AI and coding agents.
- jaegertracing/jaeger - CNCF distributed tracing platform; the standard trace backend for production systems.
- FailproofAI/failproofai - Observability and enforcement for AI agents in production.
- latitude-dev/latitude-llm - Observability platform for AI agents and LLM apps.
- pydantic/logfire - OpenTelemetry-native observability for AI applications.
- lmnr-ai/lmnr - LLM observability and evaluation platform (Laminar).
- pezzolabs/pezzo - Prompt management and LLM observability toolkit.
- future-agi/future-agi - AI evaluation and observability platform.
- Arize-ai/openinference - OpenTelemetry instrumentation for AI applications.
- Scale3-Labs/langtrace - OpenTelemetry-based LLM observability SDK.
- traceroot-ai/traceroot - Observability for self-improving AI agents.
- databufflabs/databuff - AI-native APM built on OpenTelemetry.
- traceloop/openllmetry-js - Open-source LLM observability for JavaScript.
- evilmartians/agent-prism - Trace visualization for AI agent runs.
- Jwuthri/Tracely-ai - Trace-native CI/CD pipelines for AI agents.
- langfuse/langfuse-js - Langfuse observability SDK for JavaScript.
- langfuse/oss-llmops-stack - Open-source LLMOps stack with observability.
- niklasfrick/spark-dashboard - GPU and LLM inference monitoring dashboard.
- inferock/inferock-bench - LLM cost-tracking proxy with verifiable receipts.
- tma1-ai/tma1 - Local-first observability for AI agents.
- Netis/heron - Agent API performance monitoring via packet probe.
- splunk/token-meter - Dashboard for AI agent token usage and cost.
- theagentplane/tokenops - Token governance for AI agent fleets.
- T-Sunm/rag-ops - LLMOps template for RAG applications.
- llmops-build/llmops - Toolkit for LLM operations.
- dunetrace/dunetrace - Reliability layer for AI agents.
- last9/gpu-telemetry - GPU observability for inference workloads.
- TechNickAI/claude_telemetry - OpenTelemetry wrapper for Claude Code CLI.
- overmind-core/overmind - Platform for continuously improving AI agents.
- alibaba/loongsuite-pilot - Telemetry collector for coding agents.
- slavaZim/episodiq - Economical human-readable logs for agentic trajectories.
- comet-ml/opik - Open-source LLM evaluation and observability platform.
- VTSTech/ACP-Agent-Control-Panel - Lightweight monitoring sidecar and web UI for AI agents.
- SigNoz/signoz - Open-source APM and observability with OpenTelemetry.
- grafana/grafana - Dashboards and alerting for metrics, logs and traces.
- MonishGosar/opentrajectory - Local-first observability for AI agents and coding-agent workflows.
- acailic/agent_debugger - Local-first agent debugger with replay and drift detection.
- KNeegcyao/agent-inspect - DevTools for AI agents: pause, step, and fork counterfactual branches.
- nyktora/noctrace - Network-tab-style waterfall visualizer for Claude Code agent workflows.
- umairinayat/AgentTrace - Open-source observability platform for tracing multi-agent AI systems.
- rajudandigam/agent-inspect - Local-first debugger capturing AI agent execution trees for CI checks.
- raghuece455/AgentMesh - Self-hosted observability, trace replay, and cost analytics for AI agents.
- AgentOps-AI/agentops - Python SDK for AI agent monitoring, cost tracking and benchmarking.
Evaluation
- promptfoo/promptfoo - Test prompts and agents like code; CI-friendly with red-teaming.
- confident-ai/deepeval - Pytest-style unit tests for LLM outputs.
- vibrantlabsai/ragas - Evaluation metrics for RAG pipelines.
- openai/evals - OpenAI's eval framework and registry.
- EleutherAI/lm-evaluation-harness - Standard framework for few-shot LLM evaluation.
- open-compass/opencompass - OpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI.
- UKGovernmentBEIS/inspect_ai - LLM evaluation framework from the UK AI Security Institute with 200+ agentic benchmarks.
- microsoft/promptflow - Microsoft's toolkit for building, tracing, and evaluating LLM flows and agents.
- fivetran/great_expectations - Data validation framework for testing eval datasets and pipeline data quality.
- dataelement/bisheng - Open-source LLM application DevOps platform with evaluation.
- coze-dev/coze-loop - Agent development and optimization platform with evaluation loops.
- microsoft/prompty - Prompt asset format for creating, managing, and evaluating prompts.
- plurai-ai/intellagent - Diagnose AI agents through synthetic interactions.
- JudgmentLabs/judgeval - Continuous improvement stack for AI agent evaluation.
- relari-ai/continuous-eval - Data-driven evaluation for LLM applications.
- rhesis-ai/rhesis - Collaboration layer for reviewing agent behavior.
- alphadl/AdaRubrics - Evaluator for AI agent trajectories.
- AgentEvalHQ/AgentEval - Evaluation toolkit for .NET AI agents.
- google/litmus - LLM testing and evaluation tool with a friendly UI.
- HiThink-Research/GAGE - Unified evaluation engine for LLMs and multimodal models.
- skillberry-ai/cap-evolve - Optimize agent skills against your own evals.
- dokimos-dev/dokimos - LLM and agent evaluation for Java and Kotlin.
- TIGER-AI-Lab/RewardHarness - Self-evolving reward framework for agent evaluation.
- meshkovQA/Eval-ai-library - AI model evaluation framework with 15+ metrics.
- cristianodabc/aludel - LLM evaluation and observability for Elixir.
- paradime-io/dbt-llm-evals - Warehouse-native LLM evaluation for dbt.
- maida-ai/maida - Pre-merge behavioral regression gate for AI agents.
- ProofAgent-ai/proofagent-harness - Test harness for AI agents in CI.
- tolitius/cupel - Discover LLMs punching above their weight.
- cvs-health/uqlm - Python package for uncertainty quantification and hallucination detection in LLMs.
- raindrop-ai/workshop - Lets coding agents write and run agent evals.
- truera/trulens - Evaluation and tracking for LLM experiments and AI agents.
- Agenta-AI/agenta - LLMOps: prompt playground, management, evals and observability.
- Kiln-AI/Kiln - Build, evaluate and optimize AI systems with a desktop UI.
- open-compass/VLMEvalKit - Evaluation toolkit covering 220+ multimodal models and 80+ benchmarks.
- perezjoan/UVLM - Unified Python interface for reproducible vision-language model benchmarking.
- google-research/true - Code and data for re-evaluating factual consistency metrics.
- METR/vivaria - Platform for running agent evaluations and elicitation research.
- langwatch/langwatch - Open platform for LLM evaluations and AI agent testing.
- jameswniu/self-hosted-llm-evals-lab - Self-hosted LLM eval lab with systematic prompt ablation.
- sdivyanshu90/build-your-own-llm-evals - Typed monorepo for reproducible offline LLM and RAG evals.
- strands-agents/evals - Agent evaluation with judges, simulations, traces, and red teaming.
- vostride/agent-qa - Open-source self-improving QA agent for software teams.
- CodeEmperor7/rag-eval - Production-style AI assistant with evaluation and monitoring.
Benchmarks
Security & red-teaming
Fairness & governance
Code review