Awesome AI Agent Stack
Interpretability, Alignment & Research
Understanding and steering model behaviour.
Interpretability
shap/shap
- A game theoretic approach to explain the output of any machine learning model.
TransformerLensOrg/TransformerLens
- A library for mechanistic interpretability of GPT-style language models.
decoderesearch/SAELens
- Train and analyze sparse autoencoders on language models.
openai/transformer-debugger
- Tool for investigating specific behaviors of small language models.
ndif-team/nnsight
- Interpret and manipulate the internals of deep models.
adamkarvonen/SAEBench
- Benchmark suite for evaluating sparse autoencoders on language models.
EleutherAI/sparsify
- Sparse autoencoder training library for interpretability research.
hijohnnylin/neuronpedia
- Platform for exploring and interpreting neuron activations.
jessevig/bertviz
- Attention-head visualization tool for transformer models.
inseq-team/inseq
- Attribution and interpretability toolkit for sequence models.
meta-pytorch/captum
- Model interpretability and attribution library for PyTorch.
PAIR-code/lit
- Learning Interpretability Tool for visual model analysis.
stanfordnlp/pyvene
- Causal interpretability toolkit for neural network internals.
TransformerLensOrg/CircuitsVis
- Interactive visualizations of transformer circuits.
constsynth/loupe
- Research library for SAE-centered interpretability with feature dashboards.
gwenlake/nanoSAE
- Small, fast library for training sparse autoencoders on text-model activations.
ZaheerAbbasOrakzai/transformer-internals-lab
- Interactive suite for attention maps, head ablation, and logit-lens tracing.
mi-for-the-rust-of-us/candle-mi
- Rust toolkit for mechanistic interpretability on the candle ML framework.
jacobgil/pytorch-grad-cam
- Explainability methods for computer vision models, including CNNs and ViTs.
Steering & model editing
stanfordnlp/pyreft
- Python library for representation fine-tuning of LLMs.
zjunlp/EasyEdit
- Easy-to-use knowledge editing framework for LLMs.
binhu02/repsteer
- Python toolkit for representation engineering and activation steering.
levashi/reprobe
- Memory-efficient linear probes and activation steering for safety research.
codemage05/llm-truth-steering
- Linear probing and activation steering to detect truthful LLM representations.
graphs4ai/LLM-Lobotomy
- Framework for analyzing and steering political stance in LLM activations.
smgpulse007/llm-steering
- Local-first starter kit for activation steering on small language models.
p-e-w/heretic
- Automatic censorship removal for language models through abliteration.
RL & post-training
huggingface/trl
- Train transformer language models with reinforcement learning.
modelscope/ms-swift
- Scalable fine-tuning and deployment framework for large models.
NVIDIA-NeMo/RL
- Scalable RL post-training library for large language models.
OpenRLHF/OpenRLHF
- Open-source RLHF framework for training aligned LLMs.
PKU-Alignment/align-anything
- Alignment framework for any modality with human feedback.
PrimeIntellect-ai/prime-rl
- Distributed RL infrastructure for training reasoning models.
verl-project/verl
- Hybrid RL training framework for large language models.
OpenPipe/ART
- Train multi-step agents for real-world tasks with reinforcement learning.
oumi-ai/oumi
- Fine-tune, evaluate and deploy open models with SFT and RL.
transformerlab/transformerlab-app
- Open source research environment for training and evaluating LLMs.
alibaba/ROLL
- Scaling library for reinforcement learning with large language models.
langfengQ/verl-agent
- Extension of veRL for training LLM agents with RL.
Gen-Verse/OpenClaw-RL
- Train any agent with reinforcement learning from conversations.
natolambert/rlhf-book
- Textbook on reinforcement learning from human feedback.
Alignment
PKU-Alignment/safe-rlhf
- Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback.
AI45Lab/OpenART
- Open agentic red teaming toolkit for AI safety evaluation.
AlignmentResearch/defense-in-depth-demo
- Accompanying code and demo for the Defense in Depth safety paper.
AlignmentResearch/safety-gap
- Attacks and evals measuring the safety gap between mitigated and helpful-only LMs.
AlignmentResearch/obfuscation-atlas
- Maps where honesty emerges in RLVR using deception probes.
AlignmentResearch/deception-evasion-honesty
- Code for preference learning with lie detectors inducing honesty or evasion.
AlignmentResearch/puc
- Harness for studying persuasion-under-control by misaligned AI assistants.
AlignmentResearch/AttemptPersuadeEval
- Eval measuring LLM attempts to persuade across benign to harmful topics.
AlignmentResearch/vlmrm
- Uses vision-language models as zero-shot reward models for reinforcement learning.
Research
deepseek-ai/Engram
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models.
ShishirPatil/gorilla
- Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls).
dair-ai/AI-Papers-of-the-Week
- Highlighting the top ML papers every week.
karpathy/nanoGPT
- Minimal, hackable GPT training codebase for LLM training research.
state-spaces/mamba
- Selective state-space sequence models; the leading Transformer alternative.
google-research/google-research
- Code for Google Research papers.
NVIDIA/cosmos
- NVIDIA's open world models, datasets and tools for physical AI.
This site needs JavaScript. The full list is also in the
README on GitHub
.