Awesome AI Agent Stack
Vision & Multimodal
Image/video understanding and vision-language models.
Vision-language models
- QwenLM/Qwen3-VL - Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
- haotian-liu/LLaVA - Large Language-and-Vision Assistant; open GPT-4V-style model.
- OpenGVLab/InternVL - Open-source multimodal model family rivaling proprietary vision-language models.
- LLaVA-VL/LLaVA-NeXT - Open-source large multimodal models for image and video understanding.
- OpenBMB/MiniCPM-V - Pocket-sized multimodal LLM for efficient image and video understanding on phones.
- NVlabs/Eagle - Frontier vision-language models trained with data-centric strategies.
- NVlabs/VILA - Vision-language models for multimodal AI across edge, data center, and cloud.
- OpenSenseNova/SenseNova-U1 - Native unified multimodal model for understanding, reasoning, and generation.
- m87-labs/moondream - Tiny vision-language model for fast image understanding on any device.
- ATH-MaaS/Ovis - Multimodal LLM architecture structurally aligning visual and textual embeddings.
- om-ai-lab/VLM-R1 - Reinforcement learning framework for visual understanding with VLMs.
- eulogik/TinyDoc-VLM - 256M-parameter document VLM that runs on CPU with ONNX export.
- andreagemelli/baguettotron-vlm - Fully reproducible sub-1B vision-language model with open training.
- microsoft/Magma - Foundation model for multimodal AI agents acting in digital worlds.
- apple-aiml-research/ml-ferret - Apple's region-aware VLM for referring expression grounding.
- jingyaogong/minimind-v - Train a small vision-language model from scratch in a couple of hours.
Detection, segmentation & tracking
Vision libraries & backbones
Face analysis
Video understanding & agents
Depth & 3D reconstruction