Tools
Open source libraries for every stage of AI engineering.
Data
| Tool | Notes |
|---|---|
| Hugging Face Datasets | Large collection of ML datasets with a simple, unified loading API |
| LlamaIndex | Data ingestion, indexing, and query layer for LLM applications |
| Unstructured | Parse and preprocess unstructured documents — PDF, HTML, DOCX, images |
| Docling | Document parsing with layout understanding; strong on PDFs and tables |
| dbt | SQL-based data transformation for analytics engineering |
Evaluation
| Tool | Notes |
|---|---|
| RAGAS | RAG pipeline evaluation — faithfulness, answer relevancy, context recall |
| DeepEval | LLM evaluation framework with 14+ metrics and CI integration |
| Promptfoo | CLI tool for testing, comparing, and red-teaming LLM prompts |
| LangSmith | Tracing, debugging, and evaluation for LLM applications |
| Inspect AI | LLM evaluation framework from AISI; good for safety and capability evals |
Building
| Tool | Notes |
|---|---|
| LangChain | Framework for chaining LLM calls, tools, and retrieval steps |
| LangGraph | Multi-agent orchestration with stateful, cyclic graphs |
| Haystack | End-to-end NLP framework for search, QA, and RAG pipelines |
| DSPy | Programming framework for optimizing LLM pipelines over metrics |
| Instructor | Structured outputs from LLMs using Pydantic; built on top of any OpenAI-compatible API |
| Outlines | Constrained generation — force LLMs to output valid JSON, regex patterns, etc. |
Retrieval
| Tool | Notes |
|---|---|
| pgvector | Vector similarity search extension for PostgreSQL |
| ChromaDB | Embedded vector database; easy to get started locally |
| Qdrant | High-performance vector search engine with filtering |
| Weaviate | Vector database with hybrid (keyword + vector) search |
| Milvus | Scalable vector database built for billion-scale similarity search |
Deploying
| Tool | Notes |
|---|---|
| vLLM | Fast LLM inference and serving with PagedAttention |
| Ollama | Run LLMs locally; wraps llama.cpp with a clean API |
| BentoML | Model packaging and serving framework; good multi-model support |
| Ray Serve | Scalable model serving on top of Ray; handles batching and scaling |
| Triton Inference Server | NVIDIA’s high-throughput inference server; supports TensorRT, ONNX, PyTorch |
Monitoring
| Tool | Notes |
|---|---|
| MLflow | Open source ML lifecycle management — experiment tracking, model registry, serving |
| Weights & Biases | Experiment tracking, hyperparameter sweeps, model versioning |
| Evidently | ML model monitoring and data drift detection |
| Arize Phoenix | LLM observability and evaluation; traces, spans, and evals in one place |
| OpenTelemetry | Vendor-neutral tracing and metrics for distributed systems |