8 Best LLM-eval-survey Alternatives in 2026 (Open Source)
LLM-eval-survey — The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".. Comprehensive survey and curated collection of LLM evaluation papers and resources organized by what, where, and how to evaluate
These 8 open-source tools do the same job. They are ordered by how closely they match LLM-eval-survey, with live GitHub data so you can see which projects are actively maintained.
| Tool | GitHub stars | Stars / 30d | Last commit |
|---|---|---|---|
| LLM-eval-survey(original) | 1.6k | +3 | 2026-09-13 |
| OpenAI Evals | 19.5k | +231 | 2026-04-14 |
| DeepEval | 18.5k | +676 | 2026-09-29 |
| UpTrain | 2.4k | +4 | 2024-07-29 |
| Ragas | 15.9k | +443 | 2026-02-24 |
| Langfuse | 35.2k | +1,821 | 2026-09-30 |
| langwatch | 4.9k | +277 | 2026-09-30 |
| Agenta | 4.8k | +131 | 2026-09-30 |
| phoenix | 11.7k | +418 | 2026-09-30 |
1. OpenAI Evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Best for: Teams systematically evaluating LLM performance across model versions; Prompt engineers needing no-code YAML-based evaluation workflows; Organizations building quality assurance pipelines for LLM applications
2. DeepEval
The LLM Evaluation Framework
What sets it apart: Most comprehensive open-source LLM eval framework with 30+ research-backed metrics including agentic, RAG, multi-turn, MCP, and multimodal — vs Ragas (RAG-only) or custom eval scripts
Best for: Teams needing comprehensive LLM/agent evaluation pipelines; CI/CD integration for LLM app quality gates; RAG pipeline evaluation and optimization
3. UpTrain
UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform ro
What sets it apart: vs generic eval tools: 20+ preconfigured evaluations with customizable prompts, few-shot examples, and scenario descriptions — all running locally for data privacy with root cause analysis on failures
Best for: RAG system evaluation and quality assurance; LLM application testing before production deployment; Safety and security testing for prompt injection vulnerabilities
4. Ragas
Supercharge Your LLM Application Evaluations 🚀
What sets it apart: vs manual LLM evaluation: Purpose-built evaluation framework with both LLM-based and traditional metrics, automated test generation, and seamless integration with popular LLM frameworks
Best for: Evaluating RAG pipeline quality with automated metrics; Generating comprehensive test datasets for LLM apps; Building continuous evaluation feedback loops
5. Langfuse
🪢 Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
What sets it apart: Unlike LangSmith (LangChain-specific) or Helicone (proxy-based), Langfuse is fully open-source, framework-agnostic, and self-hostable, combining tracing, prompt management, evaluations, and datasets in a single platform built on ClickHouse for scalable production use.
Best for: Teams operating production LLM applications who need tracing, prompt management, and evaluation in one platform; Organizations requiring self-hosted LLM observability for data privacy compliance
6. langwatch
The platform for LLM evaluations and AI agent testing
What sets it apart: Unified platform combining agent simulation, evaluation, observability, and prompt optimization with OpenTelemetry-native design — vs separate tools for tracing (Langfuse), eval (DeepEval), and prompt management
Best for: Teams wanting eval + observability + prompt management in one tool; Agent simulation testing before production deployment; Organizations needing OpenTelemetry-native LLM observability
7. Agenta
The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.
What sets it apart: Unified open-source LLMOps platform combining prompt playground, version control, 20+ evaluators, and OTel-native observability in one tool — vs separate tools for each
Best for: Teams needing integrated prompt management + evaluation + observability; Product teams collaborating with SMEs on prompt engineering; Organizations wanting open-source LLMOps alternative
8. phoenix
AI Observability & Evaluation
What sets it apart: Full-stack AI observability (tracing + eval + datasets + prompt management) in one open-source platform — vs LangSmith which is closed-source and LangChain-specific
Best for: Debugging and monitoring LLM applications in production; Systematic prompt engineering and experiment tracking