8 Best agentic-radar Alternatives in 2026 (Open Source)

agentic-radar — A security scanner for your LLM agentic workflows. The first dedicated security scanner specifically designed for agentic AI workflows, combining static analysis with runtime adversarial testing and automatic prompt hardening — no other tool maps agent vulnerabilities to OWASP AI security frameworks

These 8 open-source tools do the same job. They are ordered by how closely they match agentic-radar, with live GitHub data so you can see which projects are actively maintained.

ToolGitHub starsStars / 30dLast commit
agentic-radar(original)1.1k+202025-11-27
Promptfoo25.6k+1,1172026-09-30
Superagent6.8k+422026-08-25
LLM Guard3.2k+762026-07-08
LangKit997+32024-11-22
UpTrain2.4k+42024-07-29
DeepEval18.5k+6762026-09-29
langwatch4.9k+2772026-09-30
Langfuse35.2k+1,8212026-09-30
  1. 1. Promptfoo

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, Llama, and more. Simple declarative configs with command line and

    What sets it apart: Unlike LangSmith (production observability) or Langfuse (logging), promptfoo is the only open-source tool combining eval + red teaming + CI/CD code scanning — now backed by OpenAI while remaining fully MIT-licensed

    Best for: Teams hardening LLM apps against prompt injection and jailbreaks with automated red teaming; Engineering teams adding LLM eval regression tests to CI/CD pipelines

  2. 2. Superagent

    Superagent protects your AI applications against prompt injections, data leaks, and harmful outputs. Embed safety directly into your app and prove compliance to your customers.

    What sets it apart: YC-backed AI safety SDK that pivoted from general agent building to focused safety tooling — provides guard, redact, and scan capabilities with open-weight models for self-hosting, filling the gap between building agents and securing them

    Best for: Teams adding safety layers to production AI agents; Enterprises requiring PII redaction and prompt injection protection; Security-focused AI deployments with compliance requirements

  3. 3. LLM Guard

    The Security Toolkit for LLM Interactions

    Best for: Enterprise teams deploying LLMs in production needing security guardrails; Organizations with strict data leakage prevention requirements; Applications handling sensitive user data through LLM interfaces

  4. 4. LangKit

    🔍 LangKit: An open-source toolkit for monitoring Large Language Models (LLMs). 📚 Extracts signals from prompts & responses, ensuring safety & security. 🛡️ Features include text quality, relevance m

    What sets it apart: Open-source text metrics toolkit for LLM monitoring with built-in security detection (jailbreaks, prompt injection), quality scoring, and whylogs integration

    Best for: llm-output-monitoring; detecting-prompt-injection; text-quality-observability

  5. 5. UpTrain

    UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform ro

    What sets it apart: vs generic eval tools: 20+ preconfigured evaluations with customizable prompts, few-shot examples, and scenario descriptions — all running locally for data privacy with root cause analysis on failures

    Best for: RAG system evaluation and quality assurance; LLM application testing before production deployment; Safety and security testing for prompt injection vulnerabilities

  6. 6. DeepEval

    The LLM Evaluation Framework

    What sets it apart: Most comprehensive open-source LLM eval framework with 30+ research-backed metrics including agentic, RAG, multi-turn, MCP, and multimodal — vs Ragas (RAG-only) or custom eval scripts

    Best for: Teams needing comprehensive LLM/agent evaluation pipelines; CI/CD integration for LLM app quality gates; RAG pipeline evaluation and optimization

  7. 7. langwatch

    The platform for LLM evaluations and AI agent testing

    What sets it apart: Unified platform combining agent simulation, evaluation, observability, and prompt optimization with OpenTelemetry-native design — vs separate tools for tracing (Langfuse), eval (DeepEval), and prompt management

    Best for: Teams wanting eval + observability + prompt management in one tool; Agent simulation testing before production deployment; Organizations needing OpenTelemetry-native LLM observability

  8. 8. Langfuse

    🪢 Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more. 🍊YC W23

    What sets it apart: Unlike LangSmith (LangChain-specific) or Helicone (proxy-based), Langfuse is fully open-source, framework-agnostic, and self-hostable, combining tracing, prompt management, evaluations, and datasets in a single platform built on ClickHouse for scalable production use.

    Best for: Teams operating production LLM applications who need tracing, prompt management, and evaluation in one platform; Organizations requiring self-hosted LLM observability for data privacy compliance