8 Best agentic-radar Alternatives in 2026 (Open Source)
agentic-radar — A security scanner for your LLM agentic workflows. The first dedicated security scanner specifically designed for agentic AI workflows, combining static analysis with runtime adversarial testing and automatic prompt hardening — no other tool maps agent vulnerabilities to OWASP AI security frameworks
These 8 open-source tools do the same job. They are ordered by how closely they match agentic-radar, with live GitHub data so you can see which projects are actively maintained.
| Tool | GitHub stars | Stars / 30d | Last commit |
|---|---|---|---|
| agentic-radar(original) | 1.1k | +20 | 2025-11-27 |
| Promptfoo | 25.6k | +1,117 | 2026-09-30 |
| Superagent | 6.8k | +42 | 2026-08-25 |
| LLM Guard | 3.2k | +76 | 2026-07-08 |
| LangKit | 997 | +3 | 2024-11-22 |
| UpTrain | 2.4k | +4 | 2024-07-29 |
| DeepEval | 18.5k | +676 | 2026-09-29 |
| langwatch | 4.9k | +277 | 2026-09-30 |
| Langfuse | 35.2k | +1,821 | 2026-09-30 |
1. Promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, Llama, and more. Simple declarative configs with command line and
What sets it apart: Unlike LangSmith (production observability) or Langfuse (logging), promptfoo is the only open-source tool combining eval + red teaming + CI/CD code scanning — now backed by OpenAI while remaining fully MIT-licensed
Best for: Teams hardening LLM apps against prompt injection and jailbreaks with automated red teaming; Engineering teams adding LLM eval regression tests to CI/CD pipelines
2. Superagent
Superagent protects your AI applications against prompt injections, data leaks, and harmful outputs. Embed safety directly into your app and prove compliance to your customers.
What sets it apart: YC-backed AI safety SDK that pivoted from general agent building to focused safety tooling — provides guard, redact, and scan capabilities with open-weight models for self-hosting, filling the gap between building agents and securing them
Best for: Teams adding safety layers to production AI agents; Enterprises requiring PII redaction and prompt injection protection; Security-focused AI deployments with compliance requirements
3. LLM Guard
The Security Toolkit for LLM Interactions
Best for: Enterprise teams deploying LLMs in production needing security guardrails; Organizations with strict data leakage prevention requirements; Applications handling sensitive user data through LLM interfaces
4. LangKit
🔍 LangKit: An open-source toolkit for monitoring Large Language Models (LLMs). 📚 Extracts signals from prompts & responses, ensuring safety & security. 🛡️ Features include text quality, relevance m
What sets it apart: Open-source text metrics toolkit for LLM monitoring with built-in security detection (jailbreaks, prompt injection), quality scoring, and whylogs integration
Best for: llm-output-monitoring; detecting-prompt-injection; text-quality-observability
5. UpTrain
UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform ro
What sets it apart: vs generic eval tools: 20+ preconfigured evaluations with customizable prompts, few-shot examples, and scenario descriptions — all running locally for data privacy with root cause analysis on failures
Best for: RAG system evaluation and quality assurance; LLM application testing before production deployment; Safety and security testing for prompt injection vulnerabilities
6. DeepEval
The LLM Evaluation Framework
What sets it apart: Most comprehensive open-source LLM eval framework with 30+ research-backed metrics including agentic, RAG, multi-turn, MCP, and multimodal — vs Ragas (RAG-only) or custom eval scripts
Best for: Teams needing comprehensive LLM/agent evaluation pipelines; CI/CD integration for LLM app quality gates; RAG pipeline evaluation and optimization
7. langwatch
The platform for LLM evaluations and AI agent testing
What sets it apart: Unified platform combining agent simulation, evaluation, observability, and prompt optimization with OpenTelemetry-native design — vs separate tools for tracing (Langfuse), eval (DeepEval), and prompt management
Best for: Teams wanting eval + observability + prompt management in one tool; Agent simulation testing before production deployment; Organizations needing OpenTelemetry-native LLM observability
8. Langfuse
🪢 Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
What sets it apart: Unlike LangSmith (LangChain-specific) or Helicone (proxy-based), Langfuse is fully open-source, framework-agnostic, and self-hostable, combining tracing, prompt management, evaluations, and datasets in a single platform built on ClickHouse for scalable production use.
Best for: Teams operating production LLM applications who need tracing, prompt management, and evaluation in one platform; Organizations requiring self-hosted LLM observability for data privacy compliance