Khandaq
Sīra Labs · AI security

Khandaq

الخندق

A command post for authorised AI red teaming.

Orchestrate the open-source tools, unify their findings into one model mapped to MITRE ATLAS and the OWASP LLM Top 10, keep a tamper-evident chain of evidence, and watch your AI for drift long after the engagement ends. Self-hosted. Scope-locked. Yours.

Self-hosted Apache-2.0 Scope-locked Signed evidence ATLAS · OWASP · NIST · EU AI Act
The problem

The tools exist. The campaign doesn't.

An AI security professional today runs a dozen excellent, single-purpose open-source tools — garak, PyRIT, promptfoo, ART, model scanners, MCP scanners. Each speaks its own language, writes its own report, and forgets everything the moment it exits.

No shared findings model. garak writes JSONL, promptfoo writes its own JSON, a model scanner writes SARIF if you are lucky. Nothing deduplicates across tools, agrees on severity, or maps a result to a MITRE ATLAS technique and an OWASP LLM ID.

No chain of evidence. The prompts, responses and artefacts that prove a finding scatter across temp directories. Nothing binds them to an engagement, timestamps them, or makes the record tamper-evident for a report a client — or a regulator — will rely on.

No scope discipline. Running offensive tools against the wrong target is a legal event, not a bug. The tools trust whatever host you type. There is no shared, enforced authorisation boundary.

No memory. An engagement is a snapshot. The moment a model, a system prompt or a guardrail config changes, the snapshot is stale — and nothing is watching to tell you a defence that passed last month regressed today.

Khandaq is the layer the tools have been missing: the trench that organises the defence. It doesn't reinvent an attack — it commands the ones that already work, and keeps the record straight.

What Khandaq is

One command post, end to end.

A self-hostable platform that runs an authorised engagement across all ten phases of the offensive-AI lifecycle, then keeps verifying the result.

Engagement & scope

Every run belongs to an engagement with authorised targets and rules of engagement. A hard scope lock refuses anything outside it; an append-only log records who ran what, when.

Tool adapters

garak, PyRIT, promptfoo, DeepTeam, ART, model and MCP scanners — each wrapped in a thin, version-pinned adapter, each isolated in its own container. No dependency wars. Add your own.

Unified findings

Every tool's output lands in one schema — a SARIF superset — with a stable fingerprint, a shared severity, and a mapping to ATLAS, OWASP LLM & Agentic, and NIST. Deduplicated across tools.

Signed evidence ledger

Prompts, responses and artefacts are stored append-only and bound into a hash-chained, optionally Sigstore-signed bundle. The report is reproducible and tamper-evident.

Campaigns & monitoring

Schedule a suite to re-run against a target. Khandaq diffs each run against the last and alerts when a defence regresses or a new finding appears. Your snapshot becomes a watch.

Reports that map

One click turns an engagement into a management summary and a technical report, each cross-walked to OWASP LLM, MITRE ATLAS, NIST AI RMF and the EU AI Act obligations.

How it works

Scope it. Run it. Prove it. Watch it.

  1. Open an engagement

    Declare the authorised targets and the rules of engagement. Khandaq locks the scope and opens an append-only audit log.

  2. Launch the adapters

    Pick a suite — prompt-injection, agent & MCP, model supply-chain, adversarial ML. Each tool runs isolated in its own pinned container.

  3. Collect into one model

    Results are normalised, fingerprinted and deduplicated, then mapped to ATLAS and OWASP. Evidence is written to the signed ledger.

  4. Report & keep watch

    Export the cross-walked report, then schedule the campaign. Khandaq re-runs it and alerts you the day a result gets worse.

Coverage

Built around the ten offensive-AI phases.

Khandaq is organised around the lifecycle an offensive-AI-security professional works through — the same ten phases the EC-Council C|OASP curriculum is built on. For each, it orchestrates the best open-source tools and lands their output in the shared model.

#PhaseOrchestrated toolsRelease
01Methodology & taxonomyMITRE ATLAS, OWASP LLM & Agentic Top 10R1
02Recon & AI attack surfaceAI-BOM generators, inventory & discovery adaptersR2
03Scanning & fuzzinggarak, promptfoo, DeepTeam, GiskardR1
04Prompt injectionPyRIT, promptfoo, garak; AgentDojo via InspectR1
05Adversarial ML & privacyART, TextAttack, ML Privacy MeterR2
06Data & training pipelineART defences, cleanlab, dataset & provenance gatesR2
07Agentic, MCP & A2ACisco mcp-scanner, Snyk agent-scanR1
08Infrastructure & supply chainModelAudit, ModelScan, signing (OMS), CycloneDXR2
09Guardrails & hardeningGuardrail regression harness (NeMo, LlamaFirewall)R2
10IR & forensicsOTel GenAI & Langfuse trace import, timelinesR3

R1 proves the spine — engagement, scope, the findings model, the evidence ledger and the first adapters. R2 widens coverage across the ML and supply-chain phases. R3 adds the forensics and monitoring that nothing open-source does well today.

Architecture

Thin adapters, one spine, sealed evidence.

A Python control plane orchestrates isolated tool containers; a small Rust core canonicalises findings and keeps the hash-chained ledger; results live in Postgres and an append-only object store. Self-hosted on one server with docker compose, or on CapRover.

Web consoleReact SPA khandaq CLIoperators & CI ProxyTLS · OIDC Control plane · FastAPI API · engagementsscope lock · authz · audit Worker · schedulerlaunches adapters · campaigns Rust corenormalise · dedup · ledger Isolated tool adapters garak PyRIT promptfoo ART mcp-scan + yours PostgreSQLfindings · runs · authz Evidence storeappend-only · signedS3 / RustFS Authorised targetin-scope only · the scope lock

Design decisions are recorded as ADRs in the repository. Nothing in the control plane executes an attack itself — it launches the published tools, inside the scope you authorised, and keeps the record.

Speaks the standards

Every finding, cross-walked.

A finding is only useful if it lands where your obligations live. Khandaq maps each one to the frameworks a client, an auditor and a regulator already use.

MITRE ATLAS

Techniques and tactics for attacks on AI systems. Export a Navigator layer straight from an engagement.

OWASP

LLM Top 10 (2025 & 2026) and the Agentic (ASI) Top 10, with both ID sets kept in parallel.

NIST AI RMF

Findings tied to the Govern, Map, Measure and Manage functions for programme-level reporting.

EU AI Act

High-risk obligations referenced so a report shows where a finding touches a compliance duty.

Boundaries

What Khandaq deliberately is not.

Not a new attack. It orchestrates published, maintained tools. It does not ship novel exploits or weaponise anything.
Not for unauthorised use. The scope lock and audit log exist to keep every run inside an engagement you are allowed to perform.
Not a SaaS. There is no hosted control plane. You run it on your own infrastructure; your evidence never leaves it.
Not a replacement for the tools. garak, PyRIT and the rest stay upstream. Khandaq commands them and keeps their findings together.
The name

Khandaq — الخندق, the trench.

يَا أَيُّهَا الَّذِينَ آمَنُوا إِذَا جَاءَكُمْ فَاسِقٌ بِنَبَإٍ فَتَبَيَّنُوا

In the fifth year after the Hijra, a confederate army marched on Medina. On the counsel of Salmān al-Fārisī — an idea carried in from outside the city — the Muslims dug a trench, a khandaq, across the exposed approach. During the same siege the Prophet ﷺ sent Ḥudhayfa ibn al-Yamān alone into the enemy camp at night to learn their condition. The defence of the city turned on seeing the attacker clearly, and on preparing the ground before the attack came.

That is the posture of this project: attacker thinking in the service of defence. You run the offensive tools against your own AI, on ground you have prepared and within a boundary you have drawn, so that the day a real adversary comes, the trench is already dug.

Khandaq joins the Sīra Labs family: Thawr (the cave — a private network, protection), Tabayyun (verifying a report before acting on it), and now Khandaq (preparing the defence by thinking like the attacker).