Skip to content

Security benchmark · SOC operations

How big is the attack? The security AI says 14. It’s 100

In a SOC the questions are: how big is this, which assets are affected, is it still live, and is this CVE real. We tested the AI approaches that actually ship in security — RAG retrieval, a tool-calling MCP agent, an LLM analyst — against qbrin on exactly those. qbrin is exact, complete, current and reconciled, with zero silent-wrong.

Dated 2026-07-20Self-hosted Wazuh 4.9 + real community MCP server + RAGModel held constant (DeepSeek-V4)
Answer vs. ground truthLive Wazuh data

How big is the attack? · alerts from one attacker

Ground truth100
qbrin100
Wazuh RAG14

The RAG agent saw 14 of 100 — it retrieves only the top-K logs, so it undercounts the attack by 86%.

Which assets to patch? · affected-asset list

Ground truth212
qbrin212
Wazuh RAG10

The RAG agent listed 10 of 212 — it cannot see beyond its retrieval window, so 202 affected assets go unlisted.

The gap, in one chart

RAG can’t count what it can’t retrieve.

Wazuh’s documented AI is retrieval-based: it fetches the top-K relevant logs, then an LLM answers over just those. That’s fine for “find me a suspicious login” — and useless for “how big is this, and which assets are hit.”

−86%

Attack size undercounted

RAG saw 14 of 100 alerts from one attacker. qbrin counted all 100.

Measured

202

Affected assets missed

RAG listed 10 of 212. qbrin listed every one.

Measured

Not one tool — the architecture

The same collapse across the agent ecosystem.

We ran the identical benchmark on three official agents in three unrelated domains — Wazuh (security), Fleet/osquery (endpoint), Netdata (observability). Below is the share of reality each platform’s retrieval AI actually captured on a simple “how many” question. The rest, it never saw.

Security / SIEM

Wazuh

14%

saw 14 of 100 alerts · qbrin: all 100

Endpoint / IT

Fleet + osquery

2.0%

saw 20 of 982 packages · qbrin: all 982

Observability

Netdata

0.8%

saw 20 of 2,520 charts · qbrin: all 2,520

Different products, different data, one root cause: top-K retrieval can’t count, list or verify. qbrin queries deterministically — 100% on every one. The remaining platforms (OpenTelemetry, Zabbix, Velociraptor, Falco, Rudder, GLPI) share the exact architecture, so they share the exact failure.

The full scorecard

Nine SOC tasks, one honest table.

Every row is a real task on live Wazuh data. The AI column shows what the shipping approach actually returned — wins, ties, and the honest boundary, all on the same card.

SOC taskAI approachAI resultqbrinVerdict
Scope the attack
alerts from one attacker · truth 100
RAG (top-K)14 — undercounts100qbrin
Complete affected list
all web-01 level≥10 IDs · truth 212
RAG (top-K)10 — missed 202212qbrin
Is it still live?
alert arrives after indexing
RAG (top-K)stale — “no such user”live · sees itqbrin
Nonexistent entity
alerts for a fake user
RAG (top-K)“none exist” ✓noneTie
Count at scale
~2,000 alerts · truth 1,950
LLM + tool1,950 ✓ ×31,950Tie
Incomplete evidence
a page of alerts silently dropped
LLM + toolundercounts 2/3 (1,850)detects · abstainsqbrin
CVE attribution — naive deploy
58 findings · 9 traps
LLM analyst6.9 silent-wrong / 1000 silent-wrongqbrin
CVE attribution — fully specified
handed the procedure + tools
LLM analyst57/58 · 0 silent-wrong58/58Tie
CVE attribution — real Wazuh MCP agent
community server + LLM
MCP agentCVE dropped → hallucinatedexact + reconciledqbrin

0

qbrin silent-wrong across the entire security domain.

It abstains on missing or inconsistent evidence rather than guessing. The moat is not raw capability — a fully-specified LLM ties qbrin on evidence-internal reasoning. It is determinism, completeness, freshness and abstention by default.

What the shipping AI actually does

Three approaches, tested honestly.

RAG — Wazuh’s documented design

Vectorized logs, top-K retrieval, LLM answers. Built to find relevant logs, not count, list or verify them — so it undercounts the attack, misses affected assets, and a snapshot index goes stale.

100→14 · 212→10 · stale

Real Wazuh MCP agent

14 tools, none for external advisories or inventory. It surfaced findings as “CVE: Unknown” (the CVEs are in the index) — so the LLM hallucinated real-world CVEs. One of two runs failed outright.

no CVE · fabricated · unreliable

LLM analyst — naive deploy

Given a generic “triage this finding” prompt, it silently mis-certified vulnerabilities — flagging patched or wrong-package assets, marking a real vuln safe — on package identity, distro backports and swapped scanner CVEs.

4 silent-wrong / 58

qbrin — verification + reconciliation

Queries the live store deterministically; reconciles finding → asset → installed version → package identity → advisory → verdict; abstains on missing evidence.

exact · complete · current · 0 silent-wrong

What this does — and doesn’t — claim

The honest boundary.

The method behind these numbers is written up in full: how we score the trust layer with 120 trap questions, and what happened when the same gates supervised a live plant loop.

Security teams

Put an exact answer in front of your SOC.

Bring one question your analysts keep re-asking — how big, which assets, is it live. We connect one source, read-only, and show the exact answer with its evidence in a 20-minute walkthrough.

  • One source connected, read-only
  • Your real question answered, with sources
  • Nothing changes in your tools