Skip to content

Benchmarks

Measured. Sample size on every row.

Measured benchmark results for qbrin against traditional RAG, Glean’s published metrics, mem0, Graphiti and HippoRAG 2 — with the caveat on every row, and the runs where we do not lead.

MEASURED — we ran itMODELED — measured tokens × list pricesPUBLISHED — a competitor’s own number

How much it reads for one answer

A short briefing, not a pile of documents.

qbrin reads ~687 tokens per answer on decision questions. The up-to-20× figure compares qbrin’s lightweight path to a heavy production stack; ~9× is the floor that holds against a lean k=10 baseline.

Input tokens per answer

qbrin~687
RAG · k=10 (lean)~6.5K
RAG · multi-query ×3~18.5K

RAG baselines: k=10 ≈ 6,500; multi-query ×3 / 20-chunk + reranking ≈ 18,500. qbrin’s full-retrieval path measures ~3,081.

687

Input tokens

Measured

~43

Output tokens

Measured

$0.00012

Cost per answer

Comparable model’s list price ($0.15/M in, $0.60/M out).

Modeled

~1.2s

Latency p50

Measured

Measured on real 10-K filings

Grounded in the source, not guessing.

Ask a closed-book LLM about a company’s filings and it invents financial figures about a third of the time. qbrin reads the actual document and answers from it, so it’s right, or it tells you it isn’t sure. The same cited-or-abstain contract is what the REST API returns.

Closed-book LLM vs qbrin’s retrieval-grounded answers on FinanceBench, real public-company 10-K filings, each graded against the source passage. Preliminary sample (n=16).

Made-up figures (confidently wrong) · lower is better

Closed-book LLM31%
qbrin (grounded)0%

Answer accuracy

Closed-book LLM56%
qbrin (grounded)75%

Answered honestly: right, or “I’m not sure”

Closed-book LLM69%
qbrin (grounded)100%

Enterprise RAG Benchmark

qbrin’s own enterprise benchmark on real company Q&A.

500 real company questions (470 answerable + 30 unanswerable), the most product-representative test we run. The flagship result is safety and citations you can trust — not a blanket accuracy crown.

0/281

trap questions answered wrong

Every trap question rejected. MEASURED.

86%

of citations actually hold

A cited source genuinely backs the answer 86% of the time; raw search: 6%. MEASURED.

~87%

decisions correct, held out

The tuned policy scores 88.4% on the bank it was tuned against; ~87.2% transfers to held-out questions, so that is the number we quote.

Measured coverage

Validated across the board.

One story holds across all of it: qbrin has the lowest rate of confident wrong answers and the highest citation-trust, and the lowest cost per answer. Its retrieval is competitive-to-leading once tuned to a corpus, with one architectural exception: relational-graph traversal, where it trails. On raw all-answerable public Q&A it runs a calibration posture — it abstains rather than guess.

19

datasets

Measured

14

competing systems

Measured
Leadsqbrin leads on this axisCompetitiveties or holds its ownTradestrades raw accuracy for calibration: abstains more, by design—not measured
DatasetSafetyCitation-trustCurrent-factRetrievalAnswer accuracyCost / tokensMultilingual
CTP-Bench (temporal/contested) n=12Leads
Temporal vs mem0 (open-source) n=8Leads
Temporal vs Zep / Graphiti n=6Leads
PrecisionMemBench (memory-precision) n=77LeadsLeads
BrainBench (relational retrieval) n=145Trades
Multi-hop vs HippoRAG 2 n=200Leads
MuSiQue (hard multi-hop) n=300Competitive
multihop-rag (retrieval) n=2255LeadsLeads
multihop-rag (answer / unanswerable) n=120 / 301LeadsTrades
ragbench n=100TradesCompetitive
financebench n=150LeadsTrades
Enterprise RAG Benchmark (ERB) n=500LeadsLeadsLeadsTradesLeads
G3 contested enterprise Q&A n=60LeadsCompetitiveLeads
Multilingual (en / hi / te / ta) n=188LeadsCompetitiveLeads
Knowledge-map compression (SOTA) n=80CompetitiveLeads
Hallucination / fake-entity set n=60Leads
Broad RAG general Q&A n=120CompetitiveCompetitive
Internal enterprise benchmark n=40CompetitiveCompetitiveCompetitive
False-premise safety (real / fake) n=55Competitive

Measured, not modeled: each row is a benchmark we ran on qbrin’s own pipeline (competitor figures are their own published numbers or faithful re-runs). Several wins are corpus-specific and several safety A/Bs are honest ties; the suites below carry the exact figures and caveats.

Proof, measured

Built not to make things up.

0

Made-up answers

Across 120 trap questions on four corpora, Qbrin declined, or corrected the false premise with the cited real fact.

86%

Citations you can trust

of the sources Qbrin cites genuinely back the answer, with its double-checking on. Plain keyword search manages 6%.

100%

Always the current fact

Right after a fact changes, Qbrin gives the new value. A popular open-source memory tool gave the old one 3 times in 4.

88%

Finds the right source

of the time Qbrin finds the right source document before it answers.

Numbers come from real benchmark runs on qbrin’s own pipeline: the ERB set of 500 real company questions (470 with a known correct source), plus trap questions built to bait a wrong answer. The full SOC write-up is on the security benchmark page, and the architecture-level differences against Glean, including where Glean is stronger, are in qbrin vs Glean.

Head-to-head

Against named systems, same metric.

vs Glean

A bigger enterprise-context edge than Glean reports

6.8× / 9.7×qbrin · measured
2.0× / 1.6×Glean · published

Answers preferred over GPT-4o / Claude on private company questions (higher is better)

Glean’s headline is enterprise context: its answers are preferred ~2× over ChatGPT and 1.6× over Claude. Measured the same way on private company questions, qbrin’s grounded, cited answers are preferred 6.8× over GPT-4o and 9.7× over Claude-Sonnet-4.5.

Methodology & caveats

MEASURED (qbrin) vs PUBLISHED (Glean). Glean ships no public API or dataset, so its 2.0× / 1.6× are Glean’s own reported figures (Enterprise AI Context Benchmark); this compares the same METRIC TYPE on different data, NOT identical data. qbrin: 42 enterprise questions, LLM-judge preference vs gpt-4o and claude-sonnet-4.5, each given the same graceful-abstain instruction. The corpus is public, so the LLMs may have memorized it, which makes qbrin’s margin conservative. On the answerable questions the raw correctness gap is 23 vs 4 (GPT-4o) / 2 (Claude).

vs Traditional RAG

Won’t make up an answer when there isn’t one

~1%qbrin
~20%traditional RAG

Hallucination rate on unanswerable questions (lower is better)

Asked 301 questions whose answer isn’t in the data, qbrin with its full verification-gate stack fabricated an answer ~1% of the time (3 of 301); traditional and hybrid RAG made one up 18–20% of the time.

Methodology & caveats

MEASURED, N=301 unanswerable multihop-rag questions, MiMo (xiaomi/mimo-v2.5-pro) as the neutral answer model + a strong independent verifier. With the full verification-gate stack: 1.0% (3/301) on a single strong validator, 0.3% (1/301) with dual-validator consensus. qbrin’s abstain-discipline alone measures 5.3% (16/301); trad-RAG 20.3%, hybrid 17.6%, rerank 18.9%, same answer model, same set.

vs mem0

Returns the new fact, not the stale one

100%qbrin
25%mem0 (open-source)

Correct after a fact changes

When you tell the system a fact changed (a project went from active to paused), does it serve the new value? qbrin does every time; mem0 OSS keeps the outdated fact 3 times out of 4.

Methodology & caveats

MEASURED, n=8, same models for both. This is mem0 OPEN-SOURCE (free); mem0 markets temporal reasoning as a PAID feature we did not test. qbrin’s 100% is its CTP-Bench reference.

vs Zep / Graphiti

Reliably retires superseded facts

100%qbrin
~1 / 6Graphiti

Clean fact-invalidations

Against the one rival that genuinely tracks when facts become valid or invalid: Graphiti got it right once in six scenarios and formed no link at all in half of them.

Methodology & caveats

MEASURED, 6 scenarios, edges read directly from the graph store (Graphiti’s own search path errored). Graphiti’s bi-temporal model IS real but inconsistent; better prompting might lift it.

vs HippoRAG 2

Finds chained facts across documents

97.0qbrin
96.0HippoRAG 2

recall@5 (higher is better)

For a question that needs facts chained across documents, how often is the right source in the top 5? On this news corpus qbrin edges out HippoRAG’s graph algorithm.

Methodology & caveats

MEASURED, same 609-doc corpus, 200 questions, same gold. CORPUS-SPECIFIC: multihop-rag is lexically friendly so qbrin’s keyword arm shines. Win is +1.0pt at R@5 and a tie (99.0) at R@10, not a universal retrieval win.

Token economics

Tokens, cost and time per answer.

Tokens read (lower is better)

qbrin687
qbrin full3,081
RAG k=10~6,000
RAG k=12~7,200

On decision questions every answer qbrin reads is a short briefing (~687 tokens). A standard setup that dumps 10–12 raw chunks into the prompt reads 6,000–7,200.

qbrin’s production answer model at $0.14 in / $0.28 out per 1M tokens (list price). Embedding is self-hosted. All dollar figures are MODELED from measured tokens × published prices, not metered invoices.
PathTokensCost / answerType
qbrin, typical answer687$0.00011Measured
qbrin, full stack (heaviest)3,081$0.00046Measured
Traditional RAG, 10 raw chunks6,000$0.00085Modeled
Traditional RAG, 12 raw chunks7,200$0.00102Modeled

How long does an answer take?

We only show timings we actually measured. A typical answer takes ~2–4 seconds, and almost all of that is the model writing the answer, not the search. Finding the source takes ~0.15s.

MEASURED, median (p50) per-query timing on the real pipeline, n=60 each. The 90th-percentile tail runs 6.8–8.5s. We deliberately do not publish competitor seconds we did not measure.
PathFind sourceWrite answerTotal (typical)
Map only (lightest)~0.15s~2.3s2.3s
Hybrid chunks~0.15s~3.1s3.2s
Source selection~0.15s~4.1s4.2s
Full stack (heaviest)~0.15s~4.0s4.0s

Safety & traps

Made-up answers on trap questions.

We feed the system questions deliberately built to bait a wrong answer — made-up names, false premises, missing facts — and count how often it invents an answer instead of saying “I don’t have that”. In the committed clean runs, qbrin invented zero.

0/16

Discriminating traps

Measured

0/30

Fake-entity premises

Measured

0/281

ERB powered negatives

Measured

On a separate set of 301 unanswerable questions, qbrin fabricated an answer 5.3% of the time vs ~18–20% for traditional / hybrid RAG — roughly 3.5× lower false-accept. MEASURED, neutral third-party answer model + independent third-party judge.

Example trap question

“Send the PDF Maya approved in July 2025 regarding the Q3 budget shifts.”

Ordinary AI · makes something up

“Attached is the Q3 Budget PDF showing Maya’s approval…” (actually pulled a June draft from another author)

qbrin · grounded guardrail

“I cannot find a Q3 budget PDF approved by Maya in July 2025. I found a June draft by Sam, and a July thread where Maya requested budget edits, but no final approved PDF.”

Catching subtle wrong answers

Beyond outright invention, a deterministic gate (under a second, no extra AI call) catches sneaky errors.

Error typeVerifier aloneCombined gatePlain meaning
Made-up entity62%100%Every fabricated name caught
Citation contradicts source60%88%Most bad citations caught
Off-by-one / date drift27%86%+59 points over verifier alone
Methodology & caveats

MEASURED, 174-row labeled benchmark, deterministic gate. HONEST LIMIT: this is NOT “zero hallucination” — about 9% of unsupported answers still slip through and ~19% of correct answers are over-flagged until polish lands.

Benchmark suites

Every suite, with its caveats.

Download the full benchmark report (PDF): every suite, methodology, and where we don’t lead.

vs Glean, measured on Glean’s own published metrics

Glean ships no public API or dataset, so we can’t run its system. We run the same metric type Glean has published on qbrin’s own enterprise benchmark and place it next to Glean’s self-reported number. Different data, like-for-like method.

Metricqbrin (measured)Glean (published)Notes
Enterprise-context preference6.8× / 9.7×2.0× / 1.6×Preferred over GPT-4o / Claude-Sonnet-4.5 on private company questions. 42 Qs; single-LLM-judge, fixed order; hardened rerun in progress.
Token efficiency~18×1.3× (23% fewer)qbrin sends a compact knowledge map (~500–800 tokens/query) instead of raw context. Glean reports 23% fewer tokens vs off-the-shelf MCP — a different baseline, shown for context only.
Engine overhead (qbrin only)53 ms—Retrieval + compression overhead before the LLM call, on our test corpus. Not comparable to end-to-end product latency.
Methodology & caveats

We have NO access to Glean’s API or private benchmark: every Glean figure is Glean’s own reported number. The preference study used a single LLM judge with a fixed answer order; a position-swapped, dual-judge rerun is in progress. We removed a writing-quality row and a latency comparison that earlier drafts carried. We would rather show two honest rows than four impressive ones.

ERB — enterprise company Q&A (500 questions)

Whether the right source is found, whether the final decision is correct, and how often the system confidently gives a wrong answer.

SystemAccuracyCoverageConfident wrongRiskNotes
qbrin, full verification~87%87.0%05.3%Zero false-accepts on the trap questions (281/281 rejected). Accuracy is the held-out figure.
qbrin, high-stakes citation mode—43%—precision 86%A cited source backs the answer 86% of the time (raw search: 6%). ~$0.034/query.
Naive RAG (always answers)~83%100%many11.8%High blanket accuracy on a 94%-answerable workload, but confidently wrong 2× as often.
Hybrid search, no double-check—100%—precision 6%Finds the source 88% of the time, but only 6% of citations back the claim.
Methodology & caveats

On a 94%-answerable workload, abstaining COSTS net accuracy vs naive RAG; qbrin’s real win is risk/safety. The verification policy was selected against this same question bank (in-sample 88.4%); ~87.2% transfers to held-out questions, and that is what we publish.

HotpotQA held-out — four systems on real human gold

700 held-out questions with human-labeled gold, plus 500 realistic-name fabricated-entity traps and 150 corrupted-value premises. Four systems, one run each, same day, identical corpus, answer model and embeddings.

SystemPrecision when answeringCoverageMade-up-entity hallucinationNotes
qbrin (full verification pipeline)93.7%74.9%0 / 500Hand-verified: 499 explicit refusals plus 1 response that flagged the name mismatch and answered about the real, cited person.
qbrin (retrieval pinned to k=8)93.4%69.4%0 / 499*Same retrieval budget as the baselines — the safety comes from verification, not a bigger search. (*1 API-error row excluded.)
LlamaIndex (default, top-k 8)84.7%96.4%11 / 5008 direct fabrications + 3 silent substitutions. 0-vs-11 is significant (one-sided p < 0.001).
Naive RAG (top-k 8)80.5%97.3%≥155 / 500Its prompt never invites abstention; treat this as a floor.
Methodology & caveats

MEASURED 2026-07-18/19, N=1,350 per system. Explicit non-claims: qbrin is NOT zero-wrong (93.7% precision = 33 wrong of 524 answered) and does NOT lead on coverage — it abstains on roughly 1 in 4 answerable questions rather than guess. Corrupted-value premise acceptance was 0 for all four systems. The earlier v1 numbers (56%→0%, 62→82) are superseded by this audited run.

Multi-hop retrieval — head-to-head vs HippoRAG

SystemRecall@5Recall@10Notes
qbrin, keyword (BM25) arm97.099.0Edges out HippoRAG on this lexically-friendly news corpus. +1.0pt @5, tie @10.
qbrin, hybrid arm94.098.5Blends keyword + meaning.
qbrin, meaning (vector) arm90.095.0Self-hosted embedding.
HippoRAG 2 (graph algorithm)96.099.0Full-KG, API mode, same 609-doc corpus / 200 questions / same gold.
Methodology & caveats

Corpus-specific: this corpus is lexically friendly, so the keyword arm excels — not a universal retrieval win. On MuSiQue, qbrin’s query-decomposition arm scores 64.4 vs HippoRAG’s published 51.9, but that comparison is not matched.

Knowledge-map compression vs the prior-art field

Answer accuracy (%) on the same real inbox graph, same model, judge, questions and budgets. qbrin is built to win at large budgets and on decision questions — it does not sweep every budget.

BudgetqbrinCommunity-summaryRAPTORGraphRAGVanilla RAG
1,024 tokens10.07.58.81.36.3
2,048 tokens11.312.511.32.57.5
4,096 tokens12.522.516.33.812.5
8,192 tokens33.827.518.85.012.5
Decision / relationship Qs @8,19230.020.022.52.512.5
Methodology & caveats

MEASURED, n=80, bootstrap 95% CIs. qbrin wins at 8,192 and on decision/relationship questions and clearly beats GraphRAG at every budget; community-summary wins the mid-budgets and RAPTOR is competitive. At most budgets the top three have overlapping CIs — n=80 is underpowered to separate them.

Public benchmarks — answer accuracy

DatasetqbrinTraditional RAGNotes
financebench20.0%31.3%Strict independent-judge baseline (pre-upgrade embedder). The 2026 embedder upgrade lifted retrieval +18pp and accuracy +9pp; re-benchmark pending.
multihop-rag72.5%77.5%Strict baseline. The upgrade lifted multi-hop retrieval to 80% gold-recall@8; re-benchmark pending.
Methodology & caveats

All-answerable public sets: qbrin keeps a calibration posture (abstains rather than guess), so it favors precision over coverage. Its safety edge is untestable on all-gold corpora.

How we measure

How to read these numbers.

Every figure is labelled MEASURED, MODELED or PUBLISHED. We never invent a per-suite second or accuracy. Where a comparison isn’t strictly like-for-like, or a sample is small, the footnote says so.

Quality

  • Overall accuracy
  • Answered accuracy
  • Exact match / task score
  • Coverage

Safety

  • False accept rate
  • Hallucination rate
  • Abstain rate
  • Citation validity

Economics

  • Input tokens
  • Output tokens
  • Total tokens
  • Latency p50 / p95
  • Cost per correct answer

Diagnosis

  • Retrieval contains gold?
  • Verifier rejected?
  • Model failed despite gold?
  • Wrong entity / wrong period?
  • Wrong source?

See it answer

See it answer your hardest question.

Bring one real question your team keeps re-asking. We’ll connect a source, read-only, and show you the answer, sourced, in seconds, in a 20-minute walkthrough. Nothing changes in your tools.

  • One source connected, read-only
  • Your real question answered, with sources
  • Nothing changes in your tools