Skip to content

How much it reads for one answer

qbrin reads ~687 tokens per answer, a short briefing, not a pile of documents

up to20×fewer tokens
per answer

vs a heavy production RAG pipeline (multi-query ×3 / 20-chunk + reranking)  ·  ~8× vs even a lean 10-chunk RAG call

qbrin~687
RAG · k=10 (lean)~6,500
RAG · multi-query ×3~18,500
Input tokens687 inMEASURED
Output tokens~43 outMEASURED
Cost / answer~$0.00012MODELED
Latency p50~1.2sMEASURED
qbrin measured at ~687 in / ~43 out tokens on decision questions (the full-retrieval path measures ~3,081); RAG baselines modeled (k=10 ≈ 6,500; multi-query ×3 ≈ 18,500). Cost on a comparable model’s list price ($0.15/M in, $0.60/M out). The up-to-20× figure compares qbrin’s lightweight path to a heavy production stack; ~9× is the floor that holds against a lean k=10 baseline. Latency is the same token saving in time, not a separate multiplier.Separately, with its full verification-gate stack qbrin drops the false-accept rate on unanswerable questions to ~1% (3 of 301), vs ~18-20% for traditional RAG. It abstains instead of guessing. A safety benefit, not a token multiplier.
Measured on real 10-K filings

Grounded in the source, not guessing.

Ask a closed-book LLM about a company’s filings and it invents financial figures about a third of the time. qbrin reads the actual document and answers from it, so it’s right, or it tells you it isn’t sure. The same cited-or-abstain contract is what the REST API returns.

Closed-book LLMqbrin (grounded)
Made-up figures (confidently wrong)lower is better
31%
0%
Answer accuracy
56%
75%
Answered honestly, right, or “I’m not sure”
69%
100%

Closed-book LLM vs qbrin’s retrieval-grounded answers on FinanceBench, real public-company 10-K filings, each graded against the source passage. Preliminary sample (n=16).

Enterprise RAG Benchmark

qbrin’s own enterprise benchmark on real company Q&A.

500 real company questions (470 answerable + 30 unanswerable), the most product-representative test we run. The flagship result is never confidently wrong plus citations you can trust, not a blanket accuracy crown.

0 / 281trap questions answered wrongNever confidently wrong, every trap question rejected.MEASURED
86%of citations actually holdA cited source genuinely backs the answer 86% of the time, raw search: 6%.MEASURED
~87%decisions correct, held outOn 500 real company questions, with full verification. The tuned policy scores 88.4% on the bank it was tuned against; ~87.2% is what transfers to held-out questions, so that is the number we quote.MEASURED

Honest read: the real win here is risk / safety and citation-trust, not blanket accuracy. On this 94%-answerable workload a naive RAG that always answers scores ~83% on the answerable subset, qbrin’s edge is that it’s never confidently wrong and its citations hold.

Measured coverage

Validated across the board.

19datasets
14competing systems
7dimensions

One story holds across all of it: qbrin is the most trustworthy (lowest rate of confident wrong answers, highest citation-trust) and the cheapest per answer, and its retrieval is competitive-to-leading once tuned to a corpus, with one architectural exception: relational-graph traversal (a rival’s home turf), where it trails. On raw all-answerable public Q&A it runs a calibration posture (abstains rather than guess); the 2026 embedder upgrade lifted that retrieval substantially (FinanceBench gold-in-context +18pp same-harness). Every cell below is a real run.

Leads, qbrin leads on this axis, no competitor beats itCompetitive, ties or holds its ownTrades for calibration, abstains more, trails raw accuracy, by design not measured on that axis
DatasetSafetyfewer made-up answersCitation-trustcited source actually holdsCurrent-factretires stale / contestedRetrievalfinds the right sourceAnswer accuracycorrect when it answersCost / tokenstokens read per answerMultilingualparity across languages
CTP-Bench (temporal/contested)n=12Leads
Temporal vs mem0 (open-source)n=8Leads
Temporal vs Zep / Graphitin=6Leads
PrecisionMemBench (memory-precision)n=77LeadsLeads
BrainBench (relational retrieval)n=145Trades for calibration
Multi-hop vs HippoRAG 2n=200Leads
MuSiQue (hard multi-hop)n=300Competitive
multihop-rag (retrieval)n=2255LeadsLeads
multihop-rag (answer / unanswerable)n=120 / 301LeadsTrades for calibration
ragbenchn=100Trades for calibrationCompetitive
financebenchn=150LeadsTrades for calibration
Enterprise RAG Benchmark (ERB)n=500LeadsLeadsLeadsTrades for calibrationLeads
G3 contested enterprise Q&An=60LeadsCompetitiveLeads
Multilingual (en / hi / te / ta)n=188LeadsCompetitiveLeads
Knowledge-map compression (SOTA)n=80CompetitiveLeads
Hallucination / fake-entity setn=60Leads
Broad RAG general Q&An=120CompetitiveCompetitive
Internal enterprise benchmarkn=40CompetitiveCompetitiveCompetitive
False-premise safety (real / fake)n=55Competitive
CTP-Bench (temporal/contested)n=12
Current-factLeads
Temporal vs mem0 (open-source)n=8
Current-factLeads
Temporal vs Zep / Graphitin=6
Current-factLeads
PrecisionMemBench (memory-precision)n=77
SafetyLeadsCitation-trustLeads
BrainBench (relational retrieval)n=145
RetrievalCalibration
Multi-hop vs HippoRAG 2n=200
RetrievalLeads
MuSiQue (hard multi-hop)n=300
RetrievalCompetitive
multihop-rag (retrieval)n=2255
RetrievalLeadsCost / tokensLeads

Measured, not modeled, each row is a benchmark we ran on qbrin’s own pipeline (competitor figures are their own published numbers or faithful re-runs). Several wins are corpus-specific and several safety A/Bs are honest ties; the per-suite tabs below carry the exact figures and caveats. The honest headline: no competitor matches qbrin on trust, safety-under-abstention and cost, and its retrieval is competitive-to-leading once tuned to a corpus (except relational-graph traversal, where it trails), on raw answerable-Q&A accuracy it runs a calibration posture (trades coverage for caution), and the 2026 embedder upgrade lifted that retrieval substantially.

Proof, measured

Built to never make things up.

0Made-up answersAcross 120 trap questions on four corpora — invented projects, false dates, wrong values — qbrin invented zero: it declines, or corrects the false premise with the cited real fact. (Measured 2026-07: 0 invented / 120 traps, groundedness-audited.)
0%Citations you can trustWith its double-checking layer on, a cited source genuinely backs the answer 86% of the time, raw keyword search manages just 6%. (Measured, ERB.)
0%Always the current factWhen a fact changes, qbrin returns the new value every time; mem0 (open-source) keeps the stale one 3 times in 4. (Measured, n=8.)
0%Finds the right sourceRetrieval lands the correct source document in the candidate set 88% of the time, and the double-checking layer then keeps the cited ones honest. (Measured, ERB.)

Numbers come from real benchmark runs on qbrin’s own pipeline, the ERB set of 500 real company questions (470 with a known correct source), plus trap questions built to bait a wrong answer. Each figure below is labelled MEASURED (we ran it), MODELED (computed from measured tokens × published prices), or PUBLISHED (a competitor’s own number). No per-suite second-by-second timings are shown unless we measured them. The full SOC write-up, qbrin against a Wazuh RAG pipeline, an MCP agent and a naive LLM analyst, is on the security benchmark page, and the architecture-level differences against Glean, including where Glean is stronger, are set out in qbrin vs Glean.

vs Glean

A bigger enterprise-context edge than Glean reports

6.8× / 9.7×qbrin
2.0× / 1.6×Glean (published)
answers preferred over GPT-4o / Claude on private company questions (higher is better)

Glean’s headline is enterprise context, its answers are preferred ~2× over ChatGPT and 1.6× over Claude. Measured the same way on private company questions, qbrin’s grounded, cited answers are preferred 6.8× over GPT-4o and 9.7× over Claude-Sonnet-4.5.

Methodology & caveatsMEASURED (qbrin) vs PUBLISHED (Glean). Glean ships no public API or dataset, so its 2.0× / 1.6× are Glean’s own reported figures (Enterprise AI Context Benchmark) and this compares the same METRIC TYPE on different data, NOT identical data. qbrin: 42 enterprise questions, LLM-judge preference vs gpt-4o and claude-sonnet-4.5 (flagship models, each given the same graceful-abstain instruction). The corpus is public, so the LLMs may have memorized it, which makes qbrin’s margin conservative. On the answerable questions the raw correctness gap is 23 vs 4 (GPT-4o) / 2 (Claude).
vs Traditional RAG

Won't make up an answer when there isn't one

~1%qbrin
~20%traditional RAG
hallucination rate on unanswerable questions (lower is better)

Asked 301 questions whose answer ISN'T in the data, qbrin, with its full verification-gate stack on, fabricated an answer only ~1% of the time (3 of 301); traditional and hybrid RAG made one up 18-20% of the time. qbrin says “I don't have enough information” instead of guessing.

Methodology & caveatsMEASURED, N=301 unanswerable multihop-rag questions, MiMo (xiaomi/mimo-v2.5-pro) as the neutral answer model + a strong independent verifier. WITH the full verification-gate stack (claim-token + premise-guard + citation-support verifier): 1.0% (3/301) on a single strong validator, 0.3% (1/301) with dual-validator consensus. qbrin’s abstain-discipline alone measures 5.3% (16/301); trad-RAG 20.3%, hybrid 17.6%, rerank 18.9%, same answer model, same set, self-consistent cross-arm test.
vs mem0

Returns the new fact, not the stale one

100%qbrin
25%mem0 (open-source)
correct after a fact changes

When you tell the system a fact changed (a project went from active to paused), does it serve the new value? qbrin does every time; mem0 OSS keeps the outdated fact 3 times out of 4.

Methodology & caveatsMEASURED, n=8, same models for both. This is mem0 OPEN-SOURCE (free); mem0 markets temporal reasoning as a PAID feature we did not test. qbrin’s 100% is its CTP-Bench reference.
vs Zep / Graphiti

Reliably retires superseded facts

100%qbrin
~1 / 6Graphiti
clean fact-invalidations

Against the one rival that genuinely tracks when facts become valid or invalid, does it consistently retire old ones? Graphiti got it right once in six scenarios and formed no link at all in half of them.

Methodology & caveatsMEASURED, 6 scenarios, edges read directly from the graph store (Graphiti’s own search path errored). Graphiti’s bi-temporal IS real but inconsistent; better prompting might lift it. qbrin’s 100% is the CTP-Bench reference.
vs HippoRAG 2

Finds chained facts across documents

97.0qbrin
96.0HippoRAG 2
recall@5 (higher is better)

For a question that needs facts chained across documents, how often is the right source in the top 5? On this news corpus qbrin edges out HippoRAG’s graph algorithm.

Methodology & caveatsMEASURED, same 609-doc corpus, 200 questions, same gold. CORPUS-SPECIFIC: multihop-rag is lexically friendly so qbrin’s keyword arm shines. Win is +1.0pt at R@5 and a tie (99.0) at R@10, not a universal retrieval win.
See it answer

See it answer your hardest question.

Bring one real question your team keeps re-asking. We'll connect a source, read-only, and show you the answer, sourced, in seconds, in a 20-minute walkthrough. Nothing changes in your tools.

20-minute walkthrough
  • One source connected, read-only
  • Your real question answered, with sources
  • Nothing changes in your tools
Pick a timeor email hello@qbrin.com