How much it reads for one answer
qbrin reads ~687 tokens per answer, a short briefing, not a pile of documents
per answer
vs a heavy production RAG pipeline (multi-query ×3 / 20-chunk + reranking) · ~8× vs even a lean 10-chunk RAG call
Grounded in the source, not guessing.
Ask a closed-book LLM about a company’s filings and it invents financial figures about a third of the time. qbrin reads the actual document and answers from it, so it’s right, or it tells you it isn’t sure. The same cited-or-abstain contract is what the REST API returns.
Closed-book LLM vs qbrin’s retrieval-grounded answers on FinanceBench, real public-company 10-K filings, each graded against the source passage. Preliminary sample (n=16).
qbrin’s own enterprise benchmark on real company Q&A.
500 real company questions (470 answerable + 30 unanswerable), the most product-representative test we run. The flagship result is never confidently wrong plus citations you can trust, not a blanket accuracy crown.
Honest read: the real win here is risk / safety and citation-trust, not blanket accuracy. On this 94%-answerable workload a naive RAG that always answers scores ~83% on the answerable subset, qbrin’s edge is that it’s never confidently wrong and its citations hold.
Validated across the board.
One story holds across all of it: qbrin is the most trustworthy (lowest rate of confident wrong answers, highest citation-trust) and the cheapest per answer, and its retrieval is competitive-to-leading once tuned to a corpus, with one architectural exception: relational-graph traversal (a rival’s home turf), where it trails. On raw all-answerable public Q&A it runs a calibration posture (abstains rather than guess); the 2026 embedder upgrade lifted that retrieval substantially (FinanceBench gold-in-context +18pp same-harness). Every cell below is a real run.
| Dataset | Safetyfewer made-up answers | Citation-trustcited source actually holds | Current-factretires stale / contested | Retrievalfinds the right source | Answer accuracycorrect when it answers | Cost / tokenstokens read per answer | Multilingualparity across languages |
|---|---|---|---|---|---|---|---|
| CTP-Bench (temporal/contested)n=12 | Leads | ||||||
| Temporal vs mem0 (open-source)n=8 | Leads | ||||||
| Temporal vs Zep / Graphitin=6 | Leads | ||||||
| PrecisionMemBench (memory-precision)n=77 | Leads | Leads | |||||
| BrainBench (relational retrieval)n=145 | Trades for calibration | ||||||
| Multi-hop vs HippoRAG 2n=200 | Leads | ||||||
| MuSiQue (hard multi-hop)n=300 | Competitive | ||||||
| multihop-rag (retrieval)n=2255 | Leads | Leads | |||||
| multihop-rag (answer / unanswerable)n=120 / 301 | Leads | Trades for calibration | |||||
| ragbenchn=100 | Trades for calibration | Competitive | |||||
| financebenchn=150 | Leads | Trades for calibration | |||||
| Enterprise RAG Benchmark (ERB)n=500 | Leads | Leads | Leads | Trades for calibration | Leads | ||
| G3 contested enterprise Q&An=60 | Leads | Competitive | Leads | ||||
| Multilingual (en / hi / te / ta)n=188 | Leads | Competitive | Leads | ||||
| Knowledge-map compression (SOTA)n=80 | Competitive | Leads | |||||
| Hallucination / fake-entity setn=60 | Leads | ||||||
| Broad RAG general Q&An=120 | Competitive | Competitive | |||||
| Internal enterprise benchmarkn=40 | Competitive | Competitive | Competitive | ||||
| False-premise safety (real / fake)n=55 | Competitive |
Measured, not modeled, each row is a benchmark we ran on qbrin’s own pipeline (competitor figures are their own published numbers or faithful re-runs). Several wins are corpus-specific and several safety A/Bs are honest ties; the per-suite tabs below carry the exact figures and caveats. The honest headline: no competitor matches qbrin on trust, safety-under-abstention and cost, and its retrieval is competitive-to-leading once tuned to a corpus (except relational-graph traversal, where it trails), on raw answerable-Q&A accuracy it runs a calibration posture (trades coverage for caution), and the 2026 embedder upgrade lifted that retrieval substantially.
Built to never make things up.
Numbers come from real benchmark runs on qbrin’s own pipeline, the ERB set of 500 real company questions (470 with a known correct source), plus trap questions built to bait a wrong answer. Each figure below is labelled MEASURED (we ran it), MODELED (computed from measured tokens × published prices), or PUBLISHED (a competitor’s own number). No per-suite second-by-second timings are shown unless we measured them. The full SOC write-up, qbrin against a Wazuh RAG pipeline, an MCP agent and a naive LLM analyst, is on the security benchmark page, and the architecture-level differences against Glean, including where Glean is stronger, are set out in qbrin vs Glean.
See it answer your hardest question.
Bring one real question your team keeps re-asking. We'll connect a source, read-only, and show you the answer, sourced, in seconds, in a 20-minute walkthrough. Nothing changes in your tools.
- One source connected, read-only
- Your real question answered, with sources
- Nothing changes in your tools