Skip to content
Menu

How we measure

If we claim the assistant will not answer without a source, we have to be able to show it. This page explains exactly what we measure, on what, how — and where the limits of that measurement are.

Why measure at all

We test whether the assistant answers questions covered by its knowledge base and declines questions the sources cannot cover. The result below comes from a completed run on our own RobiFox help documentation.

The test set

RobiFox’s own help documentation, including the RobiFox_Kezikonyv.docx manual dated 22 August 2026, and 52 prepared questions in Hungarian. This measures one Hungarian support knowledge base; it does not measure English or cross-language answer quality.

40

40 answerable questions

Questions covered by the help documentation. Each has an expected source, usually a numbered section in the uploaded Word manual.

12

12 unanswerable questions

Questions whose requested answer is absent from this test knowledge base, including private account balances, individual billing details and advice outside product support.

One example of each
Answerable: “Which document formats can I upload?” — the manual covers PDF, Word and spreadsheet files. This is an English translation of a Hungarian test question.
Unanswerable: “How many credits remain in my own account?” — the help documentation cannot disclose an individual visitor’s balance. This is an English translation of a Hungarian test question.

What we measure

Correct refusal

How many of the 12 unanswerable questions it declines. This is the most important number: failing here means confidently telling a prospective customer something untrue. The target is 100%, and below 100% the run counts as a FAILURE, not a "good result".

Expected source cited

For questions with a recorded expected source, we check whether that exact reference appears in the answer’s citations. A file-section reference is not a public web link. This check does not test HTTP availability or the factual correctness of the answer.

Answer rate

How many of the 40 answerable questions receive an answer. Answering at all does not prove that the answer is correct; review the actual wording against the source.

Latency and cost

The run records retrieval plus model latency for each question and the credits charged. We show the median response time and the run’s actual credit total. Provider load can change the timing.

How the measurement runs

  • On the LIVE system, end to end: the same retrieval, the same model call, the same answer schema a visitor gets. Not a laboratory copy.
  • Questions run sequentially in batches of up to 50. A 52-question set is collected into one completed run.
  • The table shows the latest completed run. Repeat runs can reveal variability, but a single percentage is not a guarantee for future answers.
  • Compare the same questions, knowledge version and model when assessing a change. Date, model and real output matter; runs on different days cannot establish that a code change caused faster answers.

The latest result

Measured: 2026. October 1., on the live system, with the "openai/gpt-5.6-luna" language model.

Our own RobiFox help project, tested in Hungarian. These are internal measurements, not customer results or an independent audit. Answer/refusal status and exact source matching are automatic checks; semantic answer correctness requires reading the output.

What we measure Result Target
Correct refusal (unanswerable questions) 11/12 (92%) 100% — below this is a failure
Expected source reference present 35/40 (88%) 100% — below this is a failure
Answered the answerable ones 37/40 (93%) no hard target
Response time (median) 4.2 s no hard target
Credits charged for the full run 38 no hard target
Execution errors 0 no hard target

Download the dated test results (JSON) — 2026-10-01, v2

Actual responses from the test

The question and response below are the original Hungarian output, not a simulated conversation or an English-language test. Check the source and the answer together.

Milyen formátumú dokumentumot tudok feltölteni új anyagként?

Új anyagként PDF-, Word- vagy táblázatfájlt tudsz feltölteni. Feltöltés előtt ellenőrizd a dokumentáció idevágó részét, mert az útmutató idővel frissülhet.

References returned by the assistant:

  • fajl://RobiFox_Kezikonyv.docx#szakasz=39
  • fajl://RobiFox_Kezikonyv.docx#szakasz=38

Mennyi kreditem maradt jelenleg a saját fiókomban?

A saját fiókod aktuális kredit-egyenlegét nem látom, ezért nem tudom megmondani, mennyi maradt. Bejelentkezés után a „Kredit-egyenleg és felhasználás” lapon ellenőrizheted a rendelkezésre álló kreditet; kérlek, döntés előtt ott ellenőrizd az aktuális adatot.

References returned by the assistant:

  • fajl://RobiFox_Kezikonyv.docx#szakasz=85

What these numbers do NOT mean

  • They are NOT a promise about your site. They come from one specific website's knowledge base. Your result is decided by the quality of your content: a contradictory, outdated website will not produce a good assistant — at best one that honestly says it does not know.
  • They are NOT customer references. We measured this ourselves, with our own tool. It is not an independent audit and we do not present it as one.
  • They are NOT a guarantee that the assistant never errs. The language model can be wrong, and your knowledge base can contain mistakes. That is why every answer carries a source and every conversation is logged: we do not promise there are no errors — we promise they SURFACE.
  • Response time is not a service level commitment. It depends on the language model provider and the knowledge base; the displayed median describes this run only.

Measure your own

The same measurement can be run against your knowledge base, and you see the result on the Quality page in the panel. The refusal rate and the missing-knowledge list together show where your content stands — often the most valuable thing that introducing an assistant gives you.

Start free

FAQ · How it works

Practical chatbot guides

Explore our English guides to source content, answer quality and human handover.

We always use the cookies required for the site to work: robi_fox_session, XSRF-TOKEN, felulet_nyelv, ss_suti_valasztas Beyond those we would like to measure how the site is used (Google Analytics) — that needs your permission. We use no advertising cookies, and we never measure on the trial pages. Detailed cookie notice