ask-retrieval@calibration
Historical pass · n = 25
95% interval: 0.672–0.969
Aggregate measurements
- hitAt5Rate
- 0.889
- ndcgAt5
- 0.663
- mrr
- 0.838
- inCorpusN
- 18
- oodN
- 4
- adjacencyN
- 3
Cookies · Your choice
This site uses Google Analytics to count visits and see which pages get read. Those cookies stay off until you allow them. Everything the site needs to work runs without them. Cookie Policy.
This site uses Google Analytics to count visits. You can opt out at any time on Your Privacy Choices. Cookie Policy · Your Privacy Choices
Evaluation record
This page records the tests behind the fit reader, including a failed earlier evaluation. The scoring policy has changed since that run. Those older numbers don’t tell us how accurate the current reader is.
On September 9, 2026, the JD policy changed to credit transferable skills and a reasonable learning ramp. To count related experience, the reader must explain which existing skill applies and what Kevin would still need to learn. Years, credentials, eligibility and past projects remain factual claims.
Release checks cover representative fit outcomes, protected facts, local handling of pasted text, and working browser flows. Existing Ask Kevin guards remain in place. The earlier JD statistical gate is historical evidence; its failed result below remains a failure under that earlier standard.
Fit policy 2026-09-09.1
Release policy 2026-09-09-transferable-fit-v1
Verified 2026-09-10T19:11:04.431Z: 6 release suites passed, including 12 representative fit checks. Browser checks covered local semantic retrieval, fallback, handoff, clear, and desktop/mobile layout. No marker-bearing requests were observed during the fit reader’s privacy checks.
Representative functional and privacy checks, not a statistical accuracy estimate or a guarantee about browser extensions. Historical JD measurements are unchanged.
Read the local verification receipt →Recorded 2026-09-09T22:20:34.367Z. Suite sizes include all examples; retrieval intervals use the in-corpus subset, while negative-intent precision uses fired claims. JD intervals are document-cluster bootstrap bounds on extraction recall. These small, different samples are not a single overall accuracy score.
The historical JD failure includes missed requirements and mistakes distinguishing factual constraints from technical capabilities. Those are real limitations, beyond the change in how transferable skills count. Read the original role beside the annotations and discuss consequential requirements directly.
Historical pass · n = 25
95% interval: 0.672–0.969
Historical pass · n = 15
95% interval: 0.722–1.000
Historical pass · n = 21
No aggregate measurements were published for this record.
Historical pass · n = 389
95% interval: 0.987–1.000
Historical pass · n = 389
95% interval: 0.987–1.000
Historical pass · n = 14
95% interval: 0.976–0.996
| Reference | hard gate | quantified experience | tool vendor | capability outcome | domain | preference | unclassified |
|---|---|---|---|---|---|---|---|
| hard gate | 43 | 0 | 0 | 0 | 0 | 0 | 0 |
| quantified experience | 0 | 37 | 0 | 0 | 0 | 0 | 0 |
| tool vendor | 0 | 0 | 36 | 9 | 0 | 1 | 3 |
| capability outcome | 0 | 0 | 5 | 91 | 0 | 0 | 57 |
| domain | 0 | 0 | 0 | 0 | 2 | 1 | 3 |
| preference | 0 | 0 | 0 | 0 | 2 | 57 | 43 |
| unclassified | 0 | 0 | 0 | 6 | 0 | 0 | 51 |
Historical failure · n = 12
95% interval: 0.891–0.959
| Reference | hard gate | quantified experience | tool vendor | capability outcome | domain | preference | unclassified |
|---|---|---|---|---|---|---|---|
| hard gate | 35 | 1 | 0 | 1 | 0 | 0 | 2 |
| quantified experience | 0 | 30 | 0 | 0 | 0 | 0 | 2 |
| tool vendor | 0 | 1 | 29 | 4 | 0 | 0 | 8 |
| capability outcome | 0 | 1 | 1 | 63 | 0 | 0 | 63 |
| domain | 0 | 0 | 0 | 0 | 3 | 0 | 2 |
| preference | 1 | 0 | 0 | 0 | 0 | 41 | 52 |
| unclassified | 4 | 0 | 0 | 7 | 0 | 0 | 32 |
Historical pass · n = 14
95% interval: 0.957–0.994
| Reference | hard gate | quantified experience | tool vendor | capability outcome | domain | preference | unclassified |
|---|---|---|---|---|---|---|---|
| hard gate | 41 | 0 | 0 | 0 | 0 | 0 | 1 |
| quantified experience | 0 | 50 | 0 | 0 | 0 | 0 | 0 |
| tool vendor | 0 | 0 | 33 | 5 | 0 | 1 | 7 |
| capability outcome | 0 | 0 | 10 | 85 | 0 | 0 | 138 |
| domain | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| preference | 0 | 0 | 0 | 0 | 0 | 41 | 47 |
| unclassified | 0 | 0 | 0 | 5 | 0 | 2 | 72 |
The measured corpus differs from the corpus used to author labels. Corpus, rubric, model and policy changes limit comparisons across runs. This snapshot preserves the recorded measurements; no new holdout run or relabeling is implied.
The JD reader uses local keyword matching and optional local embeddings. It does not send the pasted role to a generative model. These retrieval and policy checks do not measure every provider response in Ask Kevin.
The current model supply-chain lock pins Xenova/bge-small-en-v1.5 at revision ea104dacec62c0de699686887e3f920caeb4f3e3. Model files are self-hosted and SHA-256 pinned; the container build verifies the committed lock. This describes the build contract, not a new measurement of the historical report.
Fincel Design, LLC · Jacksonville Beach, FL · Est. 2008