Cookies · Your choice

Analytics cookies, only if you say so

This site uses Google Analytics to count visits and see which pages get read. Those cookies stay off until you allow them. Everything the site needs to work runs without them. Cookie Policy.

This site uses Google Analytics to count visits. You can opt out at any time on Your Privacy Choices. Cookie Policy · Your Privacy Choices

Evaluation record

What the checks show.

This page records the tests behind the fit reader, including a failed earlier evaluation. The scoring policy has changed since that run. Those older numbers don’t tell us how accurate the current reader is.

The current release standard

On September 9, 2026, the JD policy changed to credit transferable skills and a reasonable learning ramp. To count related experience, the reader must explain which existing skill applies and what Kevin would still need to learn. Years, credentials, eligibility and past projects remain factual claims.

Release checks cover representative fit outcomes, protected facts, local handling of pasted text, and working browser flows. Existing Ask Kevin guards remain in place. The earlier JD statistical gate is historical evidence; its failed result below remains a failure under that earlier standard.

Fit policy 2026-09-09.1
Release policy 2026-09-09-transferable-fit-v1

Current functional checks

Verified 2026-09-10T19:11:04.431Z: 6 release suites passed, including 12 representative fit checks. Browser checks covered local semantic retrieval, fallback, handoff, clear, and desktop/mobile layout. No marker-bearing requests were observed during the fit reader’s privacy checks.

Representative functional and privacy checks, not a statistical accuracy estimate or a guarantee about browser extensions. Historical JD measurements are unchanged.

Read the local verification receipt →

Historical measurements

Recorded 2026-09-09T22:20:34.367Z. Suite sizes include all examples; retrieval intervals use the in-corpus subset, while negative-intent precision uses fired claims. JD intervals are document-cluster bootstrap bounds on extraction recall. These small, different samples are not a single overall accuracy score.

The historical JD failure includes missed requirements and mistakes distinguishing factual constraints from technical capabilities. Those are real limitations, beyond the change in how transferable skills count. Read the original role beside the annotations and discuss consequential requirements directly.

ask-retrieval@calibration

Historical pass · n = 25

95% interval: 0.6720.969

Aggregate measurements
hitAt5Rate
0.889
ndcgAt5
0.663
mrr
0.838
inCorpusN
18
oodN
4
adjacencyN
3

ask-retrieval@holdout

Historical pass · n = 15

95% interval: 0.7221.000

Aggregate measurements
hitAt5Rate
1
ndcgAt5
0.565
mrr
1
inCorpusN
10
oodN
3
adjacencyN
2

policy-goldens@calibration

Historical pass · n = 21

Aggregate measurements

No aggregate measurements were published for this record.

negative-intent@calibration

Historical pass · n = 389

95% interval: 0.9871.000

Aggregate measurements
precision
1
controlN
92
controlFired
0
targetN
297
targetFired
297
targetRecall
1

negative-intent@holdout

Historical pass · n = 389

95% interval: 0.9871.000

Aggregate measurements
precision
1
controlN
91
controlFired
0
targetN
298
targetFired
298
targetRecall
1

jd-requirements@calibration

Historical pass · n = 14

95% interval: 0.9760.996

Aggregate measurements
documents
14
goldAtoms
453
extractionRecallMacro
0.987
extractionRecallLow
0.976
extractionRecallHigh
0.996
protectedTypingPrecision
1
protectedTypingRecall
1
coordinationCollapses
1
omitted
0
Requirement typing · rows: reference; columns: predicted. Counts cover matched atoms.
Referencehard gatequantified experiencetool vendorcapability outcomedomainpreferenceunclassified
hard gate43000000
quantified experience03700000
tool vendor00369013
capability outcome005910057
domain0000213
preference000025743
unclassified00060051

jd-requirements@holdout

Historical failure · n = 12

95% interval: 0.8910.959

Aggregate measurements
documents
12
goldAtoms
414
extractionRecallMacro
0.928
extractionRecallLow
0.891
extractionRecallHigh
0.959
protectedTypingPrecision
0.904
protectedTypingRecall
0.930
coordinationCollapses
3
omitted
0
Requirement typing · rows: reference; columns: predicted. Counts cover matched atoms.
Referencehard gatequantified experiencetool vendorcapability outcomedomainpreferenceunclassified
hard gate35101002
quantified experience03000002
tool vendor01294008
capability outcome011630063
domain0000302
preference100004152
unclassified40070032

jd-requirements@calibration-2

Historical pass · n = 14

95% interval: 0.9570.994

Aggregate measurements
documents
14
goldAtoms
549
extractionRecallMacro
0.979
extractionRecallLow
0.957
extractionRecallHigh
0.994
protectedTypingPrecision
1
protectedTypingRecall
0.989
coordinationCollapses
0
omitted
0
Requirement typing · rows: reference; columns: predicted. Counts cover matched atoms.
Referencehard gatequantified experiencetool vendorcapability outcomedomainpreferenceunclassified
hard gate41000001
quantified experience05000000
tool vendor00335017
capability outcome00108500138
domain0000000
preference000004147
unclassified00050272

Where the numbers came from

The measured corpus differs from the corpus used to author labels. Corpus, rubric, model and policy changes limit comparisons across runs. This snapshot preserves the recorded measurements; no new holdout run or relabeling is implied.

Report revision
fe25dd3d877da4b6257b17946a8289bb20a1b0d7
Report SHA-256
3e4785a02f26b3d9ecda2e2f1d76d07fcaba589b4b84c7523081eb0d4fff79dd
Measured corpus
a600e51263fec471f946df0c99d9b11d2cd4531b31053e0bfc2b2527fcee058b
Label corpus
1420d480e1522b0a1a07215766217c6aa8e9d4ca0d2eb0ab189d3b903523e256
Embedding model
Xenova/bge-small-en-v1.5@q8 rev ea104dacec62
Embedding configuration
bge-small-en-v1.5:q8:cls:norm:tok512
Recorded rubric
0.2.0-draft

The JD reader uses local keyword matching and optional local embeddings. It does not send the pasted role to a generative model. These retrieval and policy checks do not measure every provider response in Ask Kevin.

The current model supply-chain lock pins Xenova/bge-small-en-v1.5 at revision ea104dacec62c0de699686887e3f920caeb4f3e3. Model files are self-hosted and SHA-256 pinned; the container build verifies the committed lock. This describes the build contract, not a new measurement of the historical report.

Fincel Design, LLC · Jacksonville Beach, FL · Est. 2008

PrivacyCookiesYour Privacy Choices

§ Cookies

Choose what this site may keep in your browser. Details for every item are in the Cookie Policy. You can change this at any time from the bottom of any page. Cookie Policy.