← Documentation indexdocs/system-one-benchmark.md

System One benchmark protocol

benchmarks/system_one.py is the repeatable typed-decision benchmark. It uses the same fixed labeled corpus for routing, an explicitly supported clarification, learning evidence, tool-output passages, and eight competing memory candidates. It does not send the corpus text to diagnostics or store it in the result file.

Each repetition uses the same cases. Decision rows report correctness, accepted-decision coverage, abstentions, false decisions, decision latency, token usage, Brier score, ECE, reliability bins, and repeated-decision agreement. Memory rows report precision@3, recall@3, MRR@3, NDCG@3, and effective rerank rate. Tool rows report precision, recall, specificity, F1, warning rate, and quarantine rate. When --url is supplied, the existing paired end-to-end harness additionally reports full-turn median/p95 latency, LLM calls/tokens, answer agreement, cost, fallbacks, provider errors, and warm/cold setup time.

The baseline is the current path: it abstains from typed decisions and keeps the original memory order. Kev and Jev use the same questions and labels. Run order is independent for typed calls; full-turn comparisons alternate the backend order and create a fresh session for every side. Setup time is kept outside turn latency.

# Reproducible three-repeat comparison (local Kev plus live Jev when configured)
python benchmarks/system_one.py --backends baseline,kev,jev --repeats 3 \
  > benchmarks/results/system_one_$(date +%F)_baseline_kev_jev.json

# Fit backend/action policies only after collecting a fixed labeled corpus.
# The command uses an 80/20 calibration/holdout split and writes only policies
# with pass/fail gates to ~/.kyrozen/system_one_calibration.json; only passing
# policies are applied at runtime.
python benchmarks/system_one.py --backends kev,jev --repeats 3 --calibrate

# Jev requires the environment key; it is never read from command-line text.
TYPESAFE_API_KEY=... python benchmarks/system_one.py \
  --backends baseline,jev --repeats 5 \
  > benchmarks/results/system_one_$(date +%F)_baseline_jev.json

# Full-turn paired measurement against one configured OpenKyrozen server.
python benchmarks/system_one.py --backend kev --repeats 3 --url http://127.0.0.1:8000 --mode ask

Current calibrated comparison (2026-09-29)

This run expanded the fixed corpus to eight routing cases, three explicit clarifications, six evidence cases, six tool-output cases, and three distinct memory tasks. Each backend ran three repetitions. Calibration used grouped 80/20 splits, with tool labels stratified so both malicious and benign passages appear in the holdout. The calibration raw rows and policies are in benchmarks/results/system_one_2026-09-29_baseline_kev_jev_calibration.json. The second pass below measures the policies actually applied at runtime: benchmarks/results/system_one_2026-09-29_baseline_kev_jev_applied.json. The canonical dataset hash is 7438998a5681eaed.

Applied workload Baseline Kev-0.8B Jev
Typed raw accuracy — 91.3% 100%
Typed accepted precision / coverage — / 0% 100% / 47.8% 100% / 87.0%
Typed Brier / ECE 0.250 / 0.370 0.1025 / 0.2325 0.0037 / 0.0319
Typed repeated agreement — 100% 100%
Decision median / p95 0 / 0 ms 70.6 / 81.8 ms 777.2 / 1,398.6 ms
Decision tokens (input + output) 0 7,200 26,352
Memory precision@3 / recall@3 0 / 0 0.556 / 0.556 1.000 / 1.000
Memory MRR@3 / NDCG@3 0 / 0 1.000 / 0.666 1.000 / 1.000
Memory effective rerank / promotion 0% / 0% 100% / 100% 100% / 100%
Tool precision / recall / specificity / F1 0 / 0 / 1 / 0 1 / 1 / 1 / 1 1 / 1 / 1 / 1
Tool quarantine / promotion 0% / 0% 50% / 50% 50% / 50%

The applied Jev pass accepted no false typed decisions and all five Jev action policies passed their held-out gates: routing, clarification, learning evidence, memory ranking, and tool quarantine. Kev passed routing, memory ranking, and tool quarantine; its clarification and learning-evidence policies remain advisory because their held-out coverage/quality gates did not pass. The memory result is materially better than the previous 0/3 smoke reranks, but it is still based on three labeled tasks and should be expanded before treating it as a general accuracy guarantee. Tool scores use the calibrated positive-class threshold and separate false-positive/recall gate.

The persisted probability thresholds are action-specific and come from this dataset: Jev uses 0.98 for routing, 1.00 for clarification, 0.78 for evidence, 0.70 for memory relevance, and 0.050001 for tool quarantine; Kev uses 0.5694, 0.9225, 0.71, 0.70, and 0.305501 respectively. The last two tool thresholds are the separating points just above the strongest benign score in each calibration split. Kev clarification and evidence thresholds are recorded but not active because their holdout gates failed.

Full-turn DeepSeek comparison (2026-09-29)

These runs used the configured paid DeepSeek provider, Ask mode, fresh sessions, and alternating off/on order. They measure full OpenKyrozen turns rather than only the decision endpoint. Reply text is not stored; the raw rows contain a short answer hash for repeated-agreement measurement.

Metric System One off Kev Jev
Matched turns / keyword checks 8 / 8 8 / 8 8 / 8
Median / p95 turn latency 3,781 / 9,771 ms 4,470 / 6,180 ms 4,282 / 5,681 ms
LLM calls 8 8 8
LLM tokens 16,203 14,386 13,451
Median System One decision latency 0 ms 121.1 ms 1,140.7 ms
Repeated answer agreement 0.625 0.625 0.625

Kev increased median turn latency by 18.3% but reduced recorded LLM tokens by 11.2% for this small paired workload. Jev increased median latency by 9.0% and LLM tokens by 0.7%. Answer agreement did not improve for either backend. The token reduction is an observed workload result, not a general cost guarantee; the primary System One evidence is decision correctness and calibrated uncertainty rather than speed.

Raw full-turn rows: Kev · Jev.

The live Jev server discovered jev-latest with release date 2026-09-10T18:38:01.391457+00:00; the /v1/systemone response resolved to jev-1.13.0. Discovery is refreshed at startup and after 24 hours, and a failed discovery safely falls back to the stable alias.

Promotion rules

Policies are backend and action specific. Each policy records confidence and probability thresholds, a probability margin, minimum calibration coverage, and the ordinary-path fallback (llm, user, candidate, existing_order, or advisory). A policy also needs at least three distinct labeled cases; repeated calls cannot substitute for new cases. It is written only after its held-out result passes the declared gate: routing precision at least 95%, learning supportive precision at least 99% with no accepted contradictions, memory NDCG@3 improvement of at least 0.10 absolute or 10% relative without a precision@3 drop, and tool quarantine with zero false positives and at least 90% malicious-passage recall. Otherwise the result remains advisory and the current OpenKyrozen path is used. The confidence value is recorded as a model signal; it is not treated as an accuracy guarantee.