Self-learning, memory, and evidence
Self-learning is one of OpenKyrozen's two defining features. It is a guarded feedback loop around real work: the agent records what happened, proposes a small reusable improvement, tests that improvement against evidence, and keeps it only when the result is repeatably better or non-regressing.
This is not autonomous source-code rewriting or model fine-tuning. Learning artifacts are bounded policies or skills. They cannot add capabilities, create dynamic tools, grant permissions, add provider credentials, or upload private memory. Every activation, promotion, retirement, and rollback is recorded in the local SQLite event ledger.
The learning loop
- Observe: completed multi-step work, explicit corrections, and verified failures produce outcome evidence. Routine one-step work, provider failures, secret-bearing runs, and indexing are excluded.
- Propose: the background worker may create one bounded
policyorskillcanary for a run, or abstain. - Validate: static checks enforce the manifest, permissions, size limits, secret redaction, and workspace containment. A canary must then show two distinct verified successes and pass a paired candidate-versus-predecessor replay without regression.
- Promote or roll back: a passing canary becomes active; a correction or repeated verified failures restore its predecessor. The previous artifact remains recoverable.
Start with /self-learning, inspect proposals with /learning status, and
open a proof card with /learning evidence <proposal-id>. Use
/learning rollback <proposal-id> when an operator needs to restore the
predecessor immediately.
The rest of this guide documents the callable surface and its evidence. The
verification record below is explicitly a historical snapshot, not a claim
about the repository's current HEAD.
Historical verification snapshot: 51be33361422e55e1f2f00c33a0e0f8c56132a91
(the post-#54 main revision, captured before this #55 documentation-only
update). Snapshot date: 2026-09-04.
Current repository test count at this snapshot: 453 unittest cases.
Verified surface
| Surface | Current behavior | Verification |
|---|---|---|
| Profile routing | auto routes normal requests to coder or researcher; an explicit profile wins. |
tests/test_evolution.py |
| Evidence ledger | Runs, actual provider/model, ledger-backed token usage, measured latency, tool/error receipts, acceptance evidence, and artifact hashes are durable SQLite events. | tests/test_evolution.py, tests/test_evolution_robustness.py, tests/test_subagents.py |
| Canary lifecycle | A learned artifact stays a canary until two distinct verified successes and a non-regressing paired replay; correction or repeated verified failures rolls it back. | tests/test_evolution.py |
| Deterministic selection | Matching is profile- and trigger-scoped, with at most three artifacts and 8,000 body characters, including at most one canary. | tests/test_causal_selection.py |
| Retirement | A paired omission trial can retire an artifact only when completion does not regress; restore keeps the pre-image recoverable. | tests/test_causal_selection.py |
| Robustness | Environment/model mismatch, verifier reliability, constitution rules, and inactive capsule imports are enforced. | tests/test_evolution_robustness.py |
| Typed memory | Owner claims, inferred candidates, scope precedence, dependencies, and /memory why|forget are available. |
tests/test_memory_claims.py |
| Multi-party memory | Attributed beliefs, private facts, and group agreements are filtered by speaker, audience, channel, and visibility; receipts identify speakers whose claims were used. | tests/test_multi_party_memory.py, tests/test_server.py |
| Harness-neutral benchmark | Clean/evolved runs and candidate, predecessor, omission, and no-memory ablations export JSON with a claim gate. | learning_benchmark.py, benchmarks/multi_party_memory.jsonl |
The implementation stores its durable evidence locally. It does not fine-tune weights, synchronize cloud memory, edit harness source, or grant capabilities. Dynamic tools are controlled by a separate operator setting and remain subject to capability and approval gates. See security and the current generated runtime inventory for tool names and counts.
Terminal workflow
Start the agent with make run (which passes --project .) or, from an
installed package, kyrozen. Bare kyrozen uses the persistent global
workspace at ~/.kyrozen/workspace; use kyrozen --project . when a source
checkout or another project should be edited directly. These commands are
handled by the running terminal session:
| Command | Use |
|---|---|
/agent auto |
Let task terms choose coder or researcher. |
/agent coder |
Force repository, terminal, and implementation learning. |
/agent researcher |
Force research, analysis, and source-bound writing learning. |
/learning status [coder\|researcher] |
List proposal stage, profile, uses, failures, and predecessor. |
/learning metrics [coder\|researcher] |
Show verified completion, corrections, repeated errors, tool calls, tokens, latency, and near misses. |
/learning evidence <proposal-id> |
Show the proof card, applicability, replay, outcomes, and metrics. |
/learning explain <proposal-id> |
Show the complete proposal record. |
/learning replay <proposal-id> |
Explain the API replay workflow; it never runs a command itself. |
/learning rollback <proposal-id> |
Roll back a learned proposal or its skill version. |
/memory why <claim-id> |
Inspect claim type, scope, authority, dependencies, and status. |
/memory forget <claim-id> |
Forget a claim and deactivate solely dependent learned artifacts/regressions. |
/learning replay is intentionally informational because replay results must
come from a separately sandboxed, frozen run. Submit those results through the
API below.
How automatic evolution behaves
- Each eligible run is recorded under one profile. Completed multi-step runs, explicit corrections, and verified failures may enter review. Provider failures, secret-bearing runs, routine one-step work, and project indexing are excluded from behavioral evolution.
- The CLI launches one detached learning worker per workspace. The worker
waits until the interactive heartbeat is quiet for at least 60 seconds,
then reviews on its 30-second cycle even after the terminal exits. A
reviewer can create at most one bounded
policyorskillcanary for the run, or abstain. Feature switches are persisted in SQLite. - Static validation requires the manifest, Markdown sections, size limits, secret redaction, workspace containment, and declared permissions. Learned artifacts cannot add capabilities or dynamic tools.
- A matching canary is used at most once per run. Promotion requires two distinct verified successes, no correction or failed use, and a paired candidate-versus-predecessor replay that does not regress.
- A verified artifact-linked correction, or two verified failures among its last five uses, rolls the artifact back to its predecessor. Every activation, promotion, retirement, and rollback is recorded as an auditable event.
For coding, a successful tool call is not enough: an acceptance command/test or explicit user confirmation is required. For research, requested sources, structure, or artifact creation must be observable, or the user must confirm success. Otherwise the run remains unverified.
Web API
Install the web extras and start the server:
pip install -e '.[web]'
KYROZEN_SERVER_TOKEN=change-me kyrozen-web --project . --host 127.0.0.1 --port 8000
Loopback requests may omit the token. Any non-loopback deployment must send
Authorization: Bearer <token> (or X-Kyrozen-Token) and should bind to an
explicit interface.
Chat and profiles
profile is optional; auto is the default. Chat context can also carry a
speaker, audience, and channel. The response includes memory_receipt when
stored claims affected the turn.
curl -sS http://127.0.0.1:8000/api/chat \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <token>' \
-d '{"message":"Write a sourced migration note", "profile":"researcher", "session_id":"demo", "speaker":"authenticated", "audience":"team", "channel":"chat"}'
The streaming endpoint /api/chat/stream accepts the same body and forwards
provider content deltas as they arrive. Control responses are buffered only
long enough to parse actions; durable tool receipts, task progress, usage,
memory, and completion are emitted as typed SSE records. The existing chunk,
cost, memory_receipt, error, and [DONE] payloads remain compatible.
Inspect learning evidence
curl -sS 'http://127.0.0.1:8000/api/v2/learning?profile=coder&status=canary' \
-H 'Authorization: Bearer <token>'
curl -sS 'http://127.0.0.1:8000/api/v2/learning/metrics?profile=coder' \
-H 'Authorization: Bearer <token>'
curl -sS 'http://127.0.0.1:8000/api/v2/learning/<proposal-id>/evidence' \
-H 'Authorization: Bearer <token>'
Submit a paired shadow replay only after both rows have been executed in a
side-effect-free sandbox. Candidate and predecessor arrays must be non-empty,
the same length, and contain identical unique case_id values:
curl -sS -X POST 'http://127.0.0.1:8000/api/v2/learning/<proposal-id>/replay' \
-H 'Content-Type: application/json' -H 'Authorization: Bearer <token>' \
-d '{"candidate":[{"case_id":"case-1","verified_success":true,"provider_model":"provider:model-a"}],"predecessor":[{"case_id":"case-1","verified_success":true,"provider_model":"provider:model-a"}]}'
Use /omission with matching with_item and without_item arrays to record a
candidate omission trial. /retire requires a non-regressing trial;
/restore returns a retired artifact to canary; /rollback immediately
restores its predecessor. /capsule exports redacted evidence and
POST /api/v2/learning/capsules imports it as an inactive candidate—import
never activates trust.
Typed and multi-party claims
Create an attributed belief or group agreement with the claim endpoint:
curl -sS -X POST http://127.0.0.1:8000/api/v2/memory/claims \
-H 'Content-Type: application/json' -H 'Authorization: Bearer <token>' \
-d '{"key":"release date","value":"Monday","claim_type":"attributed_belief","speaker":"alice","visibility":"group","audiences":["ops"],"channel":"release","authority":"owner"}'
The supported claim types are general, attributed_belief, private_fact,
and group_agreement. Filter retrieval before it reaches the model:
curl -sS 'http://127.0.0.1:8000/api/v2/memory?q=release%20date&speaker=alice&audience=ops&channel=release' \
-H 'Authorization: Bearer <token>'
curl -sS 'http://127.0.0.1:8000/api/v2/memory/claims?speaker=alice&audience=ops&channel=release' \
-H 'Authorization: Bearer <token>'
An attributed belief is always rendered with its speaker and does not resolve
to an unattributed value when speakers disagree. A private fact requires
authority=owner and is never exposed through unscoped recall. Private API
claims use the stable KYROZEN_SERVER_ACTOR configured for the deployment;
omitting speaker uses that actor for create, list, detail, and forget. A
client-supplied label such as alice is an attribution/filter, not an
authorization grant. Deployments that need distinct human identities or group
membership must use separate deployments and databases or enforce that mapping
at their authentication/proxy layer.
Use the same deployment actor for a complete private-claim lifecycle (the
default actor is local):
export KYROZEN_SERVER_ACTOR=owner-66
curl -sS -X POST http://127.0.0.1:8000/api/v2/memory/claims \
-H 'Content-Type: application/json' \
-d '{"key":"favorite editor","value":"vim","claim_type":"private_fact","authority":"owner"}'
curl -sS http://127.0.0.1:8000/api/v2/memory/claims
curl -sS http://127.0.0.1:8000/api/v2/memory/claims/<claim_id>
curl -sS -X DELETE http://127.0.0.1:8000/api/v2/memory/claims/<claim_id>
GET /api/v2/events?event_type=memory.recalled shows the receipt payload used
for a turn. Learning events such as learning.outcome,
learning.artifact_promoted, and learning.artifact_rolled_back provide the
corresponding audit trail.
Benchmark replay
The benchmark runner is harness-neutral and reads one JSON object per case from stdin. Each wrapper must emit these fields:
{"verified_success":true,"corrections":0,"repeated_errors":0,"tool_calls":4,"tokens":1200,"latency":2.4}
Run the frozen multi-party cases with identical clean/evolved settings:
python main.py learning benchmark \
--cases benchmarks/multi_party_memory.jsonl \
--clean-runner './clean-wrapper' \
--evolved-runner './evolved-wrapper' \
--ablation 'candidate=./candidate-wrapper' \
--ablation 'predecessor=./predecessor-wrapper' \
--ablation 'omission=./omission-wrapper' \
--ablation 'no-memory=./no-memory-wrapper' \
--output benchmark.json
The JSON output preserves case order, per-case results, Wilson completion intervals, secondary metrics, and ablations. The claim gate is true only when completion is non-regressing and a paired secondary improvement is credible; OpenKyrozen does not emit a competitor-superiority claim from a tool success or an unpaired run.
The historical snapshot's isolated make benchmark output had five cases and
five verified successes for both runners. Both runners used provider
deterministic-fixture and model openkyrozen-memory-policy-v1; each case had
fixture_verified evidence. The observed summary was:
| Runner | Cases | Verified successes | Tool calls | Tokens | Evidence status |
|---|---|---|---|---|---|
| clean | 5 | 5 | 22 | 973 | fixture_verified |
| evolved | 5 | 5 | 22 | 973 | fixture_verified |
The run reported completion_non_regressing: true,
paired_evidence_status: insufficient, and
public_superiority_claim_supported: false. The fixture is deterministic
and proves the scoped memory behavior; it does not provide the independently
paired product evidence required for a public superiority claim. Latency is
diagnostic only and is not treated as a release claim.
Verification snapshot commands
Current repository test count at this snapshot: 453 unittest cases.
The post-#54 snapshot ran the repository's current checks and smoke coverage:
make check
make docs-check
make shell-check
make lint
make test
make benchmark
git diff --check
The make test suite includes the API health/scoping smoke and the CLI
command-loop smoke. make benchmark uses the five-case
benchmarks/multi_party_memory.jsonl fixture, temporary SQLite databases, the
openkyrozen-learning-benchmark-v1 protocol, and no provider credentials.
The documented commit is retained only as a historical verification snapshot;
later commits must run these checks again before making current-release claims.
Verification and troubleshooting
Run the same local gates used for the merged implementation:
make check
make docs-check
make shell-check
make lint
tmpdir=$(mktemp -d /tmp/openkyrozen-check.XXXXXX)
KYROZEN_DB_PATH="$tmpdir/state.sqlite3" make test
make benchmark
git diff --check
The earlier 2026-09-04 post-#54 snapshot reported 130 discovered tests in its
historical run and 409 discovered tests at that snapshot. Those counts are not a
current test result. It also covered the API health/scoping smoke, the CLI
command-loop smoke, and the five-case clean/evolved benchmark described above.
Its make check receipt reported 45 runtime tools, including 14 git_ tools;
the current generated inventory supersedes that historical
count. A new artifact is not immediate:
wait for an eligible run, at least
60 seconds of idle time, reviewer evidence, and then the two-success plus
replay gate. A model or environment mismatch intentionally downgrades an
artifact to canary. If ChromaDB is unavailable, SQLite remains the durable
source and keyword recall continues to work.