All notable changes to this project are documented here. The format
follows Keep a
Changelog. This project is pre-1.0: while 0.x, MINOR
versions may carry breaking changes.
phr-mcp now ships with the Prometheus
exporter. The metrics feature is on by default, so
cargo install phronesis-mcp includes
phr-mcp metrics and the /metrics endpoint. A
build without the HTTP stack is still available with
--no-default-features --features rhai.The nudge README that init writes showed a
capsule that could never fire. Its journey_seen
example used the window "session", which is not a window
token; the session window is "s". The example now uses
"s", and the README lists the valid windows
(s, <N>c,
<N>s/m/h/d).
After a broken rules file was repaired,
load_rules_file reported success but the MCP server kept
its stale rules and still reported the file as failing. The
server held its startup copy (rules since changed or deleted on disk
stayed loaded) and list_rules kept returning the old
load_error until the next rule-writing call.
load_rules_file now recovers the way a write does: it
clears the error and, with autopersist on, reloads the repaired file
wholesale and reports how many rules it loaded. With autopersist off the
load stays additive, so rules added in memory are kept.
An MCP add_rule phase typo did not say which
rule it was. "Pre" was refused with
phase must be "pre" or "post", got: Pre. The message now
names the rule, the field, the bad value and the allowed values, like
the other rule-shape errors.
A coverage import in progress could make hooks report a
corrupt coverage store. A hook that could not get the store
lock within its 200 ms bound (an import running back to back starves it
— flock queues no one) fell back to an unlocked read, and a read that
landed between the import’s two renames asserted
store_corrupt(coverage, digest_mismatch), which rules may
block on; the store concurrency tests failed at random on Linux CI for
the same reason. That read now reports the store as busy: the evidence
is treated as stale (coverage_stale, no gap suppressed, a
stderr note) and phr-mcp coverage select says an import is
in progress. A torn store with no one holding the lock — a crashed
import — is still store_corrupt.
cargo test passed while BDD scenarios never
ran. The cucumber runner (tests/bdd.rs) printed
its summary and exited 0 whatever it held, so all four
coverage-evidence.feature scenarios — whose steps had no
definitions — were reported skipped and the suite still went green. The
runner now fails the test process on any failed, skipped, or undefined
step, and those four scenarios have real step definitions: they run the
committed coverage-sample fixture and its real export
through the phr-mcp hooks and assert, from the action log,
which rule fired for which test and region (branch relevance, the
evidence-gap warning, demand gating, and the stale-coverage commit
warning that warns without blocking).
The packaged commit-gate rules could be bypassed with
git -C . commit, /usr/bin/git commit, or
similar. confidence-low-blocks-commit,
confidence-medium-warns-commit,
nudge-verify-before-commit, and
llm-warn-git-add-all matched a git subcommand only when it
directly followed the literal word git, so any of git’s
global options (-C <dir>,
-c <k=v>, --git-dir=…,
--work-tree=…, --no-pager,
--bare, …), an absolute or relative path to the
git binary, a backslash-escaped binary name
(\git), or a command/env FOO=bar
wrapper silently skipped every gate — the governed commit went through
with no confidence warning at all. The packaged rules now build their
bash_command_matches pattern from a shared regex
(git_invocation_prefix / git_subcommand_gate
in init.rs) that recognizes the subcommand past any of
those forms, while continuing to reject plumbing extensions
(commit-tree, merge-base do not mutate a ref
or the index the way the porcelain command does) and text that merely
mentions the word (git log --grep commit,
echo "git commit"). Existing projects pick this up via
phr-mcp init --rules-only --packs llm,confidence (or
rust/python/etc. alongside them).
A flood of low-priority context capsules could starve
governance capsules of room. CapsuleStorage::emit
reserved 32 of 128 record slots for priority >= 50 governance
capsules, but only counted records, not bytes: 32 low-priority capsules
near the 8 KiB per-record ceiling could fill the whole 256 KiB aggregate
budget by themselves, leaving 96 free slots but no room, so the next
governance capsule was rejected with CapacityExceeded.
Low-priority (priority < 50) capsules are now also capped at 192 KiB
(256 KiB * 96/128), guaranteeing governance capsules at least 64 KiB
regardless of how the low-priority pool fills up. Separately,
re-emitting (upserting) an existing capsule id bypassed the low-priority
count and byte checks entirely, so lowering an existing governance
capsule’s priority could push the low-priority pool past its 96-record
or 192 KiB reservation; an upsert is now checked as if the old version
were removed and the new one inserted fresh.
A verification result counted as proof without saying
what it proved. property-results.jsonl records
carried no confinement tier and no artifact hash, and
execute returned results with an empty property and
revision, so a hand-written kani passed line hydrated as
verification_result(…, kani, passed) for a property with no
kani encoding, a raw (unconfined) proof was indistinguishable from a
sandboxed one, and a result with an empty revision counted as “at HEAD”
whenever HEAD was unknown — silencing the first-proof obligation for
good. Result records (format v: 2) now carry the property,
the 40-hex commit the proof ran against, the verifier, the tier, and the
SHA-256 of the artifact bytes that ran, and execute fills
them (it now takes the property and verifier name, requires the artifact
to be approved for that property, and refuses a non-commit revision). A
record hydrates as verification_result only when every
binding holds and matches a property encoding for that verifier and an
allowlisted artifact for that property; otherwise it asserts
unbound_evidence(property, verifier, reason) and never
satisfies an obligation. Bound results also assert
result_tier(property, verifier, tier) so a rule can refuse
raw evidence (SPEC-property-ontology §2 “Result binding”).
Upgrade note: existing v: 1 records still
load but read as unbound_evidence(…, legacy_record) — they
no longer count as proof, so accepted properties with changed
dependencies raise property_obligation again until
re-proved. This repository’s own hand-seeded kani record is one of
them.
One bad line in the property results file silently turned
off every property rule. A malformed
property-results.jsonl line (or a malformed
properties.json) failed the whole property hydration with a
stderr note, dropping the property facts, stale-proof warnings, and
proof obligations. The hook now asserts
store_corrupt(properties, <reason>) — mirroring
store_corrupt(coverage, …) — so a rule can warn or block,
prints a stderr warning naming the file, and derives everything else as
if the corrupt file held no evidence, so obligations still fire. Loading
also rejects a result status outside the closed
passed | failed | inconclusive | timeout | unknown
set.
Codex’s adapter could silently allow what
pre-check/post-check would block or warn
on. After PR #83 gave codex-hook the same
rule-network builder as pre-check, three edges kept the
old, looser behavior: an update_agenda() failure was
discarded (let _ = ...), so a verdict derived purely from
fact matching could vanish without a trace instead of denying the tool
call; a bash_command_matches fact-assertion failure was
downgraded from a Claude-parity block (“command pattern check failed”)
to a warning at PreToolUse, and silently dropped entirely
at PostToolUse; and apply_patch’s per-file
content read treated any disk-read error (permissions, a dangling
symlink) as empty content, so a rule scanning file content saw nothing
instead of denying the call. All three now match
hook/pre.rs / hook/post.rs:
fire_verdict propagates an agenda-update failure as a deny
(pre) or advisory (post); the command-pattern check blocks at
PreToolUse and warns (not silently) at
PostToolUse; and evaluate_patch_file fails
closed on a genuine read error while still treating a missing file (a
patch’s own Add File target) as empty.
Coverage and property gap rules misfired on real editor
payloads. A Claude Code Edit sends a one-line
old_string/new_string, and the hooks diffed
that snippet as if it were the file: pre-check found no changed region,
so a gap rule never fired before the edit, and post-check compared the
snippet with the whole file, so every function — tests included — was
reported as a changed, untested region. The integration tests passed
only because they sent the whole file as old_string. The
hooks now diff whole files: pre-check applies the edit to the file on
disk (honouring replace_all, MultiEdit order,
Gemini’s expected_replacements, and Write
content), post-check reconstructs the pre-edit file by reversing the
edit (checked by re-applying it) or from Claude Code’s
tool_response.originalFile. When the edit does not match
the file, or the pre-image cannot be recovered, the whole file counts as
changed — gaps are over-reported, never missed. Property obligations use
the same regions. A file that exists but cannot be read, or is over the
PHRONESIS_MAX_FILE_BYTES read cap, or over 1 MiB for region
mapping, counts as one whole-file region file:<path>
instead of silently mapping to nothing; a stray non-UTF-8 byte no longer
hides a file. Region mapping is now linear in file size: a pre-check on
a large file that took minutes (tree-sitter parent walks and a full
line-diff table, both quadratic) returns promptly.
Coverage evidence joined different functions that shared
a name. Region ids were the bare leaf name
(fn:new, branch:safe_divide:<anchor>),
so a test that executed new in one file counted as evidence
for every other new:
region_without_dynamic_evidence stayed silent for code no
test had run, phr-mcp coverage select picked tests that
never touched the edited function, and two identical if
conditions in one function shared one branch id. Region ids are now
unique per code site — fn:<file>::<item-path>
and
branch:<file>::<item-path>:<anchor>[.<n>],
qualified by file, module, impl type (generics included), and trait,
with a source-order ordinal for repeated conditions
(SPEC-coverage-evidence §3.2). Hooks make the edited path repo-relative
first, so the absolute file_path Claude Code sends —
including one through a symlinked project root — joins the store; an
edit outside the project root names no region. The importer rejects the
old ids, so re-importing an old export fails — run
phr-mcp coverage collect again (see the upgrade note
below). Property depends_on entries still using the old ids
keep matching conservatively (every same-named site) until
rewritten.
A coverage import interrupted between its two writes no longer passes off new hits as the old revision’s evidence. The coverage index now records a digest and count of the records file it commits, and every read checks them — plus each record’s revision and tool — against the index. A torn or mismatched store reads as corrupt instead of fresh. Imports and reads share a lock, so a hook firing during an import sees the old store or the new one, never a spurious corruption; a hook waits at most 200 ms for that lock, so a hung import cannot hang it.
Coverage records the importer would refuse are refused on
read too, and import no longer accepts inconsistent exports.
One validator runs on import and on every store read: absolute paths,
unknown hit_kind, short revisions, and inverted line spans
are rejected wherever they appear, and hit_kind must agree
with the region id (region ⇒ fn:…,
branch ⇒ branch:…). Import now lowercases
revisions (an uppercase sha was treated as permanently stale, and
mixed-case spellings of one sha were rejected as “mixes revisions”),
collapses duplicate records instead of counting them, and rejects an
empty tool or an export that mixes tools.
Stale coverage no longer hides untested changes.
Hits imported at a revision other than HEAD never suppress
region_without_dynamic_evidence, and
phr-mcp coverage select labels them
coverage_observation_stale (with a
coverage_note in --json) instead of presenting
them as current.
A corrupt coverage store no longer silences every
coverage gap warning. The hook used to print one stderr line
and drop all coverage facts. It now asserts
store_corrupt(coverage, <reason>) for rules to warn
or block on, still warns on stderr, and reports changed regions as
having no dynamic evidence. phr-mcp coverage select says
the store is corrupt instead of “the coverage store is empty”.
Upgrade note: re-collect coverage after
upgrading. A coverage store written by an earlier version has
no records digest, so it reads as
store_corrupt(coverage, unverifiable_index): the hook warns
on stderr, reports changed regions as untested, and
phr-mcp coverage select notes the corruption. Re-importing
the old export does not help — the importer now rejects its leaf-name
region ids. Re-run phr-mcp coverage collect (it re-imports;
an export collected elsewhere then goes through
phr-mcp coverage import).
The code graph no longer invents calls and
tested_by edges by matching names. A rebuild of
this repository reported about a hundred functions calling themselves
(LogEntry::new “calling” LogEntry::new when it
really called serde_json::Map::new), and tests “directly
testing” methods they never touch (f.predicate.clone() on a
String counted as a test of
TaggerConfig::clone). Paths are now resolved the way Rust
resolves them: a Type::f() call carries its type; the first
segment of module::f() is read through the file’s
use bindings, so under use std::fs; a call to
fs::write is never the project’s own
fs::write, and crate::capsule::load() inside
context is never context::capsule::load; a
module path reaches a nested function only through a
pub use. A use written inside a function binds
only there, and an impl in another module counts for a type
only when it names that same type, so a module’s own same-named
Config no longer answers a::Config::default().
A bare call never binds to a method of the enclosing impl;
a method call on a receiver whose type is unknown stays unresolved
instead of binding to the one project method of that name; a typed
receiver must name a type the caller can see, so
reqwest::Client::get no longer binds to a local
net::Client::get and io::Error::kind no longer
binds to the project’s own Error::kind; and a
let that shadows a typed parameter clears its type, until
the end of its block. Typed function parameters and calls inside macros
now count as receiver evidence. Dropped calls are counted as unresolved.
A project whose only function identity was an unqualified name could
also fail the whole rebuild on a raw @method: hint; that is
fixed too.
A predicate provider could forge the evidence a gate rule
trusts. Providers in .phronesis/predicates/ —
which an agent can write through add_predicate_provider —
could emit_fact any predicate, so a two-line script
emitting signal_pass turned a low-confidence block into a
pass. Host-owned predicates (signal_*,
journey_*, rule_overridden, outcome, clock,
coverage, property, graph, AST, and hook content facts,
store_corrupt) are now reserved: a provider that emits one
fails, all of its facts for that event are dropped, and the pre-hook
blocks with a diagnostic naming the provider and predicate.
add_predicate_provider refuses a literal reserved emit up
front. Project vocabularies such as change_set_* are
unaffected. Only signal_* and journey_* are
reserved as whole namespaces; every other host name
(store_corrupt, context_confidence_band,
proof_outcome, …) is reserved exactly, so a provider’s own
store_opened or context_switch keeps working.
Upgrade note: an existing provider that emits a
reserved host-owned name now fails, which blocks every pre-hook
(post-hooks warn) until the provider is changed to emit a name of its
own — rename it.
Guard and provider scripts could call
eval. Only the artifact render engine disabled it,
although the crate documentation said none of the engines allowed it.
eval is now disabled in every Rhai engine, so a guard or
provider that uses it fails to parse (a guard fails closed).
A block rule whose __script__ guard errored
silently allowed the edit. A guard that failed at runtime (an
undefined Rhai function, division by zero, a non-bool
result) was dropped as “blocked” with a log line the hooks never
printed, so the hook exited 0 with no output. A guard error now fails
closed: the rule fires as matched (block → exit 2, warn → exit 1), its
message names the rule and the script error, stderr gets a
phronesis: GUARD ERROR line, and the action log records
guard_error on the consequence.
Rhai scripts errored on large host data. Rhai
sizes a value as a whole, so the injected facts array (for
a guard) or event map (for a predicate provider) counted
against the script’s 4 KiB string cap: a guard touching
facts over a realistic fact base, or a provider reading an
edit over 4 KiB, failed with “Length of string too large” — and failed
closed, blocking the edit. The sandbox limits are now the script’s own
budget on top of the injected data, so host data never trips them by
itself; a script that builds unbounded data on its own still fails
closed.
A __script__ guard could be judged before
the facts it reads existed. Guards were evaluated when the
rule’s triggering fact arrived and the activation then latched, so a
guard over facts asserted later — by a predicate provider, for instance
— saw an incomplete fact base and the verdict depended on assertion
order (a facts_count(...) <= 1 guard kept blocking after
a provider emitted two facts). Guards are now judged at fire time
against the final working memory; an activation whose guard is false is
dropped without being latched, so the next update_agenda
rediscovers and re-judges it, and agenda_snapshot lists
only activations that would fire now. A __script__
condition with no script text is now a guard error (fails closed)
instead of a guard that silently passes.
A rule removed and re-added under the same id never fired
again. remove_rule left the rule’s
fired-activation keys behind, so the new rule’s matches looked already
fired, and left its pending activations on the agenda, where firing one
failed with ProductionStateNotFound. Removing a rule now
clears both.
A rule with a typo loaded silently and allowed everything
it was written to stop. A rule meant to block
git push --force let the push through (exit 0) when its
verb was "Block", its phase was "Pre", its
when was empty, a v1 argument was not a string, a v1 rule
carried a second action or named a v2 verb as its
action_type, or a key was mis-cased ("Phase").
These shapes now fail closed at load, exactly like malformed JSON:
pre-check blocks, post-check warns,
codex-hook denies, phr-mcp audit exits
non-zero (it used to print the load error and exit 0, so
--fail-on block passed in CI), and the MCP
load_rules_file and add_rule tools return an
error. The message names the rule id, the field, the bad value, and the
allowed values. Duplicate rule ids within one file are rejected too.
Unknown predicate names are still accepted, because project Rhai
providers define predicates at run time. While the file does not load,
an edit (or Codex patch) whose target is that file is allowed with a
warning so an agent can repair it, and session-context /
interaction-context lead with the load error instead of
printing nothing.
The MCP server could erase every rule in a rules file it
could not load. On startup it loaded nothing from such a file,
and the next add_rule autosaved only the new rule over it;
a second call rotated the last good copy out of
rules.json.bak, and the now-valid file lifted the block
with every user rule gone. add_rule,
remove_rule, extract_rules and
save_rules now refuse while the rules on disk do not load,
naming the error and the file to fix, and never write or rotate
.bak; list_rules reports the error as
load_error (also with
PHRONESIS_NO_AUTOPERSIST). Once the file is fixed, the same
server reloads it — the repaired file wins over the copy the server held
— before the next write.
add_rule with an existing id, or extracting
the same guide twice, wrote a duplicate rule id. Under the
stricter loader that duplicate blocked every tool call.
add_rule now replaces a rule with the same id (keeping its
phase unless one is given) and says replaced;
extract_rules replaces by id; and every rules-file write
keeps only the last definition of an id.
A block rule whose message started with
? fired but did not block.
{"block": "?reason: force-pushing rewrites history"} — or
any message whose first word was a ?var no condition binds
— made the engine drop the action, so the rule matched and the hook
exited 0. A firing rule now always produces its consequence: an unbound
variable renders literally, and the hook prints a NOTE
naming the rule and variable when the rules load and again when the rule
fires. A message that is merely text (?? are you sure) is
left alone.
A bound short variable name could corrupt or hide a
longer one. A message like
"Function added in ?file (file var is ?f)" with only
?f bound rendered as
"Function added in <value>ile (file var is <value>)"
— substitution matched ?f as a substring of
?file — and the load-time diagnostic wrongly treated
?file as bound for the same reason, so no NOTE
warned about it. Substitution and unbound-variable detection now share
one tokenizer and match whole ?ident tokens only, so
?f never touches ?file.
Two warn rules could produce a spurious block.
Hook-derived fact ids were built by replacing punctuation with
_, so new_content_contains patterns
a.b and a-b both became
new_content_contains_a_b; an edit containing both was
blocked with duplicate fact id. The same
_-join let facts from different predicates collide
(function_clone_count for high_x versus
function_clone_count_high for x). Fact ids are
now escaped reversibly and cannot collide; identical patterns from
different rules still share one fact.
Some blocks left no trace in the action log.
When pre-check failed closed — a rules-file load error, a
fact-assertion error, a predicate provider error — it exited 2 without
writing a log entry, and the Codex adapter’s fail-closed denies were not
logged either, so log.jsonl could not explain them. Every
exit-2 entry and Codex deny is now logged with a blocked_by
list naming each reason: kind: "rule" with the rule id, or
kind: "fail_closed" with the message (and, for
pre-check, the failing stage).
A damaged code graph reported “fresh” and structural
rules silently saw nothing. Freshness compared only source-file
hashes, never the graph file, so a truncated, garbled, emptied, or
deleted .phronesis/graph.jsonl still printed “Graph is
fresh.”, the session line said “Code graph: current”, and structural
block rules kept block authority over whatever edges were left — or
quietly matched nothing. The index now records the graph file’s content
hash, written by the same rebuild or save; any mismatch reports the
graph as unverified in phr-mcp graph status,
get_code_graph_status (listed in
drifted_files), and the session context line, and
structural rules warn instead of block, with a notice naming
phr-mcp graph rebuild. Upgrade note: an
index written by an earlier version carries no graph hash, so the first
hook after upgrading reports the graph as unverified until
phr-mcp graph rebuild or the next hooked save rewrites
it.
An agent could approve its own verification
artifact. The rule that refuses agent writes to the
verification trust anchors lived only in a test fixture, so no
phr-mcp init installed it — and even that fixture matched a
directory the code never reads, letting
.phronesis/verification-allowlist.json,
.phronesis/verification.json (the
raw_execution opt-in), and shell writes like
echo > verification/templates/x.rhai through. The
llm pack in the default platform now blocks
Edit/Write/MultiEdit (Gemini
replace/write_file, and Codex
apply_patch, whose *** Move to: destination
was previously ignored) to all three anchors, matched on the path
relative to the project root (new project_path_is /
project_path_under facts), so
src/verification/templates/,
templates/verification/, or a checkout that itself lives
under such a directory is not caught. Shell commands that write an
anchor — redirects, tee, rm, in-place
sed/perl,
cp/mv/rsync with the anchor as
the destination, including through bash -c — now
warn, because shell matching is lexical and the spec
calls that seam advisory. The shell rule reads the new
bash_command_code_matches fact, the command with heredoc
bodies removed, so a commit message or a document that mentions an
anchor is not flagged; copying an anchor out, another project’s
.phronesis/, and lookalike files such as
fixtures/email-verification.json are left alone. That
heredoc stripping mistook a shell arithmetic <<
($((1<<2)), (( x <<= 1 ))) for a
heredoc operator, captured the arithmetic’s trailing digits as a bogus
delimiter, and swallowed every following line as its body — hiding a
real anchor write later in the same command from the scan.
<< inside an open ((/$((
arithmetic context is now tracked as a shift operator, not a heredoc
start. Re-run phr-mcp init --rules-only to pick the rules
up.
set_property_status could promote a property
without leaving an audit line, lose concurrent transitions, and rewrite
a store every reader rejects. A failed log.jsonl
append was ignored after the status had already been written; two
transitions at once raced on one unlocked read-modify-write and a shared
properties.json.tmp; and a store with an unsupported
version or a hostile id was rewritten as if valid. The transition is now
journaled before it commits (a journal failure refuses it and leaves the
store untouched), the update holds a lock with a unique temp file, and
the store is read through the same validator every reader uses, so a
rejected store is refused. PHRONESIS_NO_ACTION_LOG no
longer silences this journal line, an unknown property id no longer
creates .phronesis/, and a staging file left by a crashed
transition is cleaned up by the next one.
Rules could silently stop firing when an id contained
: or ,. The engine remembered fired
activations as rule:fact1,fact2 strings, so rule
a over fact b:c collided with rule
a:b over fact c, and a rule whose id contained
: never fired again after its fact was retracted and
re-asserted. Fired activations are now typed (rule, facts)
keys, and a retraction clears exactly the activations that used the
retracted fact.
Journal compaction could change the confidence band. Compaction kept only each subject’s latest outcome record, but confidence signals are “latest per kind” (compile, tests, each proof property, each bug). Dropping an older record could grant a proof signal that had been withheld (an older failing property vanished while a newer passing one survived) or drop a compile signal. Compaction now keeps the latest record for every signal each subject’s reader can see, and a randomized test checks that signals are identical before and after.
The Codex hook allowed tool calls that Claude’s
pre-check blocked. On a configuration error — a
rule naming an undefined journey selector, or a malformed
journey.json that a rule depends on —
codex-hook PreToolUse returned {} (allow)
while pre-check blocked. Both hooks now build their rule
network through one shared function, so Codex denies (and PostToolUse
warns) wherever Claude does.
scrub-payload no longer leaks
$HOME paths that contain spaces or JSON-escaped
slashes. A path such as
/Users/<name>/My Plans/q3.txt used to lose only its
first word to the placeholder, and \/Users\/<name>\/…
inside a captured tool output passed through untouched; both exited 0.
Path extent now follows a documented rule per context (path-keyed value,
quoted, bare free text with /-continuation), separators
match at any escaping depth with the placeholder written back in the
same spelling, and verify plus residual detection run over
separator-normalized text so an escaped residual fails the run. The
project root is matched as a whole component, so a sibling
…/project2 is no longer rewritten to
/home/dev/project2.
Governance switched off inside some git
worktrees. For a worktree created with
git worktree add --relative-paths, a hook run from a
subdirectory resolved the worktree’s relative gitdir
against the current directory, so it either found no
.phronesis (every rule skipped) or, in a nested checkout,
found an unrelated project’s rules. The gitdir is now
resolved against the directory holding the .git file, and
the main checkout root is canonicalized.
Stray .phronesis/journey/ directories
switched governance off for their subtree. Hooks fired from a
directory no project governed (and, before project-root discovery walked
up, from any subdirectory) created
<cwd>/.phronesis/journey/ holding only
inflight.lock, seq, events.jsonl
and events.lock. Root discovery stopped at the first
.phronesis/ it met, so every hook run from under one of
those strays found no rules and allowed everything, even with blocking
rules in the real project root. A root now counts as governed only when
.phronesis/rules.json (which phr-mcp init
always writes, --packs none included) or
.phronesis/loader.json exists; discovery skips anything
else, and pre-check, post-check,
claude-hook, codex-hook,
session-context and interaction-context
neither write nor inject anything in an ungoverned root (a
durable.md without a rules file is no longer injected).
Existing strays are not removed for you. To find them, run from your
repository root:
find . -type d -path '*/.phronesis/journey' -not -path './.phronesis/*' -exec sh -c 'p=$(dirname "$1"); [ -e "$p/rules.json" ] || [ -e "$p/loader.json" ] || echo "$p"' _ {} \;
and delete each printed .phronesis directory once you have
checked it holds nothing but journey/ (older builds also
left a log.jsonl). A printed directory with more state than
that is a copy-initialized worktree missing its rules.json;
it is now governed by its main checkout, so restore
rules.json there if it should govern itself. The Phronesis
repository had fourteen strays, among them crates/,
crates/phronesis/, crates/phronesis-mcp/src/,
crates/phronesis-mcp/tests/,
crates/phronesis-metrics/ and its src/,
tests/ and examples/,
docs/specs/, and .worktrees/.
Three behaviors change with the new definition of a governed root:
.phronesis/ holds real state (a
durable.md, say) but no rules.json used to
govern itself with no rules, so hooks under it allowed everything. It is
now skipped, and the nearest governed parent’s rules apply — an edit
there that the parent forbids is now blocked. Add a
rules.json
(phr-mcp init --rules-only --packs none) to keep the nested
project separate.phr-mcp init --hooks-only,
which writes no .phronesis/, is ungoverned: lifecycle
events are no longer recorded there and durable.md is no
longer injected. Run phr-mcp init without
--hooks-only to govern it.codex-hook PreToolUse with a
malformed apply_patch now answers {} (allow)
instead of denying, matching every other ungoverned tool call. A payload
that is not valid JSON still denies.phr-mcp kalpa start,
phr-mcp unit start and the MCP
submit_suggestion tool reported success in a directory no
project governed, creating .phronesis/journey/ and
log.jsonl there — the stray shape above — for a boundary no
hook would ever record against. They now fail with a message pointing at
phr-mcp init and write nothing.
The macOS verifier sandbox let the verifier write
anywhere. The sandbox-exec profile allowed every write and
named verification/ as a relative path Seatbelt never
matches, so a verifier run could write any file the user could —
directly, or through a daemon such as defaults write or
pbcopy — and could signal any of the user’s processes.
Writes are now denied except under a fresh per-run directory (the
artifact copy and TMPDIR), daemon lookups are denied,
signals may target only the verifier itself, and network is still
denied; the real Verus harness still proves under it.
The devcontainer tier ran whatever image the verifier
command named. The docker run line had no image
and no mount, so the first verifier word was pulled from a registry as
the image and any image that printed the summary line recorded
passed. The image now comes from
verification/templates/devcontainer.json, must be pinned by
digest (the tier is refused otherwise), is never pulled, and sees only
the artifact’s run directory, mounted read-only.
A hung verifier hung its caller forever.
Verifier runs had no time limit on any tier, although S9 promises one,
and a verifier that filled a pipe while its solver child kept running
could deadlock. Runs now get a wall-clock limit —
timeout_secs in .phronesis/verification.json,
default 300 — after which the verifier’s whole process group (z3
children included) is killed, a devcontainer run’s container is stopped
by name, and the result is recorded as timeout. Output is
drained concurrently.
The devcontainer tier failed on podman-only
hosts. Tier detection accepted podman, but the run always
invoked docker, so a host with only podman selected the
tier and then could not run anything. The run now invokes whichever
runtime the probe found, with the same confinement flags.
A tampered artifact ran as approved. Verifier execution trusted the hash the caller passed in, so edited bytes ran under an old approval, and a hand-edited allowlist entry with an empty hash or principal was accepted. Execution now computes the artifact’s SHA-256 from disk and refuses on a mismatch or when that hash is not approved, and the allowlist refuses to load when any entry is invalid, naming the entry.
Generated verification harnesses could smuggle file,
process, and network access past the rendered-body validator.
The validator scanned text with a hand-written lexer and plain substring
checks, so a char literal '"' hid every line after it,
include_str ! ("/etc/passwd") and
std :: fs :: read slipped through on whitespace, and
use std::{fs, process}; was never seen as
std::fs. Rust bodies are now lexed by a real Rust lexer and
must parse as a file (anything else is refused); denied paths and macros
are matched on tokens through whitespace, comments, grouped
use trees, globs, and renames (including
std::env!, std::os::unix::fs and
std::os::unix::net), macro definitions and $
metavariables are refused, bodies nested deeper than 64 brackets are
refused instead of overflowing the stack, and an interpolated value is
refused if any occurrence touches code rather than a string, char, or
comment.
Tests that run the phr-mcp binary reached
nothing in the code graph. An integration test that spawns its
package’s binary through env!("CARGO_BIN_EXE_<name>")
makes no Rust call the graph can follow, so test_reaches
gave it nothing — per-test coverage shows such tests executing code in
32–68 source files each — and code tested only through the binary looked
untested to no_direct_test and to the static half of
phr-mcp coverage select. A test that names the binary,
directly or through a same-file helper, now gets
tested_by(<bin>::main, test) and
test_reaches(test, <bin>::main); main’s
resolved-call closure is stored once as the new derived relation
bin_reaches(main, function), and
coverage select joins the two. The name resolves through
the new cargo_bin(package, name, target) relation, read
from [[bin]] tables, autobins,
src/main.rs and src/bin/ the way Cargo
discovers targets, and only within the test’s own package; a name with
no such bin target is dropped and counted as unresolved, never guessed.
On this repository 521 tests gain the edge and, through
main’s 1318-function closure, reach about 748 thousand
(test, function) pairs instead of 61 thousand, while the graph stays
near its old size (17.9 MB to about 18.8 MB): a rule that asks whether a
test can exercise a function must join test_reaches with
bin_reaches the same way. Cargo.toml is now a
graph freshness input, and saving one rebuilds the graph.
GRAPH_FORMAT bumps to 22, so existing graphs rebuild on
next use.
phr-mcp audit right after upgrading: it
exits non-zero and names the rule, field, and bad value if the file no
longer loads. Newly rejected shapes:
description or a comment
field) — only id, phase,
priority, audit, silent,
doc_excepted, binds, when,
then (v1: conditions, actions)
are allowed;then verbs, including those written by an older
MCP add_rule autosave as
"then": {"<custom action type>": ...} — use
block, warn, log, or
emit_capsule;"Block",
"Pre");when / conditions;priority, and null or
non-boolean audit, silent,
doc_excepted, binds;id; a non-string
phase or __script__/script; a
priority outside the 32-bit integer range;when together with conditions, or
then together with actions;action_type (including emit_capsule,
which needs the v2 then form), or an unknown key inside a
conditions[i] or actions[0] object;add_rule with an existing id appended a second copy instead
of replacing it. Keep the definition you want and delete the others; the
error names the id. Byte-identical copies (from extracting a guide
twice) still load as one rule, with a stderr warning..phronesis/rules.json and
.phronesis/loader.json (only those), and the MCP
rule-writing tools refuse rather than overwrite it. A broken layer file
outside .phronesis/ must be fixed by a human.Agent lifecycle events from Claude Code. A new
phr-mcp claude-hook <Event> adapter records prompts
(with fresh / mid_turn /
correction mode), inferred interrupts, turn stops, and
sub-agent start/stop into the journey journal and the action log.
phr-mcp init registers SubagentStart,
SubagentStop, Stop, and
SessionEnd, and repoints UserPromptSubmit and
SessionStart at the adapter; replacement is now keyed on
the command, so a hook you wrote yourself on the same event survives
init. session-context and
interaction-context keep working for settings files written
by older versions.
Commit detection from ground truth.
pre-check records HEAD before a shell call and
post-check compares it after, so a commit is recorded by
observing the repository rather than by matching command text.
HEAD movement is the ground truth and the tool’s exit code
is only a veto, so a host that reports no exit code at all — Claude
Code’s Bash is one — still gets its commits recorded,
marked detection: "no_exit_code". A commit sha the host
reports itself (Claude Code’s
tool_response.gitOperation.commit .sha) is kept as
host_sha, and stands in as
detection: "host_reported" when the HEAD probe
found no baseline. Gemini CLI’s invoke_agent tool is
derived into the same sub-agent start/stop pair.
Lifecycle events, foundation.
JournalRecord v2 with optional kind,
mode, host, turn,
agent, agent_type, kalpa; the
derive pass computes positional windows on tool records only, so
existing journey_* rules are unchanged; built-in
lifecycle:* and kalpa:* selectors; new
lifecycle module (LifecycleEvent, locked state
files, classify_prompt, detect_commit,
scrub_prompt); phr-mcp kalpa start|end|show;
prompt text is redacted from PHRONESIS_CAPTURE_DIR
captures, recursively, so a nested tool_input.prompt is
covered too. No host emits lifecycle events yet (adapters
follow).
Lifecycle events, Codex adapter.
phr-mcp codex-hook now records sub-agent start/stop (with
pairing and duration), prompts (with fresh /
mid_turn / correction mode and the scrubbed
text in .phronesis/log.jsonl), interrupts, and turn stops.
Interrupt and SessionEnd are handled and
registered by phr-mcp init; the SessionStart
matcher is now empty, so compact and fork sessions also get context.
Codex tool records take their session id from the shared
.phronesis/journey/session file, agreeing with the other
hosts. Payloads teed to PHRONESIS_CAPTURE_DIR have their
prompt text redacted.
Gemini CLI lifecycle hooks.
phr-mcp init now registers BeforeAgent,
AfterAgent, SessionStart, and
SessionEnd against
phr-mcp claude-hook <Event>, so prompts, turn stops,
and interrupts (inferred when AfterAgent is skipped by an
abort) are recorded. The BeforeTool/AfterTool
matcher is anchored and now includes invoke_agent, from
which Phronesis derives Gemini sub-agent start/stop pairs. Registrations
are replaced by command rather than by matcher, so a hook of your own
sharing a matcher survives init. A blocked
invoke_agent records a compensating
subagent_stop (blocked: true, zero duration)
so the already-durable start never dangles. The install output notes
that Gemini HTML-escapes injected context and skips project hooks until
the folder is trusted.
phr-mcp stats prints a lifecycle section — sessions,
prompts by mode, interrupts, sub-agents with median duration, and
commits with their confidence band — plus the active kalpa and the
retention boundary of the action log. --kalpa <name>
restricts the section to one kalpa, and --json now always
carries a lifecycle key alongside the rule totals.
phr-mcp kalpa show [name] reports the same counts
for a kalpa, open or closed, with the date it started and the retention
boundary.
phr-mcp journey renders lifecycle records in a
lifecycle table below the fact table (⟂ marker), with the
record’s kind/mode in place of a path, and prints the active kalpa in
its header. --lifecycle shows only those records;
--corrections lists the prompts that followed an interrupt,
oldest first, with their scrubbed text, under the same retention
boundary the other lifecycle reports print.
Prometheus:
phronesis_lifecycle_events_total{host,event,mode} and
phronesis_subagent_duration_seconds{host} (13 exponential
buckets, 1 s to ~68 min). No kalpa label and no agent_type
label — both are free text, user-typed and model-supplied
respectively.
The get_journey MCP tool takes an optional
include_lifecycle boolean (default false).
Left off, it returns exactly the bare array of fact rows it always has.
Set to true, it returns
{"facts": [...], "lifecycle": [...]} so lifecycle records
travel with the facts they explain. Prompt text is never included either
way, and phr-mcp journey --json is unchanged.
Work items and governed throughput.
phr-mcp unit start [<id>] [--spec <path>] names
the piece of work an agent is building and points it at the spec it is
built to; phr-mcp unit end closes it. Starting a unit while
one is open ends the open one first. pre_check and
post_check action-log entries now carry the open work unit,
so phr-mcp unit show [<id>] can join the journal and
the action log and report one item’s spec, window, kalpa, rules
evaluated / fired / blocked / warned with per-rule counts, grounded
evidence and confidence band, human interventions with their scrubbed
text, and commits — as text or --json.
phr-mcp kalpa show and
phr-mcp stats --kalpa <name> gain a work-item split
(explicit vs implicit), a governed count (committed,
with a rule evaluated against it, and a last-commit confidence band
above low), and interventions / work item.
Implicit work units keep working exactly as before and need no new file:
the two new lifecycle records carry everything. A work item can be named
from any of three places: the CLI, the known-bug registry
(phr-mcp unit start --bug <id> names it
bug-<id> and carries the registry’s cargo test name
and its spec — .phronesis/bugs.json entries gain optional
spec and title fields, neither of which
affects confidence scoring), or the agent itself, since the
submit_suggestion MCP tool now takes optional
spec and bug_id and records the same
unit_start event through the same code path. The spec
carries an example rule, suggest-name-the-work-item, that
nudges an agent to ask which bug or spec a session is for when no work
item is open; it is an example to copy, not a packaged rule.
scrub-payload residual-risk detection now treats the
third slash of a file:///absolute/path URL as a path
boundary, so a local absolute path inside a file URL is flagged as an
error. HTTP URL path tails (https://host/Users/...) are
still not flagged. Rescued from an unregistered July worktree.phr-mcp init in each project to get the new hook
registrations. Codex users then re-trust hooks via /hooks;
Gemini users must trust the folder, or project hooks are skipped. Until
you run init, session-context and
interaction-context keep behaving exactly as they do today
and no lifecycle event is recorded./hooks (or, for a headless run,
--dangerously-bypass-hook-trust). Check that before
debugging the adapter.agent_type in lower case whatever the host sent, so
Claude Code’s Explore is tagged
lifecycle:agent:explore. Write rule selectors in lower
case.s-window journey rules now scope to a
session. The .phronesis/journey/session file used
to be create-on-miss and never overwritten, so in practice a project’s
session id — and therefore every s window — spanned the
file’s lifetime. Each session-begin SessionStart now mints
a new id, which is what the window name always claimed. This is a
permanent semantic change, not a one-time boundary: review your
s-window rules, which now see shorter windows. Rules using
Nc or time windows are unaffected.lifecycle:* rule requires ≥
0.35: an older binary fails closed with UndefinedSelector
on the first such rule, taking every journey fact with it, and reads
lifecycle records as odd __lifecycle tool records that
shift positional windows.Non-destructive init --rules-only starter syncing
with a recorded baseline, local override/deletion preservation, and
conflict warnings.
Structural Rhai registration extraction and conservative forwarding-closure backing detection. Graph format 21 invalidates older extraction caches.
state reports it and clean --cache removes
it.phronesis-metrics crate derives bounded OpenMetrics
families from each project’s .phronesis/log.jsonl. Install
the CLI with --features metrics to enable one-shot, atomic
textfile, and loopback-only HTTP export through
phr-mcp metrics. Source paths and repository names are
never exposed as labels, rule-id cardinality is capped, and non-loopback
listeners are rejected unconditionally.audit_codebase and
phr-mcp audit evaluate builtin
facts_contain/facts_count
__script__ conditions against fresh per-file path and
extension facts. Unsupported Rhai or binding-dependent guards now
produce a diagnostic instead of making the affected audit rule appear
clean. Fixes #52.New opt-in python-patterns pack
(phr-mcp init --packs python,python-patterns; alias
py-patterns). Thirteen advisories derived from https://python-patterns.guide/, every one backed by a
new tree-sitter predicate in syntax/python.rs — no
substring or regex conditions anywhere in either Python pack. Warns:
global rebinding (python_global_statement),
globals()[...] = ... introspection assignment
(python_globals_subscript_assignment), three-argument
type(...) dynamic classes
(python_dynamic_class_creation), __new__-based
singletons (python_new_override shape
singleton), containers whose __iter__ returns
self (python_container_is_own_iterator),
multiple inheritance of concrete classes
(python_multiple_inheritance), *Mixin classes
with __init__ (python_mixin_with_init), static
delegation wrappers of 4+ forwarding methods without
__getattr__
(python_static_delegation_wrapper), mutable containers
assigned in a class body (python_mutable_class_attribute),
and == None / != None
(python_equality_with_none). Audit-only: other
__new__ overrides (Flyweight), isinstance
dispatch chains (python_isinstance_chain), and file-local
inheritance depth of 3+ (python_inheritance_depth). The
isinstance advisory requires a positive
if/elif chain dispatching on the same value
across non-builtin domain types, excluding independent input guards and
primitive/container validation. Each message cites the guide page and
states the limit of its heuristic. The base python pack is
unchanged; the guide remains a secondary source there (see the Deviation
note in SPEC-python-pack-expansion.md).
xcodebuild and swift build|test
are built-in toolchains for confidence scoring. A Swift project
running xcodebuild test through the Bash tool saw “tests
never registered”: cargo was the only built-in def, so nothing
recognized the command and the
Executed 55 tests, with 0 failures result never reached the
journal, leaving the subject at low under
confidence-low-blocks-commit. Both defs parse XCTest
summaries and Swift Testing’s
Test run with N tests … passed/failed line, treat
** BUILD FAILED ** and file:line:col: error:
as build failures (an XCTest assertion’s file:line: error:
is a test failure, not a broken build), and accept
** BUILD SUCCEEDED ** / ** TEST SUCCEEDED ** /
Build complete! as compile evidence when no exit code was
captured. phr-mcp toolchains lists them.
phr-mcp signal <compile|tests> <pass|fail>
records a confidence signal explicitly for the open work unit — the
escape hatch for a test runner with no toolchain def, or a run that
happened outside the hook. It journals the same outcome:*
tag the post-check hook stamps, so phr-mcp confidence and
the commit gate see it identically. Requires the confidence
pack.
Swift sources enter the code graph.
rebuild_code_graph now indexes .swift files
(crates/phronesis-mcp/src/graph/swift.rs), emitting
file_type, declares_module,
defines_fn (methods qualified by their type or extension
target), defines_test/tested_by for XCTest
test* methods and Swift Testing @Test
functions, and one imports edge from each file to its whole
unit. Because every file in a Swift target sees the target’s entire
namespace, tested_by resolution now accepts a whole-unit
import as visibility for every module in that unit. Previously Swift was
audit-only: syntax/swift.rs fed the
audit-swift-* rules, but query_code_graph for
file_type * swift or any Swift defines_fn
returned nothing, so no_direct_test could never vouch for
Swift code. GRAPH_FORMAT bumps to 19 so existing graphs
rebuild on next use.
calls(caller, callee) edges, canonicalized through the
unit-wide import the same way tested_by is (so
test_reachability follows Swift calls), and
calls_api(function, api) edges for a small risky-API
watchlist (SWIFT_WATCHLIST: fatalError,
preconditionFailure, exit,
unsafeBitCast, the Unsafe*Pointer
initializers/static members, Thread.sleep, and
semaphore/group wait). The structural panicking-API rule
can now fire for Swift.Package.swift targets are units. Discovery parses
.target, .executableTarget,
.testTarget (and .macro) declarations —
honouring an explicit path: and SwiftPM’s
Sources/<Name> / Tests/<Name>
defaults otherwise — so a file under Sources/App is
swift:App::… rather than
swift:project::Sources::App::…, and every file in a
.testTarget is file_type test whatever it is
called (UnitContext::test_target). import Foo
/ @testable import Foo naming another target in the
repository emits a whole-unit imports(module, swift:Foo)
edge, so tested_by and calls from a test
target into its production target canonicalize; imports of Foundation,
XCTest, or any module the repository does not define emit nothing. Xcode
.xcodeproj projects have no parseable manifest and stay on
the swift:project fallback with the filename/directory test
heuristic.Review hardening makes integration contracts
explicit. CI now runs the blocking Phronesis audit in addition
to formatting, clippy, and workspace tests. Codex documentation and
integration coverage pin its structured-JSON decision contract (process
exit 0, with logical 0/1/2 verdicts retained in the action log), and a
checked-in v1 bindings fixture proves reconciliation remains idempotent
before becoming stale. JSON Schema $ref resolution now
rejects repository-root escapes and absolute/drive paths while
normalizing Windows separators consistently.
Rust panic/debug starter rules now use syntax-tree
facts. The existing unwrap, empty-message expect,
todo!, panic!, unimplemented!,
and dbg! rule IDs now consume
rust_governed_invocation instead of source substrings.
Formatting and macro arguments no longer evade the rules, while
comments, strings, similarly named methods/macros, and non-empty
expect messages no longer cause false positives. Hook and
whole-tree audit paths share the same producer. The
&Box<T> parameter rule likewise uses the derived
function_param_is_box_ref fact and accepts whitespace
variants without matching strings or comments. The audit-only
environment-mutation rule now recognizes written
env::set_var and std::env::set_var calls
structurally; arbitrary import aliases remain outside syntax-only
evidence. The deny(warnings) crate-attribute blocker now
uses rust_governed_attribute, accepting whitespace variants
without matching comments or string literals. impl Deref
detection now consumes rust_trait_impl, recognizing
qualified trait paths and generic implementing types without matching
prose. The three match-arm audit rules now consume
rust_governed_match_arm, so multiline/whitespace variants
of empty None/Err(_) arms and
return Err(...) are recognized without matching comments or
strings. Two non-portable blockers for Phronesis-internal sync method
names were removed from the public Rust pack; downstream Rust projects
should not receive rules for this repository’s private refactor history.
The *_id: u64 and Rc<RefCell<_>>
audit rules now use parsed field/type evidence, including whitespace
variants and excluding prose. The String-ID twin intentionally remains
line-oriented so its /// field exemption keeps
working.
TypeScript any enforcement is
structural-only. The TypeScript starter pack retires the older
warn-any-in-src substring rule and keeps
warn-ts-explicit-any-ast as the canonical parser-backed
rule. Real any annotations still warn with function/count
evidence; comments and string literals containing : any no
longer produce duplicate or false-positive warnings. This intentionally
retires the legacy rule ID to avoid emitting two consequences for the
same annotation. The console.log warning also uses the
parser-backed ts_console_log_call fact, recognizing
whitespace variants while ignoring comments, strings, other logging
methods, and nested logger.console.log
expressions.
Swift crash/legacy rules now use syntax-tree
facts. try!, as!,
fatalError, mutable static var shared, legacy
geometry constructors, and legacy random APIs now consume
swift_governed_construct. Comments, strings, and
neighboring member calls stay silent; the hook and whole-tree audit
paths honor the fact’s construct discriminator.
Whole-tree AST audits honor literal fact arguments. Audit evaluation now applies the same constant-argument filtering as hook-time RETE matching, so rules sharing a structural predicate report only their intended construct.
The verify-before-commit nudge is
command-scoped. It now uses bash_command_matches
instead of generic content matching, so only a Bash command containing
git commit -m can trigger it. This remains lexical command
recognition, not a language-AST rule.
Low-confidence Git gate downgraded from
block to warn. The
confidence-low-blocks-commit starter rule (id unchanged) no
longer exits 2 on
git (commit|merge|rebase|cherry-pick|revert|pull) when
build/test/ known-bug evidence is missing or failing — it now warns
(exit 1), same as the medium band. Incomplete confidence evidence is
observability, not enforcement: a low-confidence Git mutation now always
proceeds. High confidence (3/3 signals) still passes clean, and
unrelated block rules still exit 2. See
docs/specs/SPEC-structural-rule-migration.md §“Confidence
gate severity” and the migration inventory at
docs/specs/INVENTORY-structural-rule-migration.md. Projects
that already ran phr-mcp init --packs confidence keep the
old blocking rule on disk until they re-run
phr-mcp init --rules-only --force --packs confidence
(existing project rule files are not rewritten automatically).
Breaking (library API):
graph::sync::SaveOutcome gained a diagnostics
field and is now #[non_exhaustive]. Rebuild
diagnostics record analysis a run did not perform — spec §8.2
requires the compiler provider to say that build scripts and procedural
macros were disabled rather than let a caller assume the analysis was
macro-complete. The struct is a return value from
rebuild/on_save, so this affects only code
that constructed it with a struct literal.
#[non_exhaustive] makes future additions non-breaking.
Migration: read the fields you need from the returned value; do
not construct SaveOutcome yourself.
A .rs/.json/.yaml
save no longer forces a full graph rebuild merely because
.phronesis/graph.toml exists. The rebuild now
triggers when the edited file is itself declared as a generated
artifact. Previously the config file’s mere presence made every save a
whole-repo rebuild, which would have made opting into ownership
enrichment silently expensive. Projects using data contracts should know
the trade: an unrelated .rs edit no longer refreshes
inferred bindings, which are heuristics recomputed at the next
rebuild. Explicitly declared bindings are unaffected.
Opt-in Rust ownership evidence (query-only). The
structural graph can now record where Rust source clones, filters,
awaits, mutates, and acquires synchronous locks, plus four bounded
relationships between those sites: filter_before_clone,
clone_before_await, read_before_mutation, and
lock_scope_ends_before_await. Every site carries a source
span, an evidence level, and a provider, and unavailable analysis is
recorded explicitly rather than read as a clean result. Enable per
project with [ownership.rust] in
.phronesis/graph.toml; with it absent the graph is
byte-identical to before. Query with
phr-mcp graph ownership <function-id-or-glob> or the
matching MCP tool. Findings are evidence with stated limits, not
verdicts: this ships no rule, creates no audit findings, and adds no
catalogue entry. Graph format 17 -> 18. Design after Schott,
Visualizing Ownership and Borrowing in Rust Programs
(Wuerzburg, 2024); see docs/OWNERSHIP-EVIDENCE.md.
Release-ready multilingual structural graph. CUE
packages now use canonical package identities with complete import
resolution; Lua, JSON, YAML, Helm 3, Rhai, Python, TypeScript, and Rust
extractors preserve repository-local closure and reject ambiguous
references. Query arguments support embedded * and
? globs consistently in the CLI and MCP tool.
Cross-language configuration contracts. Explicit
.phronesis/graph.toml bindings and bounded inference
connect CUE producers through tracked YAML/JSON artifacts to
deserializing Rust types, including wire-key mappings and conservative
unconsumed-key evidence.
Static test reachability. Canonical
tested_by edges record direct production-function evidence,
while test_reaches(test, function) follows resolved calls
transitively. External #[path] unit tests, public
re-exports, and bounded inherent-method calls participate without
guessing ambiguous names.
Fact and decision provenance. Consequences retain asserted and derived fact origins, and graph-backed rule findings can link their evidence to architectural decision records.
Local-state housekeeping.
phr-mcp state classifies authored, cache, history, runtime,
backup, and sensitive .phronesis state;
phr-mcp clean --cache removes only rebuildable graph
artifacts.
get_drift(source="claude_md") and
phr-mcp drift --source claude_md now scan root and
package-level CLAUDE.md and AGENTS.md files
with bounded, exclusion-aware traversal, deduplicate repeated
imperatives, and identify every source file on each finding. The frozen
phr-mcp claude-md-drift compatibility command remains
root-CLAUDE.md-only.phr-mcp audit --path <dir> and
audit_codebase(path) folded in graph-rule findings from the
whole project, so an audit scoped to one module returned violations from
unrelated files while files_scanned reflected the narrow
scope. Graph rules still evaluate against the entire graph — a covering
test may live anywhere — but reported findings are now filtered to the
requested path on segment boundaries.use crate::<module>; no longer collapses
onto the crate root. A use naming a module
directly had its last segment dropped as though it were an item, erasing
the real dependency and emitting a crate-to-crate self-edge.
Sibling-crate imports (use phr::Rule;) still resolve to the
crate root, where that is the correct target.use now produces
imports edges. The Rust extractor never walked
function bodies, so a file whose source visibly imported a module could
show no edge to it — understating fan-in and hiding import cycles.binds: false do not
bind.phr-mcp drift --source code and
get_drift(source="code") report stale rule bindings
alongside the existing prose, memory, and decision corpora.get_code_graph_status reports missing, fresh, stale, or
outdated graph state with generation, edge, file, and binding counts.
rebuild_code_graph performs a server-rooted full rebuild,
reconciles bindings, and records the generation transition in the action
log.phr-mcp archives for Linux x86-64, macOS Apple
Silicon, and Windows x86-64 after release-plz creates the matching
package release.init and every non-none pack
selection include llm, confidence,
journey, structural, and context;
language packs remain additive. none is an explicit,
mutually exclusive escape hatch.list_rules, list_facts,
get_agenda, get_consequences, and
get_action_log now return named JSON objects through both
structuredContent and compatibility text, avoiding SDK
failures on top-level arrays.drift,
stats, default-platform, graph lifecycle, and MCP envelope
surfaces.init::Pack is
#[non_exhaustive]. Adding a pack is this crate’s
most routine extension point — four landed in recent releases
(Confidence, Journey, Structural,
Context) and more languages are expected — but each one was
a semver break twice over: a variant added to an exhaustive enum, plus a
discriminant shift for every variant after the insertion point.
cargo-semver-checks flagged both on the
Context addition. Marking the enum non-exhaustive makes
future packs a non-event.
Migration: only affects code outside this workspace
that matches on Pack. Add a _ => ... arm.
Nothing inside the workspace changes, since
#[non_exhaustive] does not constrain the defining
crate.
.phronesis/context.json gets deterministic,
budgeted context packing instead of unconditional full-file reinjection.
Every renderable unit is an indivisible item measured with its headings
and separators, admitted only if it fits its kind ceiling, the shared
byte capacity, and — when configured — a soft estimated-token budget
(ceil(bytes / 3)). Current enforcement activity gets first
claim on the payload, then the kernel, then situational nudges, then
activity that overflowed its reserve..phronesis/kernel.md, the always-on
core. Written by init --packs context.
durable.md keeps its meaning as the session-level project
document and is never rewritten, repurposed, or shrunk..phronesis/nudges/*.md carry strict JSON frontmatter and a
static Markdown body, selected by positive facts through the ordinary
RETE engine. Bodies never interpolate fact arguments, so a filename or
tool payload cannot become a second-order prompt-injection channel. Only
four allowlisted predicates may trigger a capsule; adding one is a
reviewed code change, not configuration.phr-mcp context inspect | predicates | stats.
inspect is a true dry run: it writes no observation, so
reading the diagnostic cannot contaminate the data it reports. It lists
candidates, costs, ceilings, per-item omission reasons
(kind_ceiling, byte_capacity,
token_capacity, displaced_by_nudge), capsule
load failures, and fact-hydration failures. Human and
--json output are projections of one value.--packs base. Expands to every
language-agnostic pack —
llm,confidence,journey,structural,context — so the usual
shape is base,<your language>. Language packs are
deliberately excluded: several match raw substrings gated only by path,
so composing them produces cross-language false positives (the
TypeScript : any rule fires on Rust’s
: anyhow::Error).kind: "context" records in the existing rotated log carry
bytes, estimated tokens, per-kind omission counts, capsule ids, latency,
and a raw-truncation flag — no bodies, fact arguments, or user content,
and no claim about whether the model read or followed anything.## sections,
not blank lines. The section — heading, lead-in, and the list
under it — is the indivisible unit. Blank-line paragraphs let an
over-budget list be dropped while its lead-in survived, producing text
that promised content it did not deliver (“Three heuristic tools …:”
followed by nothing). A section is now delivered whole or not at
all.durable.md over 64 KiB now yields no context where it
previously yielded a 4 KiB truncation.cargo nextest, but its only
test_summary pattern was libtest’s
test result: line, which nextest never emits. The def
claimed the command, parsed nothing, and recorded no tests signal — so a
project gated on cargo nextest run stayed in the low
confidence band and had every commit blocked regardless of how green the
suite was. test_summary now accepts one pattern or several
(bare string or array, so existing toolchains.json defs are
unaffected); the first pattern that matches wins and all of its matches
are summed..phronesis/nudges/README.md was parsed as a
capsule. The file init itself writes produced a
load diagnostic on every hook invocation.read_recent parsed the entire action
log. Every hook asks for a handful of recent entries, and the
read parsed every line in the log plus its rotated predecessor to return
the last few — so the cost grew without bound as a project accumulated
history. A 3.8 MB log cost 16 ms per hook, and a real project measured
30 ms. The read now scans backward from the newest record and stops once
it has enough matching entries (counting matches, not lines, so
a filter excluding the newest records still reaches back far enough),
and only consults the rotated file when the current one cannot satisfy
the limit. This affects every caller, including the legacy context path,
stats, and trend.in with no
path. Rules that fire on a shell command rather than a file
edit log an empty file, which formatted as
- WARNED 36m ago: some-rule in — malformed text in the
prompt the model reads. The location clause is now omitted when there is
no path. This changes the legacy renderer’s output for command-rule
entries..phronesis/context.json the session and
interaction payloads are byte-identical to previous behavior, capsules
are not scanned, and no context observations are written. Pinned by
test. Two deliberate exceptions, both listed above: a
durable.md over 64 KiB is now ignored rather than
truncated, and command-rule activity bullets no longer carry a dangling
in clause.init never overwrites an existing
context.json, kernel.md,
durable.md, or nudges README.md.session.charter_max_bytes is defaulted, so
configuration written before the charter existed keeps loading.Measured on a base,rust fixture carrying this
repository’s own 3,292-byte durable file and five blocked edits,
comparing legacy against opted-in:
That latency figure describes a fresh project and is not
representative on its own. Context construction reads recent hook
decisions, and before the read_recent fix below that read
parsed the entire log: a real project with a 3.8 MB log measured 16 ms,
and one in the field measured 30 ms. Both are over the specification’s 5
ms target. The fix removes the dependence on log size; the figures above
should be read together with it.
context is therefore opt-in via
--packs, and is not yet part of the default pack set.per_test extraction remains libtest-only, so the
known-bug registry does not see cargo-nextest per-test results..ts,
.tsx, .mts, and .cts modules,
functions, imports, direct test-call coverage, and non-null assertions.
Resolution supports relative specifiers, index modules,
tsconfig.json baseUrl, and paths
aliases while excluding node_modules unconditionally.warn-import-cycle applies to resolved TypeScript module
cycles, and warn-untested-risky-call uses the narrow
! watchlist to report untested unchecked type assumptions.
Both remain advisory.query_code_graph advertised “Rust only” and a
stale identity form. The MCP tool description is what a model
reads to decide whether a tool applies, so both claims changed behavior
rather than merely being out of date. Python graphs build and query
correctly — init --packs structural in a Python project
produces a graph and graph query defines_fn returns
python:<dist>::<pkg>::<mod>::<fn> —
but the description reported the language unsupported. Its worked
example also still used the pre-0.23 crate::wme identity
form, so a caller following it queried a key nothing holds, received
zero results, and could reasonably read that as an empty graph rather
than a malformed query. The description now names both languages, gives
the current identity form with an example per language, and marks
calls_api as the Rust-only relation it is.structural pack, alias graph). A durable,
gitignored graph of architectural relations at
.phronesis/graph.jsonl, extracted by the
PostToolUse sensor and hydrated into the RETE network at
PreToolUse. Ships two warn rules:
warn-untested-risky-call (a production function calling a
panicking API with no direct test) and warn-import-cycle (a
module in an import cycle). Both join edited_file, so they
report the file in front of you rather than the whole repository on
every edit. See docs/specs/SPEC-triple-store-rete.md.<lang>:<package>[#<target>]::<module path>.
Rust resolves Cargo packages, compilation targets, and dependency
aliases including [workspace.dependencies] inheritance;
Python resolves distributions from pyproject.toml (PEP 621
and Poetry), both src/ and flat layouts, and imports across
sibling distributions in one repository.phr-mcp graph rebuild, graph status, and
graph query, plus the query_code_graph MCP
tool.event.file_rel for Rhai predicate
providers — the edited path in the repo-relative form the graph
keys files by, so provider-emitted facts can join graph facts on a path.
event.file_path remains the host’s absolute path..phronesis/graph.index records the identity scheme it was
built under. A graph built by an older version is reported as outdated
rather than fresh, and the next save rebuilds it — content hashes cannot
detect an identity change, because the files themselves do not
change.PostToolUse graph sensor never ran through
a real hook. A traversal guard rejected absolute paths, which
is the only form hosts send; the sensor was additionally gated behind
post-phase rules, which a pre-phase-only pack never has. Both are fixed,
and repo_relative now resolves symlinked project roots
(/var vs /private/var on macOS).#[cfg(not(test))] and
#[cfg(feature = "test-utils")] were classified as test
attributes, dropping production functions from
defines_fn and turning their calls into coverage edges.
cfg predicates are now parsed rather than
token-scanned.audit_codebase tool omitted the graph
merge the CLI performs, reporting zero structural debt
regardless of the graph’s contents and writing that zero into the debt
trend.phase: "pre" rules exclusively)
removed an accidental gate without adding a deliberate one, so a project
on --packs llm gained an unasked-for
.phronesis/graph.jsonl and a per-save extraction pass. The
graph’s own presence is now the opt-in signal.calls_api is deliberately empty for Python: there is no
defensible equivalent of Rust’s closed panic watchlist. Python projects
therefore fire warn-import-cycle only.tested_by matches by bare short name and
over-approximates coverage, so untested under-approximates.
That direction is chosen — a missed warning is recoverable, a false
“untested” verdict is not.warn. Promotion to block awaits
a second measured corpus.phr-mcp graph rebuild.phr-mcp codex-hook now implements the current
PreToolUse, PostToolUse, session, prompt,
compaction, and subagent contracts for Bash and
apply_patch; phr-mcp init safely merges
project hooks and project-scoped stdio MCP registration without
bypassing Codex’s /hooks trust review..phronesis/predicates/*.rhai can derive new RETE
facts from normalized hook events. MCP tools add, inspect, test, list,
and remove providers so agents can evolve the rule vocabulary alongside
rules. Multi-file operations expose a once-per-operation
event.files batch context before per-file evaluation; the
repository includes a dogfood change_set.rhai
classifier.interaction-context;
turn-context remains a compatible CLI alias and the old
Rust helpers remain deprecated wrappers. The unrelated markdown
set_section_context MCP workflow is unchanged.authored and use the payload-corpus envelope
instead of claiming unverified runtime capture provenance.Consequence::from_rule_firing takes
RuleFiringContext;
journey::derive::assert_facts takes
DeriveInput; outcomes::extract takes
ExtractInput; and
outcomes::adapter::extract_from takes
ExtractFromInput. These are breaking Rust API changes;
construct the corresponding context/input struct and pass it in place of
the former positional arguments.let consequence = Consequence::from_rule_firing(
RuleFiringContext {
rule_id,
predicate,
bound_facts,
kind,
},
&payload,
)?;
journey::derive::assert_facts(&mut network, DeriveInput {
project_root,
rules: &rules,
config: &config,
scope: WindowScope {
current_sid: session_id,
now_ts: now,
},
}).await?;
let facts = outcomes::extract(outcomes::adapter::ExtractInput {
root,
subject,
command,
output,
command_exit,
});
let (tags, subject) = outcomes::adapter::extract_from(ExtractFromInput {
project_root,
tool_name,
command,
output,
command_exit,
});PHRONESIS_CAPTURE_DIR — when set,
pre-check/post-check tee the raw stdin payload to
<dir>/payloads.jsonl before parsing (flock’d for
concurrent-hook safety, best-effort and off by default). This is how the
payload-contract corpus below gets refreshed against a real CLI’s
current payload shape.payload_scrub module +
phr-mcp scrub-payload <path> [--write] [--project-root DIR]
— anonymizes captured payloads for committing as fixtures. Operates on
JSONL in/out; --write backs the original up to
<path>.bak before overwriting;
--project-root defaults to the current working directory so
in-project paths survive scrubbing while $HOME, username,
session_id, transcript_path, and other
out-of-project paths are rewritten to deterministic, indexed
placeholders. A residual leak or unrecognized shape aborts the run
before anything is written.crates/phronesis-mcp/tests/fixtures/payloads/, each tagged
provenance: "authored" — hand-written approximations of the
real envelopes pending supersession by live captures via the tee above.
A contract runner replays every fixture through the real binary and
asserts rule liveness (a hook or tagger that silently no-ops now fails
CI) and journey-journal outcome tags. A companion hook-event registry
test suite pins init’s hook wiring to event names that
actually exist on each host CLI, including a regression pin on the
0.17.1 BeforeModelRequest incident.ToolchainDefs
(built-ins ∪ project .phronesis/toolchains.json, project
ids overriding built-ins) drive outcome detection via named-capture
regexes, with the command exit code as the authoritative build signal
and regex refinement layered on top. phr-mcp init scaffolds
example pytest/tsc defs; phr-mcp toolchains [--json] lists
the effective registry.command_exit capture. PostToolUse
payloads from shell tools (Bash,
run_shell_command) are probed for a numeric exit code
(exit_code/exitCode/returncode/code/status,
then a trailing exit code: N text fallback), journaled as
command_exit, and used to grade outcomes — a non-zero exit
with a test summary grades build-pass/test-fail (the pytest exit-1
case)..phronesis/journey/events.jsonl is bounded (16 MiB default,
PHRONESIS_MAX_JOURNAL_BYTES override, 1 GiB ceiling):
compaction retains the 10k-record tail plus the latest
outcome:* record per subject, atomically via temp+rename
with fd/inode revalidation so concurrent appends never land in a stale
file.CargoAdapter — cargo grading now flows through the same
toolchain-def registry as every other toolchain.init wired the turn-context hook into
.gemini/settings.json under
BeforeModelRequest, which is not a Gemini CLI hook event —
Gemini silently ignored it, so per-turn context injection (recent hook
decisions + durable directives) never ran in Gemini sessions. Now wired
under BeforeAgent, the per-prompt analogue of Claude Code’s
UserPromptSubmit. Re-running init (or
init --hooks-only) also removes the dead legacy
BeforeModelRequest key from existing settings. The emitted
hookEventName stays "UserPromptSubmit" for
both CLIs: Claude Code validates the field, Gemini reads only
additionalContext and ignores the echo.phr-mcp, phr, and phronesis-rhai all release as
0.17.0 — the workspace adopts lockstep versioning
([workspace.package] version); from this release one number
covers all three crates. (Previous: phr-mcp 0.16.2, phr 0.14.0,
phronesis-rhai 0.1.0; the jumps are version-line unification, not
breaking changes.)
hook.rs (1764
LOC) and syntax/rust.rs (1622 LOC) split into focused
submodules; main,
audit::run/run_profiled (deduped via a shared
core), and ~30 further functions decomposed below the let-count audit
thresholds. Audit debt drops 59 → 8 hits; the remaining 8 are
core-engine functions deferred to the embedded-consumer-gated engine
spec. Behavior-preserving; no public API changes. Implements
docs/superpowers/specs/2026-06-28-mcp-crate-decomposition-design.md.phr-mcp migrate-extracted-rules <path> [--dry-run]
— the salvage command deferred from 0.14.0. Rewrites pre-0.14.0
extract_rules output in place (with a .bak
backup): strips the bracketed extraction-time prefixes
([pattern], [anti_pattern],
[context], [problem],
[directive]) from messages, demotes block
actions to warn, and demotes to log any
extracted rule duplicating a structural Rust-pack rule (the SPEC’s
static keyword table: unwrap, clone, Deref, &String, &Vec,
thiserror). Extracted rules are detected by their
markdown_rule condition, so hand-written rules are never
touched. Idempotent. Implements the salvage path in
docs/specs/SPEC-extract-rules-defaults.md.lines: 1, 1, 1 — the
placeholder line number. FileAudit gains a
details field parallel to lines, rendered as
audit.rs — run (26 let bindings), run_profiled (32 let bindings)
in both output formats.src/.
audit-rust-let-binding-count-high /
-let-mut-count-high gain a
file_path_matches: "src" gate so examples, benches, and
tests are no longer flagged for let-count debt.audit-newtype-id-string honors doc
exceptions. The rule gains doc_excepted: true, so
a /// field doc marks an intentional string ID as an
accepted exception.serde_yml →
serde_norway. serde_yml 0.0.12 and
its libyml backend are archived and flagged unsound
(RUSTSEC-2025-0068 / RUSTSEC-2025-0067) with no fix coming.
serde_norway is the RustSec-recommended maintained
serde_yaml fork with the same API; the only call site (wiki
frontmatter parsing) changes crate path only. Removes
libyml from the dependency tree entirely.--- on its own line. The parser previously
accepted any line beginning with three dashes (----,
--- see appendix) as the closing fence, silently truncating
the YAML and leaking the line’s tail into the body. The fence search now
skips lookalikes; a page with no true fence reports “missing closing
--- fence” instead of parsing corrupted content. Pinned by
five new parser tests (lookalike lines, fence at EOF, CRLF
endings).phr-mcp 0.16.0; phr library bumps to 0.14.0 (engine changes this
round — new scripting trait, a removed method, and a new feature gate);
new phronesis-rhai 0.1.0.
Three changes that tighten the engine/embedding-host boundary ahead of a 1.0 line: an expressive scripting layer, removal of the last consumer-specific engine API, and a feature gate that makes the default public surface equal what the bundled MCP consumes.
phronesis-rhai crate + ScriptEval
trait. The core __script__ evaluator now lives
behind a ScriptEval trait
(ReteNetwork::with_script_evaluator). The new
phronesis-rhai crate provides
RhaiScriptEvaluator, a sandboxed Rhai implementation
(Engine::new_raw + StandardPackage,
operation/call-depth/string/array/map caps, sync)
supporting numeric comparisons and boolean combinators over fact
arguments — the guard expressions the builtin two-primitive DSL can’t
express. Scripts see facts (array of
#{predicate, args}) and bindings (map) and
must return bool; errors/non-bool are treated as a blocked
guard. CompositeScriptEvaluator routes builtin-DSL forms
(facts_contain/facts_count) to the builtin
evaluator and everything else to Rhai, so bundled packs and Rhai guards
coexist in one rules.json. Wired into phronesis-mcp behind
an off-by-default rhai feature (server
net::build_network seam).
Implements
docs/superpowers/specs/2026-06-01-rhai-script-evaluator-design.md.embedding-host cargo feature on
phronesis (off by default). Gates the ~10 public
ReteNetwork methods only an external embedding host needs
(restore_persistent_facts*,
execute_next_agenda_item, fact_ids_matching,
fact_count, facts_matching_predicate,
get_rules_count, get_wmes_by_condition, and
the instrumentation getters). The default surface equals what the
bundled MCP consumes, so the compiler enforces the symmetry. CI
exercises the feature config. Implements
docs/superpowers/specs/2026-06-13-embedding-host-feature-gate-design.md.ScriptEvaluator renamed to
BuiltinScriptEvaluator (implements
ScriptEval). ScriptEvaluator remains as a
backwards-compatible alias and the inherent evaluate still
returns ReteError, so existing callers are unaffected. The
misleading “Rhai” docstrings in core (the builtin is a hand-rolled DSL,
not Rhai) are corrected.ReteNetwork::get_persistent_facts and its
hardcoded PERSISTENT_PREDICATES — a downstream
consumer’s game-state vocabulary baked into a “domain-neutral” engine,
deprecated since 0.11 and now that the consumer has migrated onto
facts_matching_predicates, deleted. The remaining
consumer-flavored doc/example vocabulary in the engine and MCP fixtures
is neutralized. restore_persistent_facts* stay (generic
bulk-assert; now behind embedding-host). Implements
docs/superpowers/specs/2026-06-13-domain-neutral-persistent-facts-design.md.docs/loop-programming-guide.md) — writing recurring
/loop-driven agent workflows against phronesis, with captures from live
sessions in this repo.journey_derive scaling bench plus an
ADR recording the scaling behavior of journey fact derivation..phronesis/journey.json was fail-open: a stderr warning,
then the rule loaded anyway — and for absence-style rules
(== 0) the missing tagger looked like zero occurrences, so
the rule fired on every call. Configuration errors
(BadWindow, UndefinedSelector) now propagate —
the hook exits 2 (pre-check) / 1 (post-check) naming the offending rule
id and missing selector — while transient journal I/O errors stay
fail-open. See the decision page
2026-06-23-undefined-selector-rejection.md.Dogfooding-driven polish. The 0.13.x patch line was driven by
playtest bugs visible only after install; 0.14.0 closes the four
next-deepest friction points the same playtests surfaced. Compiled under
docs/specs/SPEC-0.14.0-dogfooding-polish.md.
journey_filtered_since_ge(target, counted, k)
aggregator — the existing journey_since_ge counts
distance over every record; a long Bash session could trip “8+ tool
calls since build” with no writes. The new aggregator emits a k-ladder
up to the count of counted records appearing after the most
recent target record. Rules can now express “8 writes since
last build” directly:
facts_count('journey_filtered_since_ge', ['build','write','8']) >= 1
with a write tagger keying on
change_type=edit|write|multiedit|replace|write_file. The
existing five aggregators are unchanged. See
docs/specs/SPEC-journey-filtered-since.md.confidence_enabled zero-arg marker
fact — asserted at every hook fire when
.phronesis/confidence.json exists, mirroring the
clock_facts.rs::business_hours_local pattern. Lets rules
condition on opt-in state via the existing
facts_count('confidence_enabled', []) == 0 absence form.
Generalizable: future packs can ship journey_enabled,
wiki_present, etc. using the same shape. See
docs/specs/SPEC-pack-opt-in-facts.md.confidence-low-blocks-commit and
confidence-medium-warns-commit now match
bash_command_matches: "git (commit|merge|rebase|cherry-pick|revert|pull)"
instead of the literal "git commit". Closes the
gate-bypass-by-merge hole surfaced during the journey-facts merge night
(5 of 6 commit-producing commands silently bypassed the gate). See
docs/specs/SPEC-gate-merge-commits.md.nudge-verify-before-commit self-deactivates
when confidence is on. The rule gained a second
when clause:
{ "__script__": "facts_count('confidence_enabled', []) == 0" }.
The confidence gate enforces the same call-chain-tracing discipline by
counting signal_pass facts; the nudge was redundant in that
mode and was double-warning on every git commit. Projects
without confidence are unaffected.extract_rules defaults action
warn, not block. A live invocation
added 27 block-action rules to a project rules.json overnight; with any
section context set, every pre-check fired 6 simultaneous
constraint_violations and exited 2 on every tool call.
Block is reserved for known-bad code shapes; pattern reminders are
advisory. See
docs/specs/SPEC-extract-rules-defaults.md.extract_rules strips the bracketed metadata
prefix ([pattern], [anti_pattern],
[context], [problem]) from the user-facing
message. Those were extraction-time discriminators leaking into
prose.phr-mcp init --packs confidence
before 0.14.0 carry the narrow gate pattern in
.phronesis/rules.json. Either re-run
phr-mcp init --rules-only --force --packs confidence
(rewrites the rule pack with the broadened pattern, backs up to
.bak) or hand-edit the two
bash_command_matches clauses.nudge-verify-before-commit rule
should add the second when clause to opt into the
supersession. Same --rules-only --force flow works.extract_rules and want to
salvage their extracted rules can apply the in-tree recipe in
docs/specs/SPEC-extract-rules-defaults.md §“Salvage path.”
A phr-mcp migrate-extracted-rules command is deferred to a
follow-up PATCH.extract_rules: per-pattern marker
conditions (Problem 3b), structural-rule skip-list (Problem 4a), and the
migrate-extracted-rules command. The umbrella spec scopes
0.14.0 to the action/prefix defaults; the rest rides a follow-up
PATCH.SPEC-gate-merge-commits open question
flags it.r) —
still phase 2 of SPEC-journey-facts.phr library version unchanged at
0.13.3. The engine wasn’t touched in 0.14.0; only
phr-mcp bumps. phr-mcp’s phr dep
stays pinned at 0.13.3.bash_command_matches taggers actually
fire. journey::tagger::tagger_facts built only
file/content facts and relied on a “tagger regex pass” implied by a
misleading comment but never implemented. The default build
tagger
({ "bash_command_matches": "cargo (build|check|test)" })
silently no-fired on every cargo invocation.
tagger_facts now walks taggers[*].when[*]
(including nested or clauses) collecting
bash_command_matches patterns, regex-matches each against
the bash command, and asserts one synthetic
bash_command_matches:<pattern> Fact per match — the
same pattern check_bash_command_patterns uses for top-level
rules (hook_facts.rs:316). Surfaced in a live playtest, not
in unit tests.HookPayload.tool_output accepts
tool_response as a serde alias. Claude Code’s
PostToolUse hook delivers Bash output under tool_response,
not tool_output. Without the alias, the field was
None / empty string, so compiled("") returned
true (no error patterns match → spurious
outcome:compile_ok) and
TEST_RESULT.captures_iter("") returned nothing
(outcome:test_pass never fired). Net effect:
confidence-scoring was wedged at “low / compile” for every real
cargo run, even when tests were green — the whole
gate-by-band feature was non-functional in production. Tests and
fixtures all passed because they synthesized payloads under
tool_output; only a live hook payload surfaced it. Backward
compatible with Gemini and existing fixtures..phronesis/journey/session was missing, the journey
fallback was the literal placeholder s-YYYY-MM-DD-fallback,
collapsing distinct sessions to the same id. Now
journey::current_sid reads-or-creates atomically in the
context::ensure_session_id format
(s-YYYY-MM-DD-<6 hex>); the placeholder is gone.current_sid
consolidated. Three independent implementations (in
hook, main, and
server::get_journey) coalesced into a single
journey::current_sid(project_root) helper. Same semantics,
one source of truth.confidence alongside journey. The
scaffolded CLAUDE.md previously enumerated journey
only.phr-mcp journey nudges on empty
config. When .phronesis/journey.json is missing or
empty, the CLI emits a stderr suggestion (“run
phr-mcp init --packs journey to scaffold one”) before
falling back to an empty config. The hook stays silent — fail-open is
advisory there, not user-facing.when was entirely __script__ clauses had no
alpha state, no terminal id, no p-state — they never reached the agenda,
because __script__ clauses are post-filters on activations
and with no other clause there were no activations to filter.
update_agenda now branches on
real_condition_count == 0 (count of
non-__script__ conditions per loaded rule) and, for
pure-script rules, evaluates the script clauses against the current fact
base with empty bindings, emitting an activation when every clause
passes. Dedupe key is <rule_id> — fire-once-ever, the
right semantics for threshold rules. Alpha/beta network and the
production network shape are unmodified; mixed-script behaviour is
unchanged. Surfaced by the journey-facts SPEC’s headline
auth-churn-without-tests rule, which is naturally two
__script__ clauses (facts_count(...) >= 5
AND facts_count(...) == 0). The journey_seen
anchor leaf added as a workaround is no longer required.journey_occurrence,
journey_count, journey_seen,
journey_since_ge, journey_distinct) with
windowed selectors (5c for last 5 calls,
30m/2h/7d wall-clock,
s for session; repo-lifetime r is phase 2).
Rule-driven derivation; the journal is the substrate, the predicates are
recomputed each cycle. See
docs/specs/SPEC-journey-facts.md.
.phronesis/journey/events.jsonl) writes a record per
post-check with subject + tags + monotonic seq. Tail-read for hot
queries (SUFFIX_HARD_CAP = 10_000 lines) and per-subject
read for outcomes folding.taggers[*].when clauses are the same DSL as rule
conditions. bash_command_matches,
new_content_contains, file_path_matches all
available. Project-defined via
.phronesis/journey.json.journey::derive::assert_facts; selector
validation rejects malformed journey config without exit-2.outcomes/ledger.rs
is gone; outcomes/cargo.rs now returns
(tags, subject) and the hook stamps them on a single
journal record. outcomes/derive::signals reads via
journey::journal::read_recent_subject. Confidence-scoring
behaviour is byte-identical; the storage is unified.SessionStart stamps
.phronesis/journey/session; pre/post-check read it.
PHRONESIS_NO_JOURNEY=1 disables both paths. Fail-open
throughout — corrupt journey.json or missing journal
degrades to “no journey facts,” never exit 2.phr-mcp journey [--json] [--explain <rule-id>]
renders the journey_* facts a derivation pass would assert
against the current journal, with --explain filtering to a
single rule’s dependencies.get_journey mirrors the same
table/JSON view so the agent can ask “what does my trajectory look like”
mid-conversation.phr-mcp init --packs journey writes a
starter journey.json and ensures it is tracked.phr and
phr-mcp move together; phr-mcp’s
phr dep bumps to match.git commit: does
it compile, do the tests pass, does it catch a known bug (a TDD test red
on the buggy baseline that goes green). See
docs/specs/SPEC-confidence-scoring.md.
build_outcome,
test_outcome, bug_check_outcome) behind a
per-toolchain adapter layer (cargo first;
pytest/tsc/go later emit the same neutral facts)..phronesis/outcomes/<subject>.jsonl) bridges the
stateless hook invocations; the pre-check re-derives
signal_pass facts and gate rules count them with the
existing facts_count(...) DSL (<=1 blocks,
==2 warns, 3 passes clean).git commit settles the open work unit..phronesis/bugs.json.phr-mcp confidence [--subject <id>] [--json] —
read-only band/signals report for the open work unit.phr-mcp init --packs confidence — writes the
commit-gate rules plus the .phronesis/confidence.json
opt-in marker and .phronesis/bugs.json registry, and carves
both back into .gitignore as tracked config.get_confidence (band/signals report) and
submit_suggestion (declare an explicit work unit, e.g. a
translation, and accrue signals to it)..phronesis/confidence.json; fail-open throughout, so
projects that haven’t enabled it are unaffected.ReteNetwork —
facts_snapshot, facts_matching_predicate,
facts_matching_predicates (predicate-set membership),
facts_matching (positional-arg filters),
fact_ids_matching, get_fact_by_id,
fact_count. Sync, owned results sorted by fact id, so
embedding hosts need not reach into wme_manager.list_facts MCP tool — the
existing predicate filter plus new predicates
(set membership) and arg_filters (positional
arg = value) params, backed by the fact-query API. Lets
coding agents query working memory by predicate set or argument, not
just list-all.bash_command_matches predicate — regex
rules over Bash/command-tool text, gated to command tools (file content
quoting the same text never fires). Ships two LLM-pack guard rules
(stage-explicitly, don’t-kill-build).python_bare_except,
python_mutable_default_arg,
python_function_param_count_high,
python_function_missing_docstring; TypeScript (TSX grammar
included): ts_explicit_any,
ts_non_null_assertion, ts_suppression_comment,
ts_function_param_count_high.phr-mcp audit and the audit_codebase tool now
explain a no-hits result when the cause is recoverable (no rules carry
audit: true, or the walker scanned 0 files) instead of
returning an empty shape indistinguishable from a failure.-D warnings + tests, on MSRV 1.90 and stable).Result<_, ReteError> replaces
Result<_, String> across the engine crate.
ReteError is a matchable enum (FactNotFound,
LockPoisoned, DuplicateFactId,
BindingConflict, …) implementing
std::error::Error;
From<ReteError> for String eases migration for
string-carrying hosts.DuplicateFactId); an identical re-assert is an idempotent
no-op. Previously a duplicate silently corrupted the predicate index
(the same fact was returned twice from
get_by_predicate).BinaryHeap-arbitrary; firing order is now
deterministic.ReteNetwork::get_persistent_facts — it
hardcodes consumer-specific predicates, which don’t belong in a
domain-neutral engine. Define your own predicate set and call
facts_matching_predicates(&YOUR_SET). Slated for
removal in 0.12.f1 no longer clobbers the refraction state of
f10 (was a substring match).get_memory_drift marks guidance
actionable only when it maps to an expressible predicate (named
command, file/path/code shape, or function shape); operational prose is
bucketed ambient. Actionable entries now also register coverage
from durable.md, so the drift list converges.SPEC-gate-merge-commits.md — broaden
the confidence gate’s bash_command_matches pattern from
"git commit" to
"git (commit|merge|rebase|cherry-pick|revert|pull)". Five
of six commit-producing porcelain commands currently bypass the gate.
Live-tested during the journey-facts merge night. PATCH-shaped change
for 0.13.x.SPEC-pack-opt-in-facts.md — pack-level
supersession via zero-arg marker facts. When confidence is
opted in, assert confidence_enabled at hook fire (mirroring
clock_facts) and condition
nudge-verify-before-commit on its absence via the existing
facts_count(...) == 0 form. Removes the double-warn on
every git commit for projects running both llm
and confidence packs. PATCH for 0.13.x.Pre-0.11 history (0.10.0 and earlier) is recorded in the git log and
docs/specs/. Notably, 0.10.0 added wiki-drift, the
block-pattern rules, and the v2 rule schema.