Teach
Five expert examples establish the food-aid coordinator’s expected behavior.
ApprenticeOS turns every expert correction into a permanent evaluation, repairs the workflow, reruns the complete test suite, and refuses to deploy until every test passes.
Regression #30 added
Expert correction stored as an eval
Minimal policy repair
avoidance phrase + known allergen → high risk
Release gate open
Policy v2 · 0 regressions
Not another chat interface. A focused control plane for a food-aid coordinator agent whose behavior becomes safer after every correction.
Five expert examples establish the food-aid coordinator’s expected behavior.
A real policy gap misreads “cannot eat peanuts” as a preference.
One expert correction becomes permanent regression #30.
A minimal compound-signal rule updates policy v1 to v2.
All 30 cases rerun, including every earlier behavior.
The release gate opens only at 30/30, then builds the Agent Pack.
Every future correction strengthens the system’s evaluation memory.
The correction outlives the conversation and protects every later release.
Every visible state is backed by policy files, evaluation records, and server-side actions. The demo remains fully reproducible without a paid API.
The input, bad output, expert answer, reason, and timestamp are retained together.
Model suggestions are Zod-validated, then checked against the deterministic safety policy.
The Agent Pack cannot be generated while even one regression remains.
Download policy, skill, API contract, evals, test summary, README, and changelog as one ZIP.
Then prove it before you deploy it everywhere.