Company eval packs and meta-harness scoring contract
This is the product-surface contract for cross-company evaluation. A company supplies an eval pack that says what good looks like; HSM-II supplies the shared runner, trace capture, KPI verifiers, scoring, frontier ranking, refiner artifacts, and promotion gates.
For the plain-language product layer on top of this contract, see Agent quality cockpit. The short version: enterprise-grade evaluation underneath, simple quality packs, agent health, run stories, and promotion gates on top.
The current implementation lives in:
scripts/meta-harness/pack_loader.py— loads pack YAML and normalizes tasks.scripts/meta-harness/meta_harness.py— runs Company OS agent-chat candidates and writesresults.json,summary.json, NDJSON traces, frontier state, and refiner artifacts.scripts/meta-harness/evaluate_turn.py— executes one task and combines process score with KPI score.scripts/meta-harness/kpi_verifier.py— deterministic Company OS state and trace verifiers.schemas/company_eval_pack.v1.json— versioned schema for the YAML contract.docs/company-os/examples/company-eval-pack.example.yaml— runnable example accepted by the current loader.
Eval pack YAML
Pack files are YAML documents that validate against schemas/company_eval_pack.v1.json. The loader currently requires company, personas, tasks, and scoring; newer optional fields are safe documentation for humans and future verifier plugins.
yaml
schema_version: hsm.company_eval_pack.v1
company:
id: 0b1aeb33-4f4e-4c70-8d83-a66d087e24c5
name: Agentsys Engineering
api_base: http://127.0.0.1:3847
personas:
- cto
- staff-engineer
- qa-release-lead
tasks:
- label: task_create_release_blocker
persona: cto
prompt: >
Create a high-priority task for the release lead to investigate the latest
production build blocker, include acceptance criteria, and set a due date.
scorer: task_state
weight: 1.25
slice_tags: [task-write, release, governance]
kpis:
task_created: true
assigned_to: qa-release-lead
has_due_at: true
priority_min: 7
spec_keywords: [production, build, blocker, acceptance]
safety:
requires_human: false
max_tool_calls: 8
- label: memory_grounded_recent_work
persona: staff-engineer
prompt: >
Summarize what we have been working on recently using Company OS memory.
scorer: memory_grounding
weight: 1.0
slice_tags: [memory, grounding]
kpis:
memory_tool_called: true
grounded_in_memory: true
answer_keywords: [Company OS, harness]
- label: validate_delivery_tools
persona: qa-release-lead
prompt: Run validate-delivery and give a pass/fail verdict with the blocking reason.
scorer: composite
weight: 1.0
slice_tags: [tool-use, qa]
kpis:
tool_sequence_contains: [validate-delivery]
answer_keywords: [pass, fail, blocker]
min_answer_length: 80
scoring:
process_weight: 0.35
kpi_weight: 0.65
min_mean_score: 0.75
min_finalize_rate: 1.0
max_error_rate: 0.0
max_mean_latency_s: 90
pareto_filter: true
promotion:
require_no_regression: true
shadow_runs: 3
rollback_on: [quality_drop, finalize_drop, safety_failure]Run it with:
bash
python3 scripts/meta-harness/pack_loader.py docs/company-os/examples/company-eval-pack.example.yaml
python3 scripts/meta-harness/meta_harness.py --pack docs/company-os/examples/company-eval-pack.example.yaml --iterations 1Field semantics
| Field | Meaning | Current implementation |
|---|---|---|
schema_version | Versioned contract identifier. | Documented by schema; current loader tolerates packs without it for backward compatibility. |
company.id | Company OS company ID used for API calls and task context. | Overrides the default --company-id unless CLI explicitly sets a different company. |
company.api_base | Company OS API base for state verifiers. | Passed to evaluate_turn.py --api-base. |
personas | Allowed persona slugs for tasks. | pack_loader.py rejects tasks whose persona is absent. |
tasks[].label | Stable task ID used in trace filenames and reports. | Must be unique in the pack. |
tasks[].prompt | User-facing task prompt sent to agent-chat. | Passed directly to the Company OS chat route. |
tasks[].scorer | Verifier binding. | Must be one of task_state, memory_grounding, tool_sequence, stigmergic, answer_keywords, or composite. |
tasks[].kpis | Scorer-specific assertions. | Passed to kpi_verifier.py. |
tasks[].weight | Relative importance. | Loaded today; weighted aggregation is a next implementation slice. |
tasks[].slice_tags | Analysis slices such as memory, tool-use, finance, or handoff. | Schema-defined; slice aggregation is a next implementation slice. |
scoring.process_weight | Share of final turn score from process/trace behavior. | Passed to evaluate_turn.py. |
scoring.kpi_weight | Share of final turn score from post-run KPI verification. | Passed to evaluate_turn.py; weights must sum to about 1.0. |
scoring.min_mean_score | Offline quality gate. | Applied by meta_harness.py through --min-mean-score. |
scoring.min_finalize_rate / max_error_rate / latency and cost caps | Promotion gates for reliability and efficiency. | Partly implemented (min_finalize_rate CLI exists); remaining gates belong in the promotion command. |
Scorer contract
Each scorer returns a float in [0.0, 1.0]. Deterministic checks are preferred over LLM judges when the Company OS state or NDJSON trace can answer the question.
| Scorer | Verifies | Example KPI keys |
|---|---|---|
task_state | A task was created or changed with the expected assignment, priority, state, due date, and specification content. | task_created, assigned_to, has_due_at, priority_min, spec_keywords, task_state_after |
memory_grounding | The answer used Company OS memory tools and overlaps with real shared memory. | memory_tool_called, grounded_in_memory, answer_keywords |
tool_sequence | The trace used required tools, exact tool order when needed, and non-hallucinated tool keys. | tool_sequence_contains, tool_sequence_exact, no_hallucinated_tool_key |
stigmergic | The run left useful notes or coordination residue in the Company OS graph. | note_written, note_keywords |
answer_keywords | The final answer contains required terms and enough substance. | answer_keywords, min_answer_length |
composite | Runs all applicable deterministic scorers and averages them. | Any combination of the keys above |
Turn score
For a task from a pack:
text
process_score = evaluate_turn trace score
kpi_score = deterministic verifier score
turn_score = process_weight * process_score + kpi_weight * kpi_scoreThe process score covers agent-chat behavior such as finalization, error state, tool evidence, and basic trace quality. The KPI score checks business outcomes after the turn by querying Company OS APIs and inspecting the NDJSON trace.
Aggregation contract:
text
candidate.mean_score = weighted mean(turn_score by task.weight)
candidate.finalize_rate = finalized turns / total turns
candidate.error_rate = error turns / total turns
candidate.mean_latency_s = mean(latency_s)
candidate.mean_tool_calls = mean(tool_calls)
slice[tag].mean_score = weighted mean over tasks containing tagThe in-tree aggregate() function currently computes unweighted means over result rows. Weighted means and slice aggregates are the intended next patch now that pack fields are explicit.
Frontier and promotion
The meta-harness should rank candidates on a Pareto frontier, not just a single average. A candidate is promotion-eligible only if it satisfies all hard gates:
mean_score >= scoring.min_mean_scorefinalize_rate >= scoring.min_finalize_rateerror_rate <= scoring.max_error_ratemean_latency_s <= scoring.max_mean_latency_swhen setmean_cost_usd <= scoring.max_mean_cost_usdwhen cost telemetry exists- no required safety assertion failed
Among eligible candidates, Pareto dimensions are:
text
maximize: mean_score, finalize_rate
minimize: error_rate, mean_latency_s, mean_cost_usd, mean_tool_calls when tool count is a cost proxyPromotion flow:
- Run the pack offline against a candidate.
- Write
results.json,summary.json, per-task NDJSON, and refiner artifacts. - Score deterministic KPIs and record failure tags.
- Compare against baseline and current frontier.
- If hard gates pass, run shadow or limited rollout according to
scoring.promotion.shadow_runs. - Promote only through
company_os_promote_delta.pyor a successor service that writes a promotion record and rollback condition.
The safety invariant from docs/EVAL_AND_META_HARNESS.md still applies: eval-side winners do not change live runtime behavior until an explicit integration maps them into runtime prompt policy, skill dispatch, memory policy, or guarded Company OS configuration.
Failure taxonomy
Every failed or weak turn should be tagged so the result is actionable:
| Tag | Trigger |
|---|---|
tool_missing | Required tool absent from trace. |
tool_wrong_args | Tool was called with missing or invalid key arguments. |
state_not_changed | Expected Company OS object was not created, updated, routed, or resolved. |
ungrounded_answer | Answer lacks required memory/data grounding. |
handoff_pending | Run delegated work but counted pending handoff as completion. |
finalize_missing | Agent did not emit the expected final response. |
latency_over_budget | Runtime exceeded pack latency cap. |
cost_over_budget | Cost telemetry exceeded pack cost cap. |
safety_gate_failed | Safety assertion failed or required human gate was bypassed. |
Failure tags should flow into summary.json, proposer context, refiner deltas, and promotion records so the loop can improve specific behavior instead of chasing opaque scores.
Governance & instrument integrity (the layer external eval frameworks omit)
Every published agent-eval framing (offline/online, Pareto, failure taxonomy, promotion service) assumes the verifier is ground truth. This program proved on real data that it is not: task_state read 0.25 confidently for an extended period and was 100% scaffolding artifact. The platform's distinguishing contract is therefore not "we score agents" — it is "the score is attributable and you can tell when it is lying before you act on it."
The three-place scorer invariant
scorer is one contract enforced in three places; drift between them is a governance defect, not a doc nit:
schemas/company_eval_pack.v1.json— the documented enum.scripts/meta-harness/pack_loader.py:VALID_SCORERS— the loader gate.- Postgres
CHECK ck_company_eval_packs_scorer— the persistence gate.
A scorer added to one must be added to all three in the same change. (This section exists because the schema had drifted out of sync with the live code and CHECK — found and reconciled rather than papered over.)
Provenance / drift units
A score is only governable if attributable to the revision that produced it. Three parallel drift units, same discipline:
company_eval_runs.scorer_version— git sha of the harness that scored.company_memory_entries.policy_version— memory-policy sha (Phase M0).company_eval_runs.skill_version— content hash of the skill body/path (Phase S0). NOTE:company_skills.bodyis empty in some companies (content lives on disk atskill_path); S1 must hash the real on-disk content, not a naivesha(body)which collapses to one constant.
Runs across different versions of any unit are not directly comparable; the dashboard surfaces a drift banner instead of trending across them.
Falsifiable governance, human-checked
company_eval_kill_criteria— pre-registered falsifiable conditions evaluated against run history; fire → dashboard flag, never auto-revert.company_eval_reviews— sampled human review; the override rate is the measured convincing-vs-correct signal (≥20% → scorer suspect). An override cannot be recorded without stating the correct verdict.
The online / shadow / auto-promotion wire is deliberately GATED
External guidance lists "promotion service → shadow → online evals → auto rollback" as a primitive to build next. This program's evidence says the opposite for any high-stakes domain: an automated promote/rollback loop on a verifier that can be confidently wrong compounds bad learning at machine speed (the Continual-Harness "monotonic accumulation" property, double-edged).
The substrate exists and its interlocks have fired on real data. It is Panel-1 (human-in-the-loop) by design: the refiner proposes (apply_policy: review_only), kill-criteria flag, humans promote. Enabling autonomous promotion is a customer risk decision taken only with override-rate low and kill-criteria/drift green — not a default and not a "next primitive." Selling or shipping it as automatic would be the exact failure this contract exists to refuse.