Skip to content

chore(isolation): formalize JSON Schema for acq-provider-facts/v1 and acq-neutral-model-catalog/v1 #435

Description

@wz-gsa

Summary

Fast-follow from #424's review (nexus-agents 3-role consensus vote, unanimous fast-follow). Two new JSON shapes were introduced without formal JSON Schema files:

  • acq-provider-facts/v1 — the provider self-registration facts file (/var/lib/acq/models/providers/<providerId>.json).
  • acq-neutral-model-catalog/v1 — the neutral model-catalog shape (usai-provider's scripts/normalize-catalog.mjs output and its vendored model-catalog-snapshot.json).

Both currently exist only as inline JSON examples in ADR 0004 (integrations/isolation/docs/decisions/0004-neutral-model-provider-discovery.md) and as this implementation's literal output.

Why deferred, not skipped

Per the consensus vote: writing additionalProperties: false-strict schemas now, while ADR 0004 is still unmerged and before a second real producer/consumer exists (#425, the models orchestrator kit), risks two sources of truth diverging before either is settled. The natural forcing moment is when #425 actually reads these shapes — that's when required/additionalProperties decisions can be made against real usage instead of guessed, following the existing precedent of integrations/providers/usai/catalog.schema.json (JSON Schema draft 2020-12, validated in CI).

Scope

  1. Write acq-provider-facts.schema.json and acq-neutral-model-catalog.schema.json under integrations/isolation/schemas/ (or an equivalent location — follow this repo's existing schema-placement convention), matching catalog.schema.json's style ($schema, $id, additionalProperties: false, required, per-field description).
  2. Wire CI validation for both (the existing schema-validation job that already validates catalog.schema.json should generalize cleanly).
  3. Validate the existing usai-provider kit's facts file, normalizer output, and vendored snapshot against the new schemas — retrofit, not a design change.

Security condition carried forward from the review (binding for whoever reads these shapes, not just whoever writes the schema)

The Security Engineer vote on #424 was explicit: deferring the schema file is only acceptable because it is not the actual security control — the control is defensive parsing at every READ site. Whoever builds #425 (the first real consumer) MUST NOT trust either JSON blob unsafely: no interpolating field values into a shell command or filesystem path without allow-list validation, and a hard cap on document size/array length to avoid resource exhaustion from a malformed or oversized file. This issue's schema files should encode those constraints (e.g. pattern on providerId/keyEnv, maxItems on models[]) so the schema itself becomes part of that defense, not just a shape check.

References

Activity

wz-gsa commented on Sep 24, 2026

@wz-gsa
ContributorAuthor

Picked this up and stopped before writing any schema, because the investigation turned up a blocker in the issue's own scope plus two findings worth acting on separately. AI-assisted (OpenCode), advisory. I wrote this issue, so this is me correcting it rather than reviewing someone else.

Blocker: step 3 can't be done on main

The issue says:

  1. Validate the existing usai-provider kit's facts file, normalizer output, and vendored snapshot against the new schemas — retrofit, not a design change.

None of those three artifacts are on main. In a fresh clone at a717b12:

$ grep -rn 'acq-provider-facts\|acq-neutral-model-catalog' .
(no hits)

A PR on main today would add two schemas with zero instances to validate. That is the thing this issue was written to prevent, so I'd rather fix the issue than satisfy it literally.

That exact failure mode is already in this repo

integrations/providers/usai/catalog.schema.json is validated by a hand-rolled JS validator in build-catalog.mjs:604-684, whose own comment explains why: "ajv is NOT a repo dependency ... so we hand-roll a check." But:

$ grep -rn 'build-catalog' Makefile .github/workflows/*.yml .pre-commit-config.yaml
(no hits)

validateCatalog is exported at :845 and imported nowhere. No test calls it. check-json in pre-commit is syntax-only. So that schema is enforced only when a human manually runs the script. It also re-hardcodes its constraints in JS (:616, :641, :666) and reads only schema.properties keys and schema.required from the file — so the schema and the validator can drift silently, and the schema isn't the source of truth.

Whatever #435 adds has to be wired into something CI actually runs, or this repo gets a second unenforced schema.

The wiring that works, from the observed pattern

Every schema that is enforced lives in schemas/ and is validated by Python jsonschema (already pinned, pyproject.toml:9):

Schema Validator Reached from
skill.schema.json validate_frontmatter.py:284 validate_repo.py:34 → make validate → ci.yml:171
network-tier-baseline-v1 validate_network_tiers.py:33 validate_repo.py:37 → same
kit-hybrid-v1 validate-kits.py:132,317 own ci.yml:180 step + pre-commit
usai/catalog hand-rolled, unrun nothing

Three for three: in schemas/, validated by Python, enforced. The one that isn't in schemas/ is the one that isn't enforced. So: put both new files in schemas/, add one tuple to validate_repo.py's validator list, and there's no ci.yml change at all. validate-kits.py:317 is the repo's only Draft202012Validator.check_schema call — worth copying, it catches a typo'd schema at the gate.

Two things found on the way, both separable from this issue

1. A degraded-path defect in validate_network_tiers.py:49-62. jsonschema is imported inside a try, and on ImportError the fallback checks only the host regex:

    except ImportError:
        pat = re.compile(schema["$defs"]["entry"]["properties"]["host"]["pattern"])

…then :87 prints the same ✓ network-tier baseline valid either way. A missing dependency silently downgrades full schema validation to one regex with no "I could not check" state — and jsonschema==4.26.0 is a hard pinned dep, so the fallback isn't even reachable in a correct install. Related: validate-kits.py:320-322 returns 0 on "No kits found", which is the no-evidence-is-a-pass shape. The new validator should do neither. Happy to file this separately if you'd rather keep #435 clean.

2. The cost unit is documented both ways in catalog.schema.json. Five sites, four wrong:

:135  "Per-million-token pricing. All fields optional."     <- correct
:139  "Cost per input token (opencode cost.input)."         <- wrong
:143/:147/:151  likewise per-token for output/cacheRead/cacheWrite
:110  "Base per-token pricing."   <- $ref sibling, overrides at the use site

The data settles it: catalog.json's claude_4_5_sonnet has cost.input: 3, i.e. $3 per million (per-token would be $3/token). Provenance agrees — build-catalog.mjs:253-279 copies models.dev values with a pure key rename and no arithmetic, and models.dev publishes per-million.

It's documentation-only today: grep for * 1000000, / 1e6, * 1e6, 1_000_000 across every .mjs/.py returns zero hits; every consumer passes the number through verbatim (prime-agent.mjs:43-51, opencode.mjs:50-57, sync-usai-models.mjs:515-520). But the schema is the only place the unit is written down and it says two different things, so the first consumer that actually computes a cost gets a coin flip. This is also why I'd put the unit in the field name for the new catalog schema (costPerMillionTokens) rather than repeat a bare cost.

Proposed rescope

Replace step 3, and sequence the rest:

  1. Now, on main: write both schema files in schemas/, add the validator + a tuple in validate_repo.py, and add scripts/tests/test_validate_model_provider_schemas.py validating in-memory instances (positive + negative, matching the test_validate_kits_* pattern). Also mirror test_test_cases_ci_wired.py with a 3-line assertion that validate_repo.py lists the new validator — that's what stops this becoming the second unenforced schema.
  2. When docs(isolation): add ADR 0004 for neutral model-provider discovery #423 merges: the ADR becomes the shape's source of truth, and the schema descriptions cite integrations/isolation/docs/decisions/0004-neutral-model-provider-discovery.md (per the Durable-References rule — not #435).
  3. When a real producer merges: point the validator at the actual instance files. That is the retrofit step 3 wanted, and it becomes possible then.

One open question that blocks step 1's field set, and it's the reason I'm not just writing the files: the shape in ADR 0004 is {schemaVersion, providerId, host, baseUrl, modelsUrl, keyEnv}, and I raised evidence on #436 that several of those diverge from universal practice — env is an array in all 223 live models.dev providers (14 need more than one variable, so scalar keyEnv can't express Azure/Bedrock/Vertex at all), everyone uses id not providerId, and nothing has a host field. Writing additionalProperties: false around the current shape would freeze that in, which is precisely the "two sources of truth diverging" risk this issue was deferred to avoid.

So: should the schema encode ADR 0004's shape as-is, or should the shape revision land first? I don't think I should decide that unilaterally — ADR 0004:332 says the field set is a cross-repo contract surface. Say which and I'll build it.

Meanwhile the security constraints from the original issue are unaffected and I'd encode them either way: pattern on the id and env-var names, maxItems on models[], format: uri plus an ^https:// pattern on URLs, and additionalProperties: false for authoring — with the consuming side documented as ignore-unknown-fields, per the survey.

wz-gsa commented on Oct 1, 2026

@wz-gsa
ContributorAuthor

I went to write these two schemas and stopped, because the work surfaced a conflict I created and a false premise in this issue's own text. Four questions below, each a decision I do not think is mine to make.

1. This issue's stated precedent is not enforced anywhere

The scope says to follow integrations/providers/usai/catalog.schema.json, described here as "JSON Schema draft 2020-12, validated in CI." Measured:

catalog.schema.json  -> referenced ONLY by scripts/build-catalog.mjs
                        NOT in .github/workflows/, NOT in Makefile, NOT in package.json
validateCatalog()    -> runs only inside build-catalog's main(); nothing in CI calls it
                        exported, but no test imports it

So the repo has a schema file that is documentation, and enforcement that is a parallel hand-rolled implementation in build-catalog.mjs:623, written that way because "ajv is NOT a repo dependency".

To be fair to it: I checked for drift and there is none today. The validator derives top-level allowed keys from schema.properties and required from schema.required, so the top level cannot drift; only the nested gateway/vendors/models key sets are hardcoded, and all three match the schema exactly. catalog.json passes both the hand-rolled validator (0 errors) and real jsonschema (0 errors). The defect is the absent gate, not a current mismatch — I had initially mis-grepped this as duplicated key lists and want to correct that.

And the hand-roll's stated justification no longer holds:

pyproject.toml:9   "jsonschema==4.26.0"     <- already pinned, already used by validate_repo.py

ajv is genuinely absent (so the Node-side hand-roll was reasonable), but a real schema gate needs no new dependency — it can run python-side like every other validator in this repo. I verified jsonschema.Draft202012Validator against catalog.json: 0 errors.

2. The precedent schema does not meet the bar this issue sets

This issue's security condition asks the new schemas to encode "pattern on providerId/keyEnv, maxItems on models[]". The schema it points to as the model has:

models.id              {'type': 'string'}      <- no pattern
models.vendor          {'type': 'string'}      <- no pattern
models maxItems:       NONE

So "follow the existing precedent" and "encode these constraints" are in tension. Should usai-model-catalog/v1 be retrofitted in the same pass, or left alone?

3. ADR-0004 as written would invalidate #436's shipped artifacts — and that is my fault

This is the blocking one. ADR-0004 (#423) says:

Both a normalizer's output and the aggregate carry schemaVersion: "acq-neutral-model-catalog/v1".

Per-entry provenance is required … Each entry states which providerId it came from and whether it came from a live refresh or that provider's vendored snapshot.

Measured against what #436 actually ships:

normalizer output   envelope: models,providerId,schemaVersion   model[0]: contextWindow,cost,id,maxOutputTokens
shipped snapshot    envelope: models,providerId,schemaVersion   model[0]: contextWindow,cost,id,maxOutputTokens
provenance present? false (both)

A schema honouring the ADR literally would fail both files #436 ships. One schemaVersion is being asked to cover two documents with different required fields — a per-provider normalizer output, and a multi-provider aggregate. I introduced that in e322cb0 while addressing your item 3, and it needs a decision before either schema can be written:

  • (a) Split the version — e.g. acq-neutral-model-catalog/v1 for a per-provider document, acq-neutral-model-aggregate/v1 for the orchestrator's output, each with its own required set; or
  • (b) Keep one version and make provenance required only on the aggregate, with the ADR amended to say so.

4. Half of this issue's scope is now orphaned

acq-provider-facts/v1 has zero producers — #436 removed the only one, and grep across integrations/ finds no remaining reference. That half should probably move to #456 (host-materialized provider facts), leaving this issue scoped to the catalog shape only. #424's scope item 1 also still specifies the removed heredoc verbatim.

Prior art relevant to question 3

Docker's Sandbox Kit Spec v3 (docker/sandbox-kit-spec) hit the same problem and solved it with a lever we could borrow: capability types carry their own version, independent of the document's schemaVersion. network-policy@1 and network-policy@2 coexist as published, addressable config schemas; the descriptor-wide version bump is explicitly "the last resort for what that lever cannot express." That is a cleaner answer to "two documents, different required fields" than overloading one version string.

They also close precisely the gate we lack: spec/schema_test.go pins the JSON Schema's constants to the Go validator so the two cannot drift silently, with the Go validator named as the validator of record. That is the shape of the fix for item 1 here — one implementation is authoritative, and a test asserts the other cannot diverge.

(Separately, and not an argument for adopting anything: their credential@1 keeps the real secret out of the guest entirely, which bears on #456 — I have left the detail there rather than here.)

What I would do, pending answers

Nothing yet. Writing either schema before questions 3 and 4 are settled means writing it twice, and picking (a) or (b) myself is the kind of call ADR-0004's own "What an agent must NOT decide unilaterally" section covers.

Once decided, the mechanical work is small and I will take it: schema files, a python-side jsonschema gate wired into the existing make validate path, and a drift test pinning the hand-rolled validator's hardcoded key sets to the schema so item 1 cannot recur.


AI-assisted (OpenCode), on operator instruction. Advisory; human review required.

wz-gsa commented on Oct 2, 2026

@wz-gsa
ContributorAuthor

Decisions on the four questions above, from my operator. Recording them here so whoever implements has the rationale, not just the outcome.

1. Two versions, not one — and ADR-0004 is amended

Split. acq-neutral-model-catalog/v1 for a per-provider document (a normalizer's output or a vendored snapshot); new acq-neutral-model-aggregate/v1 for the orchestrator's multi-provider aggregate.

The deciding argument was one I had not surfaced when I asked: a normalizer cannot know its own provenance. It is a pure file-to-file transform over bytes the bounded-fetch helper already retrieved, so source: live|snapshot is knowable only to its caller — a field the producer must be told and cannot verify. "Provenance required everywhere" would also have invalidated #436's shipped normalizer output and snapshot, reopening a green PR.

ADR-0004 is amended accordingly (patterns ea10fbf, CI 7/7), with the mistake written in rather than quietly corrected.

2. Provenance: aggregate-level, per-provider outcome

The aggregate records the set of providers discovered and each one's outcome — live / snapshot / rejected. A provider rejected by host-side validation must be visible as rejected, not merely absent. Entries name their origin providerId; staleness is answered once per provider rather than repeated on every row.

The per-provider document requires no provenance field.

3. Enforcement: python jsonschema, already in the repo

jsonschema==4.26.0 is already pinned in pyproject.toml and already used by validate_repo.py. Verified: Draft202012Validator passes catalog.json with 0 errors.

We considered adding ajv and retiring the hand-roll, and measured the cost first:

integrations/providers/usai/package.json              deps=0  lock=NO
integrations/isolation/acq-kits/usai-provider/...     deps=0  lock=NO
.github/linters/package.json                          deps=1  lock=yes

Both candidate homes have zero dependencies and no lockfile, so ajv would mean creating the repo's second npm dependency surface from scratch — new lockfile, new dependabot ecosystem, new npm audit gate — plus 4 transitive deps. There is also a correctness footgun: these schemas declare $schema: .../draft/2020-12/schema, and ajv 8 defaults to draft-07 (2020-12 needs the separate Ajv2020 export). The python path has neither problem.

The Node hand-roll stays for build-time feedback, with a drift test pinning its hardcoded nested key sets (gwAllowed, vendorAllowed, modelAllowed) to the schema — so item 1 of my earlier comment cannot recur.

4. Retrofit usai-model-catalog/v1 in the same PR

Verified safe before proposing a pattern, against the real 21 ids:

^[a-z0-9][a-z0-9_.-]*$   -> 0 of 21 ids fail    <- use this
^[a-z0-9][a-z0-9-]*$     -> 18 of 21 fail       (claude_4_5_sonnet, claude_4_7_opus, …)

So the retrofit lands with an underscore-permitting pattern plus maxItems on models[]. Had I guessed the stricter form it would have broken the shipped catalog — worth stating because "add a pattern" sounds safer than it is.

5. This issue: close and reopen scoped — but not yet

Agreed that the right end state is scoped issues, since much of the text above is now stale (the false "validated in CI" premise, the orphaned acq-provider-facts/v1 half, the #424 reference to a removed heredoc).

Deliberately sequenced last. Closing an issue a maintainer is engaged with, then reopening replacements, while the ADR the work depends on is still mid-review, is how a thread gets lost at the worst moment. Order:

  1. ADR-0004's version split gets reviewed (docs(isolation): add ADR 0004 for neutral model-provider discovery #423, ea10fbf).
  2. Then the schema PR: both schemas, the python gate wired into the existing make validate path, the retrofit, and the drift test.
  3. Then close this issue, referencing the merged work, and reopen anything genuinely still open.

Writing the schemas before step 1 means writing them against a version boundary that can still move.

Still owed to a maintainer, not decided by us

The acq-provider-facts/v1 half of this issue's scope belongs with #456 now that #436 removed its only producer. I have not moved it unilaterally — that is a scope call, and it is bundled into the close/reopen in step 3 rather than done piecemeal.


AI-assisted (OpenCode), on operator instruction. Advisory; human review required before any of this lands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions