Repository navigation
Add the integration seam design specs - #1084
jwrosewell wants to merge 46 commits into
Conversation
6364acc to
18b65ff
Compare
18b65ff to
616b740
Compare
The provider series design specs move to the spec-only PR (IABTechLab#1084) so they can be reviewed before the code that implements them. Three doc comments cited those files by repository path, which no longer resolves from this branch. Refer to each document by name instead, so the comment stays true whichever PR is read first.
3c7bc27 to
1800604
Compare
aram356
left a comment
There was a problem hiding this comment.
Summary
Thank you for this, and for §8 in particular. Publishing what implementing the seam actually found, including three items that say your own work is not finished, is the right instinct and it is what makes the design half of this trustworthy.
The design is accepted. §3.6's rule is correct, and it is the rule the work below implements. Things the host supplies are platform services. Things a vendor supplies are capabilities of that vendor's module. What is being changed is the order and the structure, not the direction.
Five changes to sequencing and structure follow. Tell me where the reasoning is wrong and I will change my view, but I want these settled before more code lands on the current shape. The two items that need wider sign-off are named at the end.
1. One crates/trusted-server-integrations crate first, with a discovering build.rs
Move all thirteen integrations into a single crate and generate the registration table at build time by scanning the crate's own directories, in the manner crates/trusted-server-js/build.rs already scans dist/tsjs-*.js.
This is the change that delivers the neutrality goal. Core stops naming vendors in builders(), in validate_enabled_integrations, and in migration_guards.rs, and it stops naming them because nobody writes the list, not because the list moved. Three of §1's four hardcoded vendor tables become generated output.
It also resolves several of §8's findings without further design:
- §8.2, the hand-carried module hash and its line-ending fragility, goes away because hashes are generated exactly as core's are today.
- §8.6, a module's
validaterunning nowhere, is solvable in the generated table. - §1.2, the compile-time JS map, is solved by one scan that finds Rust and TypeScript together.
- The
migration_guards.rsdrift is structurally impossible once generated. Worth noting that this drift is not hypothetical:osanois registered inbuilders()and has no entry inmigration_guards.rsat all onmain. Nobody noticed because both lists are hand-written.
prebid stays in core as protocol support, as §3.4 already says.
The runtime injection seam in §3.1 is not required for any of the thirteen integrations you have, because they are all compiled from source in this repository. It becomes the right mechanism when a genuinely external vendor crate exists, and it should be added then, with that vendor as its first consumer.
2. Crate layout follows from that
§4 proposes nine vendor crates, one PR each. We are not doing that now. One crate holds all thirteen integrations, and a vendor gets its own crate when that vendor takes over maintaining it, which none of them does today. Every integration on main imports the same roughly twelve crates, all of which core already depends on, so nothing about dependencies argues for splitting them either. Splitting later is a directory move against a module that already has clean boundaries.
Two consequences for what is in #1094 now:
- Directory name equals package name, as everywhere else in the repository. The crate is
crates/trusted-server-integrations.crates/geo/fastlyas packagetrusted-server-geo-fastlybreaks that convention, andcrates/geo/is a category folder holding one item. - Host code stays in its adapter.
FastlyPlatformGeois 52 lines that already live atcrates/trusted-server-adapter-fastly/src/platform.rs:687, and the Fastly adapter is its only possible consumer. Extracting it into a crate gains nothing and contradicts §3.6's own rule that host-supplied things are platform services. The same applies to the 103-line Fastly device provider.crates/edgecookie/contains only a README and should be deleted until a vendor crate lands.seam-probeis a test fixture and should not be a permanent workspace member.
3. Permissions policy belongs in trusted-server.toml; permissions.yaml should not exist
Policy belongs in a [permissions] section of trusted-server.toml, beside the existing [consent]. We should not add a second configuration format, and policy should not be compiled into the binary.
The permission-model spec's §3.1 gives the reason for YAML. The draft required policy to flow through the runtime config pipeline, and that activation apparatus does not exist. That conflates two things. ts config push does exist. It publishes trusted-server.toml as a blob envelope, and TrustedServerAppConfig already carries the entire Settings struct. What does not exist is the staged multi-revision activation protocol the draft designed, with fleet quiescence, admission leases, and a hash-linked journal. No other setting needs that machinery. The publisher domain, the proxy secret, the EC passphrase, and every integration config all push atomically today, and policy is not different in kind.
What the move gains:
- One place to configure everything. Today policy is the only control that bypasses
ts config validateandts config push, and it is the most compliance-sensitive one. - The validation we already have, being
deny_unknown_fields,Validate, and startup errors, instead of a bespoke parser reimplementing duplicate-key and case-collision checks. - A publisher can correct a jurisdiction rule with a config push rather than rebuilding and redeploying a WebAssembly binary. For a compliance control that may have to change on a regulator's timeline, that difference is the point.
- §3.1's own recorded defect, that the
include_str!path reaches above the crate root, disappears.
Alongside the move, please trim the fifty-three Data Uses that have no signal mapping and no enforcement point. Only two carry enforcement weight today. Moving fifty-three inert flags into trusted-server.toml would carry the misleading operator view across rather than fix it, and the spec already acknowledges this departs from its own rule that a permission appears only with both a mapping and an enforcement point.
4. The EC provider spec rebases onto 1 and 3
Rewritten against a seam that already exists and policy that lives in settings, the EC provider work is materially smaller than #1043 as it stands:
- No
RuntimeServices::ec_providerslot to add and then delete, and noresolved_ec_providersecond field. Note that #1094 currently has both, plus the registration path, so identity has three injection paths where §3.6 specifies one. §3.6's sentence that theec_providerslot "goes" is not yet honored. - No per-adapter resolution divergence. Related, and worth fixing wherever this lands: on #1094
resolve_geo_providerandresolve_device_providerreturnResultand fail loudly on a bad selector, whileresolve_ec_providerreturns a bareOptionand falls back silently. §3.6 says an unresolvable selector is a startup error. - No double configuration migration. This is the strongest reason for the reordering. As it stands, #1043 migrates operators from
[ec] passphraseto[ec.providers.hmac], and §3.6 with sign-off row 8 then migrates them again to[integrations.<id>]. Two breaking configuration changes in one release cycle, the second withdrawing the first. Reordered, operators migrate once, into the final shape.
5. Some of this belongs in edgezero, not Trusted Server
edgezero-core already owns the platform layer, being app_config, blob_envelope, config_store, secret_store, key_value_store, store_registry, the router, the middleware, and the four adapters. Hooks::stores() returning StoresMetadata is already the seam run_app uses to inject store registries. Before we build more platform machinery in this repository, we should be explicit about which side of that line each piece sits on.
- The Fastly composition gap (§8 item 7) is probably an edgezero fix rather than a Trusted Server one.
trusted-server-adapter-fastlywrapsedgezero-adapter-fastly. If a deployment composes its modules throughHooks, the pattern already exists upstream and no Trusted Server adapter needs a new library target. Adding a[lib]here would work around the absence of something that may already exist one level down. RuntimeServicesoverlaps edgezero substantially. It carries a config store, a secret store, a KV store, an HTTP client, and a backend.edgezero-coremodels all five. Adding identity, device, and geo provider slots to it builds up a second platform abstraction beside edgezero's rather than using edgezero's.PlatformGeohas no edgezero counterpart. There is no geo abstraction inedgezero-coretoday. Either a geo lookup is a host capability edgezero should expose the way it exposes a KV store, or it stays here. Whichever it is, that should be a stated call rather than a default that follows from where the code happens to have been written.TrustedServerAppimplementsHooks::name()andHooks::routes()but notHooks::stores()(crates/trusted-server-adapter-cloudflare/src/app.rs:332). The store-injection seam edgezero already provides is unused here, which is worth resolving before we design another one.
Which of these should move is a question for whoever owns edgezero. I would rather ask it now than build a second mechanism here and answer it afterwards. It also makes change 1 smaller: if composition runs through Hooks, the integrations crate needs less new machinery than either this spec or #1094 assumes.
Spec corrections, verified against the tree
- Five of the seven specs are statused "Implemented" for code that is not on
main.EdgeCookieProvider,PermissionSet, andresolve_from_clientare all absent fromorigin/main. Merging this set as written would tell every later reader that shipped behavior exists when it does not. Please status everything unmerged as "Proposed". - §4, "thirteen vendor files" in
migration_guards.rs, is wrong. There are eighteen vendor files across twentyintegrations/entries;nextjsalone contributes six. Thirteen is the count of builders. And as above,osanohas no entry at all, which §4's migration plan depends on. - §3.5, GPT diagnostics "called by name from all four adapters", undercounts. There are nine
prepare_requestcall sites plus onefinalize_response, so the hook work in §3.5 is larger than stated. - §8.3, "a provider is resolved more than once per request", does not hold on
main. No integration or proxy calls geo. A real double resolution does exist, but it is core against core:handle_auctionatauction/endpoints.rs:262and the adapter'sbuild_ec_contextboth calllookuponPOST /auction. The underlying gap, that there is no per-request provider context, is real and worth keeping. - §8.5 understates the Spin gap. Spin does not merely skip inline admin stubs; it skips all six first-party route bindings (proxy, click, sign, proxy-rebuild), which carry real traffic. Cloudflare, by contrast, has no health route and so covers everything.
- Line references to correct:
settings.rs:166should be:215;auction/mod.rs:49points at the function header rather than the list;publisher.rs:4361should be:4369.
§8 item 7 is the load-bearing finding
Your assessment is right and the cause is more fundamental than the text states. The Fastly adapter has no [lib] section and no src/lib.rs, and every module is declared bare mod, so nothing in the crate is externally reachable. This is not a visibility modifier to widen; it is a missing library target. Axum, Cloudflare, and Spin each expose pub fn routes_with_settings and pub struct TrustedServerApp; Fastly exposes neither.
As things stand the seam is proven on the dev server and unreachable on the primary deployment target. §6's acceptance criteria should require the round trip on Fastly specifically, not on any adapter. Under change 1 this stops blocking the immediate work, since in-tree integrations compile in and nothing needs runtime injection. It must still be fixed before the first external vendor crate, and per change 5 the fix may belong in edgezero rather than here.
Disposition of the open pull requests
#1043 through #1047 and #1094 are superseded rather than rejected. The design in §3.6 is accepted, and most of the provider logic carries over onto the reordered base. Please do not read this as the work being discarded.
The order, with the first two independent and able to run in parallel:
- One
crates/trusted-server-integrationscrate with a discoveringbuild.rs, againstmain. - Permissions into
trusted-server.toml,permissions.yamldeleted. - The EC provider work, rebased onto 1 and 2.
- Device and geo selection, in the same shape.
- Per-vendor crate extraction, if and when a vendor takes over maintaining one.
Happy to discuss who does what. The Fastly composition gap is core infrastructure rather than vendor contribution, and per change 5 it may not belong in this repository at all.
The jurisdiction default should not name a country
[geo] default_country is required in every mode, and it must resolve to a rule in the policy file. There are three groups today, being gdpr-eu, gdpr-uk, and us-opt-out, so a deployment with global traffic and no geo provider has to nominate one country and then apply that country's law to every visitor. That is the wrong shape. It makes a neutral product take a position on whose rules apply, and the operator's honest answer is usually "I do not know where this visitor is", not "treat everyone as French".
The default should be a policy baseline rather than a country. Two forms worth considering, and I do not mind which:
- A named group, so an operator selects
gdpr-euor a strictest-available baseline directly without claiming a location. - A per-permission baseline written inline, for an operator who wants to state exactly what is allowed with no signal.
Either way the value says what is permitted, not where the visitor supposedly is. The requires-signal floor already exists for a failed lookup, so the machinery for a location-free baseline is there.
This also removes most of what I was going to send to the task force. The UK row shipping necessary.operations.storage: granted while the EU requires a signal, and Australia mapping to us-opt-out, are still policy positions that need someone with the relevant expertise to confirm with a citation, and the shipped table should carry that citation. But nobody has to ratify a default that no longer asserts a jurisdiction.
The crate layout, the configuration format, and the merge order are not task force questions. They are engineering calls and we should settle them here.
One further item to fix wherever the permission work lands: §7 of the permission-model spec records that publisher navigation and page-bids attach user.id under the provider gate rather than the sharing pair, so a storage-only grant can still place the EC identifier in user.id on those paths while /auction pair-gates it. That is an open inconsistency in the sharing boundary, and it should be closed rather than tracked.
|
Thank you for this. It is a serious review and it moves the design forward. Taking your points in order. 1. One integrations crate with a discovering 2. Crate layout. Here we differ. Whether a vendor wrote the code, or the Tech Lab team wrote it for them, should not decide whether it gets its own crate. Separating per vendor now makes ownership legible, and a crate boundary is enforced by the compiler, while a CODEOWNERS entry only assigns review ownership and constrains nothing about what code can reach. The maintainer can be the Tech Lab team today. If the Tech Lab team built the Permutive integration, the Permutive crate is theirs until Permutive take it over, and the handover then changes one line, not the structure. We take the point that host code stays in its adapter, which §3.6 already says, and we will follow the repository's crate-naming convention. The disagreement is only on timing. 3. Permissions. This one is fundamental for us and not something we can change 🙂. Deleting the dedicated permissions file breaks a deployment model that is very hard to change later. The transparency point first. A single, dedicated permissions policy that anyone can open and read is how Trusted Server shows exactly what a publisher permits, on its own rather than buried beside proxy secrets and backend hosts. Then the mechanism. Trusted Server should not decide policy for the deployer, and there should be no automatic default. Selecting the policy is a conscious choice when an environment is set up for the first time. The entity that compiles a deployment provides the policy it wants as part of the build, which is why the policy has to be a build-time artifact and not only a runtime config section, and that same entity can expose configuration override so the operator of the compiled variant can adjust it without rebuilding. Consider a provider in Europe that compiles Trusted Server with a European base and ships it to its customers, who get an interface to tune it. Two actors and two moments, with the integrator setting the base at build time and the operator adjusting it at runtime. A single runtime config section cannot express that, and neither can a compiled-only policy, so both layers are required. The repository should also carry multiple policy files for different test scenarios, because exercising the permission model across scenarios is part of the robustness of the service. Three things follow. Policy is not locked in the binary forever. The compiled base is the layer these PRs implement, and the operator override is the follow-on the permission-model spec already records as deferred, so a jurisdiction rule can change on a regulator's timeline without a rebuild. What the shipped jurisdiction rules should say, the UK storage row and the country baselines included, is a policy question for the group to confirm with citations, which is your point in the jurisdiction section too, and we would take it there rather than settle it in code review. And which file carries the compiled-in base and which one arrives through configuration at runtime is not something engineers need to debate. That choice sits with the people deploying, and testing proves every mutation of base and override works. 4. The EC provider and one injection path. Agreed. §3.6 specifies one identity injection path and we will honor it, collapsing to the single registration path. An unresolvable selector will fail loudly, while an absent selector keeps today's behavior of running with no Edge Cookie provider. Rebased onto a seam that already exists, the double configuration migration also goes, so operators migrate once. 5. What belongs where. The distinction that matters is which repository these live in. The geo provider, the device provider and the Edge Cookie provider are part of Trusted Server, not EdgeZero. EdgeZero supplies the host plumbing, and the provider capabilities are Trusted Server's domain, which is where a vendor goes to plug in. They are also deliberately off by default. Only User-Agent device detection runs by default, geo resolves no location until a provider is configured, and the Edge Cookie provider is selected by configuration, which is the same conscious-choice principle as point 3. Vendor code already lives in Trusted Server as the registered integrations, and the provider seam exists so a vendor can be added the same way. We are asking to use it for the purpose it was built. Where you are right is the platform plumbing rather than the providers. The config, secret and KV stores that Spec corrections. Taken, and thank you for checking against the tree.
On merging. We agree with most of this and will make the changes. The differences are narrow, being per-vendor separation now rather than later and keeping a dedicated permissions file. Rather than treat the current PRs as superseded and reset, could we agree a concrete sequence and owners for landing this so the provider work merges soon? We would like to keep this a direct back and forth to a merge. Thanks again for the depth of the review. |
ChristianPavilonis
left a comment
There was a problem hiding this comment.
Summary
Requesting changes for three security or authorization failures and three contract gaps in the normative design. These findings are distinct from the existing feedback on sequencing, implementation status, Fastly composition, policy placement, and jurisdiction defaults.
Reviewed 180060402f91b4efa96f5823da8934d19c1c78f8 against d516a9e94249e10cbc36e41beb4269f9255cf407.
| APS can register from its own crate. Prebid stays in core as protocol | ||
| support rather than as a vendor integration. | ||
|
|
||
| The bid renderer contract is generalized in the same change. Today |
There was a problem hiding this comment.
🔧 P2: Add a browser renderer registration seam before moving APS
Generalizing Rust's BidRenderer is not enough to move APS out whole. TSJS core still imports parseApsRendererDescriptor from APS in core/auction.ts, imports APS dispatch in core/request.ts, and fixes AuctionBidRenderer to ApsRendererV1 in core/types.ts. Because build-all.mjs builds core/index.ts as a self-contained IIFE, those static imports remain in tsjs-core.js independently of carried integration modules. Please specify a browser renderer registry keyed by descriptor type. Integration modules should register their parser and dispatcher, core should reject unknown types safely, and a cross-language test should prove APS can be removed without changing or breaking the core bundle.
There was a problem hiding this comment.
You are right, and it is out of scope for this stack. Leaving this thread open deliberately.
Confirming your reading. core/auction.ts imports parseApsRendererDescriptor from the APS integration, core/request.ts imports the APS dispatch, and core/types.ts fixes AuctionBidRenderer to ApsRendererV1.
Two things to add, and both help your case. There is no integrations/aps/index.ts at all, so APS is not a discovered browser module in the first place. And that code is pre-existing on main and byte identical on this branch, so no PR in this stack introduces or changes it.
What I have done. Corrected the §4 APS migration row so it names the browser-side work as well as the Rust side, and added the coupling as a §8 finding. Designing the browser renderer registry belongs with the APS migration rather than here, so I have filed it as #1111 to make sure it is not lost in a review thread.
I am not resolving this one, because the deferral is a judgment you should get to disagree with. If you would rather the registry were specified before this spec merges, say so and I will take it back into scope.
Full response to all six points is in this comment.
Rewrite the resolve endpoint's origin check as an exact serialized-origin
comparison, defaulting to the single origin https://{publisher.domain}
with an optional operator-configured list, and delete the sibling
subdomain justification that did not hold. Record that the origin check is
defense in depth under the provider's envelope verification, require the
sibling subdomain, wrong scheme and non-default port rejection cases, and
close open question 7.5.
Require core to expire the ts-ecr marker on any request carrying a ts-ec
the selected provider does not own, in both the client-cycle and pluggable
provider specs, so a provider switch restarts the client cycle instead of
leaving a visitor with a marker and no identity. Record the constraints on
carrying resolved permissions to the browser without designing the
mechanism, which belongs with the first vendor module.
Require a nonempty DeviceProvider::required_permissions from a
module-supplied provider to be refused at registry resolution rather than
described as a documented no-op.
Replace the DataDome singleton-header rule that skipped the vendor call on
a repeated field, which would have created an attacker-controlled bot
detection bypass and described no shipped behavior. Core now passes
repeated field lines through as received and leaves interpretation to the
vendor integration, keeping only the protocol-level content-length rule.
Correct the APS migration row and add a seam spec finding for the
pre-existing browser-side APS coupling in core TypeScript.
DeviceProvider::required_permissions is public but has no enforcement point. Device classification runs before the permission set for a request is assembled, so there is no per-request gate to check a declaration against, and selecting a provider is an operator decision rather than a per-request permission decision. Before the module seam that was a documented no-op, because core and the host supplied the only two providers and neither declares anything. With the seam a vendor crate can implement the trait, so the method reads as a promise that is not kept. Refuse the selection at registry resolution, loudly and at startup, rather than honoring the declaration by ignoring it. The refusal lifts when a real per-request device gate exists. Raised by Christian Pavilonis in review of IABTechLab#1084.
|
Thanks. Great review and I've checked the points. All six hold up. Five are changing in this stack, and one is real but is new design work that these PRs are not the place for. In order. P1, subdomain identity fixation. Agreed, and it is worse than you say. The check in The fix is the same-origin test of RFC 6454, so the request One more point on severity. The origin check is defense in depth here, not the primary control. The primary control is the provider's envelope verification with audience and session binding, which §3.9 parks out of v1. An exact allowlist does not help against a subdomain that legitimately serves publisher pages and is later compromised, so §3.9 still has to land with the vendor scheme. P1, device provider permissions. Agreed. I've taken your second option. A nonempty P1, duplicate headers and DataDome. Agreed on the risk, and I've changed the approach rather than patched the rule. One correction on the framing first. Nothing skips the call today because a header arrived twice. I haven't taken either remedy exactly, because both still have core deciding what to do with a vendor's evidence. Core should not invent header normalization for a vendor, and it certainly should not decide to skip a vendor call. Core passes the request headers through as received, repeated field lines preserved, and the vendor integration decides how to interpret multiplicity, because that is the vendor's detection logic and not core's business. The one field I am leaving out of that is The wider point, and it is the same separation argument running through this whole review. The DataDome logic sits in core today. Vendor detection logic of that kind belongs in the vendor's own module, behind the seam, not in the shared engine. Getting it there is the direction, and this rule is a small example of why. P2, the I've taken your second remedy rather than your first. Core expires the marker on any request carrying a P2, browser permission gating. Agreed, the mechanism was undefined. It is defined and delivered now, in the permissions PR (#1045). There are two deliveries. On the server, integration request filters receive the permission state resolved for the request, next to the geo result they already get, on the Fastly adapter, which is the only adapter that runs the filter step today. On the page, that same state arrives as Because the arrival point moves, core exposes A page module is treated the same way as any other provider. It declares the permissions it requires and checks them against the state it is handed before it does anything. The client-fixed demo script does exactly that now, declaring P2, the browser renderer seam. You are right, and it is out of scope for this stack. I've corrected the §4 APS migration row so it names the browser-side work as well as the Rust side, and added the coupling as a §8 finding. Designing the browser renderer registry belongs with the APS migration, not here, and I have filed it as #1111 so it is not lost in this thread. On scope. This stack is not adding new designs or features. It implements the seam that was agreed, and where it overclaims I've corrected the text rather than widened the specs to cover work these PRs will not deliver. The browser renderer registry is genuinely new design work. It is a good agenda item for the next meeting and I'd rather settle it there than in a review thread. Five changes are in the stack now, being the exact origin check with its rejection tests, the device permission rejection, the marker expiry, the DataDome pass-through rule, and the page and filter carriage of the resolved permissions. The other one is a spec correction that says plainly what is not yet designed. |
DeviceProvider::required_permissions is public but has no enforcement point. Device classification runs before the permission set for a request is assembled, so there is no per-request gate to check a declaration against, and selecting a provider is an operator decision rather than a per-request permission decision. Before the module seam that was a documented no-op, because core and the host supplied the only two providers and neither declares anything. With the seam a vendor crate can implement the trait, so the method reads as a promise that is not kept. Refuse the selection at registry resolution, loudly and at startup, rather than honoring the declaration by ignoring it. The refusal lifts when a real per-request device gate exists. Raised by Christian Pavilonis in review of IABTechLab#1084.
aram356
left a comment
There was a problem hiding this comment.
Permissions policy should use the existing typed app-config path
Requesting changes on the permissions.yaml decision.
The argument for a dedicated YAML file does not establish a technical guarantee. YAML is editable before a build, and include_str! only records a snapshot in that binary. It does not prove who selected the policy, that the repository file matches the deployed binary, or that a later build or deployment preserves it. If an immutable distributor floor is required, that needs an explicit trust model—such as signed build provenance, an exposed policy digest, and a merge rule under which runtime configuration can only tighten the floor. The file format itself does not provide that.
The current design also does not yet deliver the two-layer model described in the discussion: permissions.yaml is compiled in, while runtime overrides are deferred. As written, operational policy changes require rebuilding and redeploying, and the most compliance-sensitive configuration bypasses ts config validate, ts config diff, and ts config push.
We already have an appropriate configuration path:
- Add
permissions: PermissionsConfigtoSettings, besideconsent. - Serialize and deserialize it through
TrustedServerAppConfig. - Validate group references, permission identifiers, acquisition values, normalized location keys, and provider enforcement coverage in
validate_settings_for_deploy. - Publish it in the existing
BlobEnvelopewithts config push. - Construct
PermissionMapsfromsettings.permissionsonce at startup.
The TOML can represent the same concepts: a location-independent default_group, named groups or baselines, country/region-to-group rules, and consent-signal mappings. A default_group also avoids pretending that a visitor with unknown geography belongs to a made-up default_country.
Please revise the spec so the effective operational policy lives under [permissions] in trusted-server.toml and participates in the existing typed configuration lifecycle. If a compiled distributor floor is still required, specify it as a separate optional layer and define exactly:
- Who is trusted to choose it.
- How a deployed binary proves which floor it contains.
- Whether runtime configuration may only tighten it—for example,
granted < requires_signal < denied. - How the effective merged policy and its digest are exposed for audit.
Without those guarantees, a separate compiled YAML file adds a second configuration system and a rebuild requirement without delivering the immutability or auditability used to justify it.
There was a problem hiding this comment.
Vendor-owned crates should be an extension path, not the default migration
This is blocking because §4 makes separate vendor crates the required structure for integrations that are currently maintained together in this repository. I agree with the goal of allowing vendors to own their integrations independently. I do not agree that we should create those ownership boundaries before the owners and independent lifecycles actually exist.
A crate boundary should represent a real maintenance boundary: an independent owner, release process, dependency or licensing requirements, security-response responsibility, and compatibility commitments. The current integrations are reviewed under the same project governance, released with Trusted Server, and use largely the same dependencies. Moving each one into its own crate does not make it vendor-maintained. Unless a vendor has explicitly accepted those responsibilities, the Trusted Server team still owns the integration—only now it also owns another manifest, release surface, CI path, dependency-update stream, and compatibility obligation.
Separate Rust crates provide compile-time package boundaries. They do not by themselves provide runtime isolation, independent deployment, or the ability to add an integration without rebuilding Trusted Server. If all of these crates remain workspace dependencies compiled into the same binary, most deployment characteristics remain unchanged while repository and release complexity increase.
The current integrations should instead live as vendor modules in one crate:
crates/trusted-server-integrations/
src/
vendor_a/
vendor_b/
vendor_c/
That crate can still provide the intended provider interface, generated registration, and shared conformance tests. The interface should also allow an independently maintained external crate to implement the same contract. A vendor contribution can therefore start as a module in the shared crate or as a genuinely external implementation, depending on who owns its lifecycle.
An integration should be extracted into its own crate when at least one concrete boundary exists, for example:
- The vendor commits to owning releases, compatibility testing, and security fixes.
- The integration has materially different dependency or licensing requirements.
- It needs an independent release cadence.
- A real external integration demonstrates that the provider seam requires the separation.
Whether Trusted Server supports independently maintained vendor integrations is a product and governance decision. Whether every existing integration needs a separate crate is an engineering packaging decision, and the current ownership and release model does not justify it.
Please revise the spec to use one trusted-server-integrations crate for the integrations maintained here, with vendor-specific crates supported as an optional extraction path when real owners and lifecycle boundaries emerge. This preserves the vendor-owned extension model without creating a separate package for every hypothetical future owner.
aram356
left a comment
There was a problem hiding this comment.
Blocking: define and sequence the EdgeZero boundary before adding another platform layer
This is blocking. §3.6 states the right high-level rule—host-supplied capabilities are platform services—but the design does not carry that rule through to repository ownership or implementation sequence. Trusted Server already runs on EdgeZero. If we add the missing host-service and composition machinery here first, we create a second platform abstraction that every adapter and future integration must understand.
The overlap is concrete on current main:
edgezero-coreownsHooks,RequestContext, app configuration,BlobEnvelope, config/secret/KV store registries, proxying, routing, and middleware.Hooks::stores()is already the portable declaration that EdgeZero adapter runners use to provision and inject stores.RequestContextalready exposes named and default config, secret, and KV handles.- Each EdgeZero adapter already owns extraction of its native request context.
- Trusted Server's
RuntimeServicesseparately carries config store, secret store, KV store, HTTP client, backend, geo lookup, and client information.
That leaves the boundary ambiguous: some platform services flow through EdgeZero, while others are rebuilt and selected in Trusted Server. The provider seam should not make that duplication permanent.
What belongs in EdgeZero
-
Portable host-service access. Config, secret, and KV handles must continue to flow through EdgeZero's store registries and
RequestContext. Generic outbound HTTP/proxy access and platform transport/backend capabilities should also have one EdgeZero-owned contract. Trusted Server may use thin type adapters where its error or data types differ, but it should not own a second service registry or a second adapter-selection mechanism for the same host resources. -
Normalized host evidence. Client IP, TLS protocol/cipher, JA4/HTTP2 evidence, and host-provided geographic data are facts extracted by the deployment adapter. The portable evidence types and capability-presence metadata belong in
edgezero-core; the Fastly, Cloudflare, Spin, and Axum implementations belong in their EdgeZero adapters. Trusted Server should consume normalized evidence rather than reaching into platform SDKs or inventing a separate cross-adapter host context. -
Generic adapter lifecycle/composition hooks. If loading an application-supplied module requires a new startup, request-extension, or composition hook, the generic hook belongs in
edgezero_core::app::Hooksand must be honored by all four EdgeZero adapter runners. Fastly cannot have a separate composition rule. The payload carried through that hook can be Trusted Server-specific, but the adapter lifecycle mechanism cannot be. -
Host capability discovery. EdgeZero should expose which optional host capabilities are available so an application can fail startup when a selected provider requires unavailable evidence or transport. The provider requirement remains a Trusted Server concept; the authoritative statement of what the current adapter supplies belongs to EdgeZero.
What remains in Trusted Server
IntegrationRegistration, vendor JavaScript, prepare/finalize hooks, auction providers, configuration selection, and conformance tests.EdgeCookieProvider,GeoProvider, andDeviceProviderbehavior, including vendor calls, selection, per-request sharing, and permission enforcement.- HMAC identity, the User-Agent classifier, and other Tech Lab-owned implementations.
- Permission and jurisdiction policy.
- Auction telemetry and other ad-tech-specific request/response behavior.
The important distinction is that a host geo lookup belongs in EdgeZero, while a vendor geo provider that interprets or augments geo data belongs behind the Trusted Server integration seam. Likewise, raw TLS/JA4/H2 evidence belongs in EdgeZero, while device classification based on that evidence remains Trusted Server behavior. Identity remains a Trusted Server provider, while the KV and secret-store plumbing it consumes remains EdgeZero infrastructure.
Required changes to this spec
Please make the dependency direction and sequencing normative:
- Define the missing portable host-evidence, transport, capability-discovery, and lifecycle/composition contracts in EdgeZero.
- Land and version those EdgeZero changes first.
- Upgrade Trusted Server to that EdgeZero version and consume the contracts through
HooksandRequestContext. - Make
TrustedServerAppdeclare its stores throughHooks::stores()instead of leaving the existing EdgeZero injection seam unused. - Keep
RuntimeServices, if it remains, as a Trusted Server request/provider context over EdgeZero handles and normalized evidence—not as an independently owned platform layer. - Require the external-registration round trip on Fastly, Axum, Cloudflare, and Spin through the same lifecycle contract. Any Trusted Server-specific library entry point can remain here, but it must not compensate for a missing generic EdgeZero adapter hook.
Until those prerequisites and ownership lines are recorded, §3.6 permits two competing platform seams and leaves each implementation PR to decide ad hoc what belongs upstream. That is exactly the architectural drift this design is supposed to prevent.
PR IABTechLab#1016 replaced the auction provider table that section 3.4 opened with a compiled auction plan, and the seam implementation follows that plan rather than adding an auction provider builder. Only the generalized bid renderer survives from section 3.4. - Section 3.4 says the change does not open the auction, keeps the renderer contract, and records what the plan keeps closed, being the fixed profile registry, the profile special cases in the plan compiler, and the single mediator the plan accepts. - Section 1 item 4, section 3.1, section 3.3, the APS row in section 4, section 8 item 4 and section 9 row 6 no longer claim an auction provider seam, and section 3.1 records that the Prebid and APS ids are reserved by the duplicate id check.
Every pluggable thing is a provider, each provider type is one top-level table, and a `provider` key inside that table selects what runs. The series specs still described the shapes that convention replaces, so a reader building against them would have written configuration the stack rejects. The providers spec now states the convention once, in a new section 2.1, and the other specs follow it. The seven types are `ec`, `geo`, `device`, `permission_signal`, `demand`, `adserver` and `integration`, all singular and snake_case. `provider` takes a string where one runs and a list where several run. A `[<type>.<name>]` table exists only where that provider has something to set, the word `providers` appears nowhere in configuration, and the name is the implementation unless the table carries `implementation`, which is how two Prebid Servers run side by side under different names. A table its type's `provider` does not select refuses startup, an implementation the build does not carry refuses startup with the list of what it does carry, a demand or ad server endpoint must be HTTPS or HTTP to a loopback host, every provider rejects settings it does not know, and a secret setting names a key in `trusted_server_secrets` rather than holding a value. What changed, by shape: - `[ec.providers.<key>]` becomes `[ec.<name>]`, and the hyphenated selection names become `host_signals`, `client_fixed`, `gpp_sale_opt_out` and `us_privacy`. The `client-fixed-demo` cargo feature keeps its name, because a cargo feature is not a provider name. - `[permission_signal] sources` becomes `[permission_signal] provider`, the list form of the same key every other type uses. - `[integrations.<id>]` becomes `[integration.<id>]`, and which vendors run is `[integration] provider`. - `[auction.providers]` and the word mediator are gone. Demand and the ad server are provider types in their own right, so the seam spec section 3.4 now shows `[demand]`, `[demand.<name>]`, `[adserver]` and the `[auction.bidders.<code>] provider` mapping that stays. APS stopped being an integration, so it is configured under `[demand.<name>]`, carries its `rendering_mode` there, and can never be named in `[integration] provider`. - `[auction]` is not a provider type and keeps `enabled`, `timeout_ms`, the creative settings and `allowed_context_keys`. All three specs that describe the seam now say the same thing about it. Core names no vendor, an implementation is registered by an integration builder, and a builder that supplies an implementation does not have to be a page integration. That is how the HMAC identity provider, the User-Agent-only device classifier and the four permission signal schemes reach configuration without any of them shipping browser JavaScript or being named in `[integration] provider`. Left as they are. The permission model's record of the draft's unadopted `regime` class still names `us-privacy`, because that is a class name from a proposal that was rejected rather than a configuration key. The `signals.us_opt_out.sources` list is policy inside the compiled `permissions.yaml`, not a provider selector, so it keeps its name. `provider-code-registry.md` needed no change, because it allocates four-character identifier codes and describes no configuration. Verified. `prettier --check` under the docs configuration, the same check the `format-docs` job runs, passes on all seven files. Every markdown link, document reference and section cross-reference in the changed files resolves, apart from one that was already broken, being `docs/guide/permission-model.md` in the permission model spec, which exists further down the stack but not on this branch, because this branch targets main directly.
Two rows of the decisions table joined their clauses with a colon and a semicolon. They now read as one sentence each, and the table is realigned to the widened cell.
This branch adds specification documents only, so main's eight new commits merge into it without a decision.
What an operator selects to supply a capability is a module, and provider keeps the one meaning it has on main, an auction provider instance. The built-in set is discovered at build time from the integration directories, so builders() is generated and core keeps no hand-written vendor list, which is the better answer the 31 August review gave. The maintainers convention gains its two practical reasons, that a vendor crate pins its own dependency versions and releases without a pull request here. A new section, 3.7, makes a page change one middleware, run only where an ordered [[fetch]] or [[serve]] entry names it, in two phases. Fetch output is what a shared template stores and is handed nothing about the reader. Serve runs per reader on that reader's copy and is never stored. The four page hook traits and their contexts go when their nineteen implementations in twelve integrations have converted, one integration per change against the differential harness. The acceptance criteria and the sign-off table follow, and the hook traits leave "What does not change". Prettier and the docs lint pass. The docs build leaves this folder out.
What an operator selects to supply a capability is a module, and provider keeps the one meaning it has on main, an auction provider instance. The seven specs in this pull request now say so throughout, matching the code in the Edge Cookie seam pull request, which selects with [ec] module. - The pluggable spec's convention (§2.1) states which key each type reads: module for ec, geo, device, permission_signal and integration, and provider for demand and adserver, with a selector column in the type table. - The selectors read [ec] module, [geo] module, [device] module, [permission_signal] module and [integration] module in every example and rule. [demand] provider, [adserver] provider and [auction.bidders.<code>] provider are unchanged, and the seam spec's §3.4 is untouched. - The plain word says module wherever it meant an Edge Cookie, geo, device or permission signal module, and the identifiers follow the code (EdgeCookieModule, ec/module.rs, build_module, module_owns_id). - The registry says module and keeps its file name, so links to it keep working, and it no longer claims a retirement condition on module_owns_id, which the code no longer carries. - The seam spec's title is "The Integration Seam", since "provider" in its old title would now read as the auction meaning. - Each spec records the change in its status note and its last-updated line, and the seam spec's revision record carries it. Checked with the docs lint and Prettier.
…ations by type A module's name is its crate path below crates/, with . between the parts, taken from CARGO_MANIFEST_DIR at build, so it cannot drift from its folder as the hand-written ids had. Sections are named as their type folders, [ec] apart, a written name may leave its section's type folder off, core's own modules take bare names, and a section selecting several modules uses modules. Page integrations are selected from the section of their type, so there is no [integration] section.
`[demand] modules`, `[ad-server] module` and `[auction.bidders.<code>] module` replace the provider key in the seam design and the pluggable modules design, the permission model design selects with `[permission-signal] modules` and names its modules `gpc`, `gpp`, `us-privacy` and `tcf`, and the rule that every name is snake_case gives way to the module name rule, under which a module from a crate is named by its folder below `crates/` and only a `[demand]` or `[ad-server]` label stays snake_case.
…tion] A builder that supplies only an implementation is selected by no section, Prebid Server, APS and the ad server mock are selected by [demand] and [ad-server] and by no section, the registry exposes every registration it was built from and not only the ones the sections select, a vendor's table is [<type>.<name>] and the sections are read into TypeSections, and the principle that every type selects with module or modules replaces the sentence that kept provider for the auction's tables.
The seam design's auction section still named the demand and ad server implementations `openrtb`, `prebid_server`, `aps` and `adserver_mock`, and said a table's name is its implementation unless it carries an `implementation` line. The implementation names are module paths, which are `auction-protocol.openrtb`, `auction.prebid-server`, `auction.aps` and `ad-server.mock`. A demand table always names its implementation on an `implementation` line, because `demand` is not the type an implementation is named under, and the ad server's name resolves within its own type, so `mock` is `ad-server.mock`. Also in the seam design - The principle in section 2 says which names are crate folders today. An integration that still lives in core carries the name its crate will have, held as a constant, so its section and table do not change when it moves out. - `[auction]` and `[proxy]` select the modules they run beside their own settings, and the text that said `[auction]` is not a provider type says what its `modules` list selects. - Where section 3.4 describes `main`, the profile ids are the ones `main` has, `standard`, `prebid-server` and `aps`. - Sign-off row 10 is reworded and row 14 added, and the revision record carries an entry for 2026-10-07. In the other specs - The pluggable modules design states the same rule for `implementation` and for `[auction]`. - The response header design names DataDome's settings under `[bot-protection.datadome]`. - The permission model design no longer cites `[integration] module`.
Section 3.6 has the three capability declarations take a name, the path under `crates/` of the crate the module lives in, and has `[ec] module`, `[geo] module` and `[device] module` read that name the way a section reads one. One registration can then supply a module of each type, each under the name of its own crate. The declarations and the device trait carry the names the code uses, `with_geo_module`, `with_device_module` and `DeviceModule`, and the composition paragraph describes the Edge Cookie module an adapter holds for itself where it named a slot the code does not have. Sign-off rows 10 to 14 carry a status, where a stray cell had left the column empty.
…etting Section 8 gains item 9. A secret reference is resolved only at the paths core's own fixed list names, DataDome's two secret settings are on that list by their path, and two more places in core's settings code name DataDome. A vendor crate with a secret of its own therefore reads it from the secret store itself, and moving DataDome out of core means changing all three. All of it is on `main`, and the series does not change it. Sign-off row 6 said no integration move needs a Rust core change and then listed moves that do. It says the Rust side is complete for a page integration that holds no secret, and names what an auction vendor, a vendor with a secret and APS each still need.
…load Section 8 item 6 said a module's validate function runs nowhere in a real deployment. The implementation gives the settings load a form that takes a deployment's builders, so the item says validate runs at startup for a deployment that loads its settings with them, and that the operator's CLI still does not carry them. Item 10 is new. The settings are validated as they load, and with the built-in implementations alone that refuses a demand source or an ad server one of the deployment's own builders supplies. The implementation carries the builders through the load and through deploy validation, and its probe supplies an ad server. A demand source from a crate outside core is not yet proven the same way.
Section 8 gains item 11. Only the Fastly adapter classifies a request and sets device signals, so a device module a crate supplies runs there and on no other adapter. On Fastly the entry point derives the signals before the request reaches the application, so the implementation hands it the module the registry resolved. Acceptance item 2's device signals can therefore be shown on Fastly and not on the dev server.
Section 8 item 10 said a demand source from a crate outside core was not yet proven. The implementation's probe supplies one, so the item says both reach the plan, the settings load and deploy validation from outside core, and that no auction is driven through either.
…run_with Acceptance item 1 and section 8 item 7 described the Fastly adapter as a binary with no entry point a vendor crate could reach. The implementation (IABTechLab#1094) makes it a library with a thin binary, and the round trip the acceptance criteria ask for runs on it under Viceroy. Item 7 now says what was built and what no test covers, and the two paragraphs that weigh the open items no longer count it among them.
The implementation moved every integration module out of core, so the passages that described the moves as work to come say what was built. - Vendor crates sit under `crates/<type>/<vendor>`, where the design said `crates/integrations/<vendor>`. - The auction is open to demand and ad server implementations from a crate, through the builder a page integration registers with. The previous revision recorded the auction plan as closed. - Section 4 says what each module's move needed, and that the moves are one commit each inside the implementation where the design planned one pull request each after it. Sign-off rows 5, 6 and 9 follow. - The source-file guard's list is generated, and the audit takes a module's section from the list of stock modules. - Section 8 gains four findings from the moves: request state for a module, a module's own secret settings, registration from the auction plan, and the builders every process that validates settings needs.
Carries the design specs for the module series and for the integration seam,
seven documents and no code. It targets
maindirectly, so the diff is onlythose files, and the specs are reviewed here before the code that implements
them in #1043 to #1047 and #1094.
The documents
2026-08-27-integration-provider-seam-design.md2026-07-30-pluggable-providers-design.md[ec] module,[device] moduleand[geo] module, and the configuration convention every type follows2026-07-30-permission-model-design.mdrules:tree in the permissions file, and permission signal modules run in a configured order2026-07-30-client-cycle-ec-resolve-design.md2026-07-30-provider-migration-rollout-design.md2026-07-30-integration-response-header-hook-design.mdprovider-code-registry.mdThe one configuration convention
Every type is a top-level table with a selector key inside it, and
[<type>.<name>]holds that selection's settings only when it has any:The key is
modulewhere a type runs one andmoduleswhere it runsseveral, for every type.
[ec],[geo],[device]and[ad-server]takemodule,[permission-signal]and[demand]takemodules, and a pageintegration's section takes whichever fits. A table's name is the
implementation it configures, unless the table carries an
implementationline, which names the implementation by its module path. A demand table
always carries one, because
demandis not the type its implementations arenamed under, and that is how two instances of one implementation run side by
side under names a deployment chooses. A name is parts joined by
., each oflower case letters, digits,
_or-. A module from a crate is named by itsfolder below
crates/, and a[demand]or[ad-server]label of theoperator's own is snake_case.
The specs also state what this buys, which is the part a reviewer should
weigh. A table the selector does not name stops startup rather than sitting
unused, an implementation this build does not have stops startup with the ones
it does have listed, a setting a module does not know is refused rather than
dropped, and demand and ad server endpoints must be HTTPS, or HTTP to a
loopback host. The operator-facing version of the same rules lands with the
code as
docs/guide/configuration-rules.mdin #1094.
The auction in the seam spec
Section 3.4 opens the three things the compiled auction plan keeps in core on
main, being a fixed list of three OpenRTB profiles, special cases for theprebid-serverandapsprofiles by name, andadserver_mockas the only adserver the plan accepts. The ad server, plain OpenRTB, Prebid Server and APS
become implementations that an integration registers on its builder, named
by module path as
ad-server.mock,auction-protocol.openrtb,auction.prebid-serverandauction.aps, and a deployment selects them with[demand] modulesand[ad-server] module.The compiled plan stays, and what it compiles changes from a fixed profile
list to whatever implementations the registered builders offer, so core names
no vendor. An auction-side vendor needs its own crate and a registration, on
the same terms as an identity, location or device module.
The bid renderer contract is unchanged. It is an open descriptor, a type tag
plus the payload the demand provider supplies, with the same serialized form,
so the response a page receives does not change.
The migration and rollout spec records the cost. The auction keys change with
no deprecation period, which is a deliberate choice for a major release, and a
configuration still carrying
[auction.providers]or[auction] mediatorisrefused with a message naming where the setting moved to.
Page changes as middleware
Section 3.7 of the seam spec. Today a page change is one of four hook traits
an integration implements, run in the integration's one position in the
pipeline, for every document. The spec replaces the four traits and their four
context types with one middleware contract, run only where an ordered
[[fetch]]or[[serve]]entry names it. An entry covers one media type,optionally only the requests under a path prefix, and lists the middleware in
order, so a publisher can run a change on part of a site, or place two
changes of one integration apart, without touching the auction's priority.
Entries carry no settings and activate nothing, which stays where it is
decided today.
The two phases make the shared-template rule structural rather than
declared. A fetch middleware is handed nothing about the reader, so what it
leaves is what the shared template stores for every reader. A serve
middleware runs per reader on that reader's copy and nothing it writes is
stored. Each integration converts in a change of its own, checked against the
differential harness, and the traits are deleted when the last one converts.
Governance question this raises
The seam spec's section 2 states a principle for the task force rather than for
the compiler. Tech Lab engineering owns core and reviews vendor crates, and
vendors ship and maintain their own integrations. That is the change that
removes the per-vendor core work the project pays for today, most recently the
LiveRamp module (#1054). Opening the auction side widens it, because a demand
source or an ad server is now a vendor crate too.
Declared interest
51Degrees is a vendor and would use this seam for its own integration (#1072).
The reasoning here applies to every vendor on the same terms, ourselves
included, and our own module follows the same route.