Skip to content

Site-Sliced Release Evidence by Scale

Last updated: 2026-05-24

Site-sliced release evidence prevents a common autonomy failure mode: a model improves globally, passes one replay suite, or succeeds at one site, then gets treated as approved everywhere. Fleet autonomy does not release to an abstract fleet. It releases to a bounded operational design domain (ODD): site, route, zone, weather, lighting, vehicle kit, sensor calibration, map state, runtime package, semantic taxonomy, task, and operating procedure.

The rule for S3+ autonomy is direct: promotion is by ODD cell, not by fleet percentage. A 5% canary on easy daytime service roads does not validate night stand entry, terminal frontage, jetblast zones, public pedestrian plazas, warehouse aisles, port quays, mine haul roads, campus crossings, or utility corridors.

Use evaluation-platform-replay-gates-by-scale.md for the underlying evaluation manifest, metric spec, replay package, runtime package smoke, shadow/canary denominator, and waiver fields that make an ODD-cell release decision comparable and auditable.


Release Unit

A release unit should be smaller than "fleet" and larger than one vehicle. It is the minimum scope where evidence can justify behavior authority.

DimensionExamplesWhy it matters
SiteAirport A, port terminal B, yard C, campus district DLocal geometry, rules, traffic mix, markings, and operational norms differ
ZoneStand, apron lane, terminal frontage, warehouse aisle, loading bay, quay, public crossingHazard classes and right-of-way rules differ
Route/taskGate-to-bagroom, tug crossing, inspection route, service-road transitA model can be safe for transit but weak for close-proximity tasks
Time/weatherDay, night, rain, fog, de-icing, low sun, snow, dustPerception and planning distributions shift
Vehicle/sensor kitLiDAR model, camera coverage, radar, compute target, firmwareRuntime and calibration behavior differ
Map/calibration stateMap package, semantic layer, changed tiles, calibration packageActive artifacts bound what was evaluated
Model/runtime packageModel version, TensorRT/ONNX/container, class order, thresholdsOffline checkpoint approval does not prove runtime package approval
Operational authorityAdvisory, shadow, supervised, limited autonomous, full autonomousEvidence depth must match the consequence of wrong behavior

The release packet should name the ODD cell explicitly. "Model v42 approved for Airport A" is too broad. "Model v42, runtime bundle R17, map M31, calibration C9, Airport A stand-entry task, daylight/dry, vehicle kit K2, supervised-to-autonomous canary" is reviewable.


Scale Ladder

MLOps scaleSite-slice postureMinimum evidencePromotion anti-pattern
S0 notebook researchSlice labels are exploratory notesRecord which site/zone examples were inspectedClaiming transfer from a few screenshots
S1 repeatable prototypeFrozen validation split has basic site/task tagsReport metrics by site, class, and known ODD slice if availableTreating one validation split as representative of operations
S2 single-product productionRelease packet names the intended site or serviceOffline holdout, replay smoke, package load, rollback targetShipping to operators from aggregate mAP alone
S3 fleet and multi-siteEvery rollout has an ODD-cell manifest and local holdoutSite/route/weather slices, replay, shadow, canary, delayed-label joinsOne global champion hides local regressions
S4 regulated safety-criticalODD-cell evidence links to safety-case claims and waiver stateClaim/evidence table, hazard slices, monitor thresholds, rollback drill, incident trigger policyWaiving a local regression without owner, expiry, or mitigation
S5 platform scaleShared release service enforces site-slice policyPolicy-as-code, evidence completeness, tenant/site isolation, scorecard APITeams bypass platform with local canary spreadsheets

Small fleets can still require S4 controls when the release affects people, aircraft, public roads, protected zones, or regulatory evidence. Large offline research programs can remain S1 if outputs do not affect operations or release evidence.


Evidence Stack

GateWhat it provesRequired for
Data contractThe requested ODD cell is named with site, route, task, weather, vehicle, map, calibration, and taxonomy scopeS2+
Offline holdoutThe candidate does not regress on independent local data, with split lineage checked through dataset-split-leakage-controls-by-scale.mdS1+
Scenario replayKnown incidents, map-change cases, hazard classes, and required maneuvers still passS2+
Runtime package testThe exact deployable artifact loads and meets latency/memory/class-order constraintsS2+
Shadow modeThe candidate behaves acceptably on live inputs without control authorityS3+
Canary by ODD cellLimited operational exposure shows acceptable health, disagreement, intervention, and delayed-label evidenceS3+
Monitor and rollback readinessRuntime monitors can detect degradation and rollback can execute for the same cellS3+
Safety-case reviewResidual risk, waivers, mitigations, and reportability are approvedS4+
Platform policy checkShared evidence and release services enforce the fields and blocks automaticallyS5

The gates are cumulative. A canary does not replace replay; replay does not replace shadow; shadow does not replace a safety-case decision when behavior authority changes.


ODD-Cell Manifest

FieldMinimum contents
release_candidate_idModel registry version, alias, and release packet ID
artifact_setModel, runtime, map, semantic taxonomy, calibration, telemetry schema, prompt/labeler/evaluator, replay package, rollback target
site_scopeSite ID, owner, customer/tenant, geography, local rules, data residency
zone_scopeRoutes, lanes, stands, crossings, restricted/protected zones, map tile IDs
task_scopeDriving, inspection, FOD detection, tug crossing, routing, labeling, map publication, advisory-only
environment_scopeWeather, lighting, surface state, de-icing, jetblast, dust, GNSS state, construction
vehicle_scopeVehicle type, sensor kit, compute hardware, firmware, maintenance/calibration status
data_evidenceTraining snapshot, split manifest, local holdout, leakage check, label QA, source-map acceptance if map-derived
evaluation_evidenceEvaluation manifest, evaluator version, offline metrics, confidence intervals, replay suite, runtime smoke, hazard-slice metrics, known failures
shadow_evidenceExposure hours, denominator, disagreement taxonomy, interventions, operator notes, trigger yield
canary_evidenceCohort, start/end time, exposure denominator, monitor thresholds, rollback triggers
delayed_label_evidenceReviewer samples, incident review, false positive/negative estimates, map-change review
decisionApproved, held, restricted, canary-only, rolled back, deprecated
expiryEvidence expiry date, required revalidation triggers, waiver owner

The manifest should be machine-readable even if the release review is written in Markdown. The same fields drive deployment, monitoring, incident response, and future invalidation.


Slice Taxonomy for Managed Sites

Slice familyAirside examplesOther non-road examplesEvidence focus
Protected peopleGround crew, security, pedestrians near terminalWarehouse pickers, yard workers, campus pedestriansMissed-person recall, false-free-space, safe stop behavior
Large movable assetsAircraft, GSE, jet bridges, baggage trainsContainers, forklifts, trailers, mining trucksClearance, occlusion, prediction, protected-zone rules
Small hazardsFOD, chocks, cones, hoses, debrisPallets, tools, rocks, cables, fallen cargoRare-class recall, close-range braking, scenario replay
InfrastructureStands, lane markings, signs, fences, polesRacks, curbs, gates, quay edges, utility cabinetsSemantic-map correctness, localization stability, route constraints
Environmental stateRain, fog, low sun, night, de-icing, jetblastDust, snow, indoor glare, wet floor, mudSensor degradation, OOD, perception health, ODD boundary
Operational phasePushback, turnaround, refuel, loading, maintenanceLoading, shift change, peak yard flow, public eventProcedure-specific behavior and operator handoff
Map stateNew construction, changed stand, temporary closureTemporary aisle closure, roadworks, site eventMap freshness, changed-tile replay, semantic release-state handling

Non-road sites often have more repeated structure than public roads, but the local procedures are stronger. A class that is harmless in one zone can be release-blocking in another. A baggage cart, forklift, maintenance cone, or stationary person can be either expected context, protected obstacle, transient map artifact, or incident signal depending on the slice.


Release State Machine

StateMeaningExit criterion
candidate_globalCandidate passed generic offline checksODD-cell manifest created
candidate_siteCandidate has target site/task scopeLocal holdout and replay suite attached
site_shadowCandidate runs without authority in the target cellShadow exposure and disagreement review pass
site_canaryCandidate has limited authority in controlled cohortCanary metrics, monitor health, and delayed labels pass
site_championCandidate is approved for this ODD cellRelease decision signed and rollback verified
restricted_championCandidate is approved with exclusions or mitigationsRestriction expires or evidence closes gap
heldEvidence incomplete or regression unresolvedMissing evidence supplied or candidate rejected
quarantined_cellField signal invalidates current approval for a cellRollback or containment plus root-cause review
deprecated_cellCell approval retired by new artifact, map, or ODD changeReplacement release or permanent withdrawal

Alias movement should respect the state. A model can be champion for Airport A daylight service-road transit and only site_shadow for Airport B night stand entry.


Training and Adaptation Choices

ApproachAdvantagesDisadvantagesBest use
One global modelOperationally simple, more data, one runtime packageHides local regressions, weak local terminology, can overfit dominant siteS2 single site or S3 when sites are very similar
Global model plus local thresholdsFast adaptation, small release deltaCan mask calibration or label-quality issues; needs per-slice evidenceConfidence/calibration differences by site/weather
Global backbone plus site adapters/LoRAStrong transfer with small local dataAdds adapter registry, compatibility, and rollback complexityMulti-site perception where local visuals differ
Site-specific modelBest local specializationFragmented evidence, harder fleet learning, more runtime variantsHigh-value sites with unique ODD or regulatory constraints
Mixture-of-experts / route-gated modelCan route by ODD cellRouting errors become safety-critical; harder auditS5 platform with strong ODD classification and policy
No promotion; collect more dataPrevents unsafe releaseSlower rollout and higher labeling costAny slice with insufficient exposure or unresolved hazard regression

The default for managed-site autonomy is global backbone plus local evidence. Local adapters are attractive for different airports, yards, or warehouses, but the release unit must include the adapter ID, training data, local holdout, runtime package, and rollback artifact.


Statistical Discipline

Site-sliced release evidence should report uncertainty, not only point estimates.

Evidence typeMinimum discipline
Offline metricsConfidence intervals, class/slice sample counts, unchanged thresholds unless justified
Rare hazardsUse scenario replay and targeted sampling; do not rely on natural exposure alone
Shadow exposureCount relevant opportunities, not only hours or kilometers
CanaryCompare to baseline/control for the same ODD cell and time window
Delayed labelsSample enough negatives and positives to estimate false-free-space and missed-object risk
MonitoringRecord denominator: active hours, routes, sites, map tiles, vehicle cohort, weather bins
WaiversName owner, expiry, operational mitigation, and revalidation trigger

Natural exposure is weakest exactly where safety matters most. FOD, personnel intrusion, aircraft proximity, construction changes, and near-conflict behavior need designed replay, targeted mining, and local review because waiting for field frequency can be unsafe and statistically slow.


Semantic Map and ML-SLAM Coupling

For semantic maps and ML-related SLAM, site-sliced release evidence must include map state:

Map/SLAM artifactRelease-slice dependency
Source mapSource-map acceptance package must cover the target tiles and sessions
Semantic layerClass taxonomy and release-state labels must match the consuming model
Map hygiene layerDynamic residual, static-transient, FOD-candidate, artifact, and unknown-review states must be preserved
CalibrationLocal sensor calibration and time sync must match training/eval/replay evidence
Changed tilesRuntime approval does not carry across map changes without replay or impact review
Pseudo-label exportsTraining labels derived from a site map inherit that site's evidence and invalidation rules

A segmentation model trained from Airport A map-derived labels may be a strong prior for Airport B, but Airport B release still needs local map acceptance, holdout, and ODD-cell evidence. Public road datasets, urban district proxies, or utility/facade benchmarks are pretraining evidence; they are not local release evidence.


Acceptance Checks

  • Every production or safety-affecting release has an ODD-cell manifest.
  • Aggregate metrics are accompanied by target site, route, weather, object, map-state, and vehicle-kit slices.
  • Training, local holdout, replay, source-map tile, and safety-holdout assignments have split IDs and leakage reports for the target ODD cell.
  • The release packet names the exact model/runtime/map/calibration/telemetry/taxonomy artifact set.
  • Replay scenarios include known incidents, local map changes, rare objects, protected people, and operating procedures for the target cell.
  • Shadow/canary evidence covers the same ODD cell requested for approval.
  • Delayed-label review samples are tied to the canary cohort and active artifact IDs.
  • Monitor thresholds, suppression rules, and rollback triggers are versioned release artifacts.
  • A local regression cannot be waived without owner, expiry, mitigation, and revalidation trigger.
  • Rollback is tested for the same vehicle kit, runtime, map, calibration, and deployment channel.
  • Expansion to a new site, task, vehicle kit, map state, or weather band creates a new release decision.

Failure Modes

Failure modeConsequenceControl
Global mAP approvalLocal rare-class or ODD regression reaches operationsRequire ODD-cell manifest and local blocker metrics
Canary by fleet percentageEasy cells dominate evidenceCanary by site, route, weather, task, map state, and vehicle kit
Shadow data from wrong ODDEvidence does not support requested releaseGate on matching ODD-cell exposure
Local holdout leaks into trainingSite evidence overstates performanceSplit lineage and leakage checks by site/task
Runtime artifact differs from evaluated artifactCanary does not test what will deployCompatibility manifest and package hash
Map change not reflected in model releaseModel runs against unseen or stale map semanticsChanged-tile replay and map/model compatibility check
Waiver becomes permanentKnown local risk remains unresolvedWaiver owner, expiry, mitigation, and review cadence
Adapter sprawlMany local models with weak evidenceAdapter registry, shared backbone policy, per-adapter rollback
Incident cannot be scopedFleet cannot know which cells are affectedActive artifact IDs in monitoring and incident evidence

  • mlops-scale-research-scope.md - MLOps maturity ladder and research backlog.
  • model-governance-release-evidence.md - registry aliases, release packets, approval, and rollback evidence.
  • mlops-scorecards-and-kpis-by-scale.md - release-blocking metrics and operating cadence.
  • mlops-reference-architectures-by-scale.md - S2-S5 release lanes and artifact interfaces.
  • evaluation-platform-replay-gates-by-scale.md - evaluation manifests, metric specs, replay/runtime gates, shadow/canary evidence, and waiver controls.
  • dataset-split-leakage-controls-by-scale.md - split manifests, local holdout leakage controls, and training/evaluation split architectures.
  • data-flywheel-airside.md - active learning, local holdouts, shadow/canary validation, and data mining.
  • ../data-platform/replay-scenario-mining-ops.md - scenario mining and replay package promotion.
  • ../ota/perception-slam-artifact-compatibility-matrix.md - model/map/calibration/runtime compatibility.
  • ../../40-runtime-systems/ml-deployment/production-ml-deployment.md - runtime packaging, shadow, canary, and rollback.
  • ../../60-safety-validation/verification-validation/shadow-mode.md - dual-stack shadow-mode architecture.
  • ../../60-safety-validation/runtime-assurance/online-perception-monitoring-odd-enforcement.md - runtime ODD and perception health monitoring.
  • ../../30-autonomy-stack/perception/overview/aggregated-map-semantic-segmentation.md - semantic-map release-state and training-export controls.

Sources

Public research notes collected from public sources.