-
Notifications
You must be signed in to change notification settings - Fork 0
Record the E002 Stage B narrow result #11
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
40 changes: 40 additions & 0 deletions
40
...ations/E002__OBSERVATION__SCALAR_FLOW_PROJECTION_ASYMMETRY__v0.1__2026-07-14.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,40 @@ | ||
| # E002 Observation O001 — Scalar Flow Projection Asymmetry | ||
|
|
||
| **Artifact class:** Emergent observation | ||
| **Version:** v0.1 | ||
| **Date:** 2026-07-14 | ||
| **Status:** Preserved — requires a new experiment to extend | ||
| **Source experiment:** `SW.EXPERIMENT.E002` | ||
|
|
||
| ## Observation | ||
|
|
||
| The fixed scalar candidate-flow law did not preserve the diagnostic value of the typed field state. | ||
|
|
||
| For wait or credit denial, `1 - F*` produced pooled ROC-AUC 0.368152, substantially below receiver congestion at 0.772791. For transfer commit, `F*` improved on the weak vacancy-only baseline in the pooled aggregate but remained near chance at 0.509871 and reached the required 0.70 AUC in no scenario. For healthy-loop stall, the pneumatic and baseline scores were exactly equal. | ||
|
|
||
| ## Architectural significance | ||
|
|
||
| The result suggests that the field's useful structure is not a single congestion gradient. Transfers depend on distinctions the CBF contract treats as constitutional: readiness, receiver capacity, fault isolation, semantic validity, freshness, provenance, and reciprocal obligation state. | ||
|
|
||
| Multiplication into one scalar creates two characteristic losses: | ||
|
|
||
| 1. A zero positive gradient suppresses candidate flow even when an equal-occupancy transfer is valid. | ||
| 2. A capacity or fault barrier can dominate a transition for different reasons that should not become numerically interchangeable. | ||
|
|
||
| The observability layer remained coherent precisely because it kept those causes typed. The scalar projection became weak when it collapsed them. | ||
|
|
||
| ## Relationship to E001 O001 | ||
|
|
||
| E001 O001 found a sharp zero-vacancy liveness boundary for strictly local receiver-issued credit. E002 O001 adds a complementary finding: occupancy gradient alone cannot explain or predict the constitutional transition surface on either side of that boundary. | ||
|
|
||
| Together, the observations favor a vector or topological account of field state over an automatic scalar pressure law. They do not authorize cyclic exchange, global visibility, or a bypass. | ||
|
|
||
| ## Governance boundary | ||
|
|
||
| This observation cannot be used to tune E002 after inspection. Any alternative vector, graph, phase, or topological projection requires a newly frozen experiment with declared targets and baselines. No causal scheduler follows from this observation. | ||
|
|
||
| ## Evidence | ||
|
|
||
| - [E002 Stage B report](../results/reports/E002__STAGE_B_RESULTS__v0.1__2026-07-14.md) | ||
| - [Aggregate analysis](../results/summary/AGGREGATE_ANALYSIS__E002__a9923615.json) | ||
| - [E001 Observation O001](../../E001/observations/E001__OBSERVATION__FULL_RING_CIRCULATION_ASYMMETRY__v0.1__2026-07-14.md) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Binary file added
BIN
+3.53 MB
...002/results/raw/E002__CANONICAL_EVIDENCE__a9923615c09a7cd1afc452c16e505d70b98c0568.tar.gz
Binary file not shown.
1 change: 1 addition & 0 deletions
1
...ults/raw/E002__CANONICAL_EVIDENCE__a9923615c09a7cd1afc452c16e505d70b98c0568.tar.gz.sha256
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1 @@ | ||
| df090ed18647b3e793ecb8dffef5bbf1af3a7f679b4ffffb2d3c00103bfd5d08 E002__CANONICAL_EVIDENCE__a9923615c09a7cd1afc452c16e505d70b98c0568.tar.gz |
155 changes: 155 additions & 0 deletions
155
experiments/E002/results/reports/E002__STAGE_B_RESULTS__v0.1__2026-07-14.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,155 @@ | ||
| # E002 Stage B Results — Noncausal Pneumatic Instrumentation | ||
|
|
||
| **Artifact class:** Observational experiment report | ||
| **Version:** v0.1 | ||
| **Date:** 2026-07-14 | ||
| **Status:** Completed — Narrow | ||
| **Evidence class:** E4-scoped negative and reconstructive evidence within the E001 model | ||
| **Experiment:** `SW.EXPERIMENT.E002` | ||
| **Exit decision:** **Do not proceed** to causal Stage C with the tested scalar flow law | ||
|
|
||
| ## 1. Result in one sentence | ||
|
|
||
| The pneumatic layer reconstructed all 72 E001 runs exactly and preserved the required diagnostic distinctions, but its fixed scalar candidate-flow projection failed the predeclared scenario-general usefulness threshold and therefore remains descriptive rather than constitutive. | ||
|
|
||
| This result is evidence about a read-only projection of the frozen three-loop traces. It is not evidence of physical thermodynamics, pneumatic equivalence, calibrated energy, neural computation, or production-scale behavior. | ||
|
|
||
| ## 2. Canonical evidence identity | ||
|
|
||
| | Property | Value | | ||
| | --- | --- | | ||
| | Implementation commit | [`a9923615c09a7cd1afc452c16e505d70b98c0568`](https://github.com/AUo959/superloop-workshop/commit/a9923615c09a7cd1afc452c16e505d70b98c0568) | | ||
| | Source E001 commit | `81e7f859f71425bdba7603a61566f9fb47c116f9` | | ||
| | Runtime | CPython 3.13.11, standard library only | | ||
| | Execution route | Manual canonical invocation after local and fetched `main` tree equality was verified | | ||
| | E002 specification SHA-256 | `badb4844084fc14adf20935868941a95feebe724e366baa3dc4afdaf297b9821` | | ||
| | Instrumentation contract SHA-256 | `ab34a99d031630a00ce28b3785330a764eab8f1a7ff5eca07cb4afc8ffb4feae` | | ||
| | Source E001 archive SHA-256 | `584cc15bfeb3cc74dec9d9069cde26e1abaa6f1350c0aa12ae10a9784bd1663b` | | ||
| | E002 aggregate digest | `41d35fba334b4f23c773d16f78b6e7fe6939b4d43ca69f7b8827ea897af82898` | | ||
| | E002 archive SHA-256 | `df090ed18647b3e793ecb8dffef5bbf1af3a7f679b4ffffb2d3c00103bfd5d08` | | ||
| | Runs | 72 | | ||
| | Telemetry records | 207,360 | | ||
| | Reconstruction accounting checks | 48,600 | | ||
| | Determinism mismatches | 0 | | ||
| | Invalid runs | 0 | | ||
|
|
||
| The Python 3.13 aggregate digest is byte-identical to the preceding Python 3.12 dry-run aggregate digest. The repository preserves the [compressed canonical evidence](../raw/E002__CANONICAL_EVIDENCE__a9923615c09a7cd1afc452c16e505d70b98c0568.tar.gz), its [checksum](../raw/E002__CANONICAL_EVIDENCE__a9923615c09a7cd1afc452c16e505d70b98c0568.tar.gz.sha256), all [run summaries](../summary/), the [aggregate analysis](../summary/AGGREGATE_ANALYSIS__E002__a9923615.json), and the [provenance record](../summary/PROVENANCE__E002__a9923615.json). | ||
|
|
||
| ## 3. Validity gate | ||
|
|
||
| The offline E002 implementation passes its hard validity gate: | ||
|
|
||
| - the frozen E001 archive, manifest, trace digests, summary self-digests, and event sequences verified before replay; | ||
| - all 72 source runs reconstructed without E001 simulator or workload imports; | ||
| - reconstructed post-tick occupancy-time matched every E001 run summary exactly; | ||
| - every canonical rational was reduced and no canonical float, infinity, or `NaN` was emitted; | ||
| - pre-tick lineage contained no same-tick or future event identifier; | ||
| - capacity blockage, fault state, semantic rejection, stale authority, and reciprocal debt remained separate typed channels; | ||
| - all ten instrumentation invariants passed for every run; and | ||
| - repeated instrumentation produced identical telemetry, run-summary, and aggregate digests. | ||
|
|
||
| The generator emitted no E001 transition and did not modify the source evidence. | ||
|
|
||
| ## 4. Hypothesis disposition | ||
|
|
||
| | Hypothesis | Disposition | Evidence | | ||
| | --- | --- | --- | | ||
| | H1 — Exact noninterference | Offline clause supported; optional shadow clause not evaluated | The replayer is a separate read-only process with no E001 runtime imports, and every source digest remained unchanged. No in-process shadow observer was added. | | ||
| | H2 — Coherent reconstruction | Supported | All 72 runs, 207,360 records, 48,600 declared occupancy checks, work ledgers, modes, readiness states, and occupancy-time identities reconstructed without an invalid run. | | ||
| | H3 — Typed diagnostic separation | Supported | Capacity, fault, semantic, stale-authority, and debt states remained separate; unavailable trust and uncertainty stayed `not_observed`. | | ||
| | H4 — Diagnostic usefulness | Not supported | No target met both the six-of-eight scenario AUC threshold and pooled lift threshold. | | ||
|
|
||
| H1 is deliberately not overstated. Offline noninterference is demonstrated by construction and digest verification; shadow-mode parity remains unevaluated and is not needed to interpret the offline negative result. | ||
|
|
||
| ## 5. Fixed diagnostic comparisons | ||
|
|
||
| Decimal AUC values below are noncanonical renderings of the preserved exact fractions. | ||
|
|
||
| | Target | Pneumatic pooled AUC | Scalar baseline AUC | Lift | Scenarios at AUC ≥ 0.70 | Threshold result | | ||
| | --- | ---: | ---: | ---: | ---: | --- | | ||
| | T1 — wait or credit denial | 0.368152 | 0.772791 | -0.404639 | 0 of 8 | Fail | | ||
| | T2 — transfer commit | 0.509871 | 0.429770 | +0.080101 | 0 of 8 | Fail | | ||
| | T3 — healthy-loop stall within five ticks | 0.989479 | 0.989479 | 0 | 3 of 8 | Fail | | ||
|
|
||
| ### 5.1 T1 — Wait or credit denial | ||
|
|
||
| The tested score `1 - F*` was materially worse than receiver congestion. Five scenarios contained no positive T1 case, so their AUC was correctly undefined rather than coerced to zero. In the three discriminating scenarios, no pneumatic AUC reached 0.70. | ||
|
|
||
| The direction matters: multiplying eligibility, vacancy conductance, and positive congestion gradient suppresses the score in several states where waits are caused by full capacity or fault boundaries. The scalar does not compress those typed causes safely. | ||
|
|
||
| ### 5.2 T2 — Transfer commit | ||
|
|
||
| Candidate flow exceeded the simple vacancy baseline in the pooled comparison by approximately 0.0801, but its pooled AUC remained only 0.5099 and no scenario reached the 0.70 threshold. The positive pooled lift is therefore not scenario-general evidence of useful transfer prediction. | ||
|
|
||
| The result rejects the temptation to select only the favorable pooled lift. The predeclared decision required both lift and scenario coverage. | ||
|
|
||
| ### 5.3 T3 — Healthy-loop stall | ||
|
|
||
| The pneumatic score and raw congestion baseline were exactly identical. Both separated the three scenarios containing positive stall cases extremely well, but the pneumatic formulation added no information or compression beyond its declared scalar input. T3 is coherent but redundant. | ||
|
|
||
| ## 6. What the negative result reveals | ||
|
|
||
| The failure is informative. A transfer in this model is not governed by a single occupancy gradient. It also depends on local readiness, capacity, fault mode, schema and provenance validity, lease freshness, and reciprocal obligation handling. Those dimensions were intentionally kept separate by the CBF constitution. | ||
|
|
||
| The tested `F* = eligibility × conductance × positive gradient` collapses that structure too aggressively. Equal-occupancy exchanges have zero gradient even when a transfer is valid, while zero receiver vacancy can represent a capacity boundary that should remain typed rather than absorbed into one pressure score. | ||
|
|
||
| This supports the original architectural caution: pneumatic language is most coherent here as a vector observability layer around constitutional interlocks, not as a replacement operating system or automatic scalar scheduler. | ||
|
|
||
| The projection asymmetry is preserved separately as [E002 Observation O001](../../observations/E002__OBSERVATION__SCALAR_FLOW_PROJECTION_ASYMMETRY__v0.1__2026-07-14.md). | ||
|
|
||
| ## 7. Outcome classification | ||
|
|
||
| The overall E002 outcome is **Narrow**. | ||
|
|
||
| Retain as noncausal observability: | ||
|
|
||
| - exact congestion and vacancy; | ||
| - typed finite resistance and zero-vacancy blockage; | ||
| - separate pressure-vector channels; | ||
| - typed capacity, fault, semantic, stale-authority, and debt states; | ||
| - integer phase indices as descriptive activity labels; | ||
| - occupancy-time and zero-vacancy-cycle diagnostics; and | ||
| - the offline reconstruction and evidence pipeline. | ||
|
|
||
| Narrow or reject as demonstrated diagnostics: | ||
|
|
||
| - reject `1 - F*` for T1 denial prediction in this model; | ||
| - narrow `F*` to an explicitly experimental descriptor, not a transfer predictor; | ||
| - treat the T3 resistance score as a restatement of congestion, not independent pneumatic evidence; | ||
| - make no usefulness claim for phase or dissipation beyond deterministic description; and | ||
| - do not fit weights or alter the frozen formula after seeing this result. | ||
|
|
||
| ## 8. Exit decision | ||
|
|
||
| **Do not proceed to causal Stage C with the tested scalar law.** | ||
|
|
||
| The evidence does not justify allowing `F*`, resistance, dissipation, or phase to authorize work or replace the minimal interlock contract. Optional shadow instrumentation is also deferred because it would prove parity for a projection that has not demonstrated sufficient diagnostic value. | ||
|
|
||
| A later experiment may test a fixed vector or topological diagnostic that preserves readiness, typed barriers, and zero-vacancy cycles without scalar collapse. That work must receive a new experiment identifier, frozen targets, and equivalent-work controls. It may not be described as a tuned rerun of E002. | ||
|
|
||
| ## 9. Limitations | ||
|
|
||
| - The evidence inherits E001's three-loop, one-direction topology and deterministic workloads. | ||
| - The eight scenarios are adversarial within the model, not independent external replication. | ||
| - AUC is undefined in scenarios with only one target class; undefined values cannot support H4. | ||
| - T3's high AUC reflects the same congestion information in both compared scores. | ||
| - Phase and dissipation were recorded but did not receive separate calibrated targets. | ||
| - No continuous fluid dynamics, physical pressure, energy, hardware cost, wall-clock behavior, or learning was modeled. | ||
| - Shadow-mode noninterference was not evaluated. | ||
|
|
||
| ## 10. Reproduction | ||
|
|
||
| After checking out the implementation commit, run: | ||
|
|
||
| ```bash | ||
| PYTHONPATH=src python tools/generate_e002_evidence.py \ | ||
| --output build/e002-canonical \ | ||
| --source-commit a9923615c09a7cd1afc452c16e505d70b98c0568 | ||
| ``` | ||
|
|
||
| Verify the preserved archive from `experiments/E002/results/raw/` with: | ||
|
|
||
| ```bash | ||
| sha256sum --check \ | ||
| E002__CANONICAL_EVIDENCE__a9923615c09a7cd1afc452c16e505d70b98c0568.tar.gz.sha256 | ||
| ``` | ||
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
For anyone using this reproduction block to verify the preserved E002 archive,
build/e002-canonicalcannot recreate the committed archive checksum:tools/generate_e002_evidence.pyembeds the exact--outputvalue intoprovenance.jsonbefore packaging, while the committed provenance records/tmp/e002-canonical-a9923615. The replay data may match, but the tarball bytes and SHA-256 will differ, making the canonical evidence artifact appear non-reproducible unless the recorded invocation/runtime is used or the instructions compare only the aggregate digest.Useful? React with 👍 / 👎.