Validation

Evidence before trust.

Urd Atlas should not ask technical customers to trust a black box. This page separates dated methodology validation from live published-data diagnostics: whether the classification rules are internally coherent and robust, and whether the current data windows have enough variation, confidence coverage and transition structure for analysis.

Validation boundary

Internally tested. Not a claim of external ground truth.

Network-state labels are explicitly defined descriptive categories. Validation tests reproducibility, internal rule consistency, threshold robustness and signal dependence. It does not claim a hidden objective regime truth, calibrated label probability, forecast or recommendation.

Validation Report v1 · 12 August 2026

The current classifier has been tested for internal consistency and robustness.

The historical consistency audit covers all dated Meta rows available at the audit date. Sensitivity and ablation results are scoped to the current active rulesets so older methodology versions are not incorrectly judged by today's rules.

Read methodology

Historical consistency

2,464

Published chain-days audited across Bitcoin, Ethereum, Arbitrum and Base.

Hard rule violations

0

No audited label-rule, confidence-range or candidate-signature consistency failures.

Baseline reproduction

2,100 / 2,100

Current-ruleset candidate labels exactly reproduced before counterfactual tests.

±5% threshold test

96.5–97.6%

Candidate labels unchanged when key threshold families were moved together by ±5%.

Threshold sensitivity

Combined ±10% perturbations changed 6.57–8.10% of current-ruleset candidate labels. Percentile boundaries were the most influential threshold family; robust z-score and momentum changes had materially smaller marginal effects.

Signal ablation

Removing individual signals changed labels in semantically expected ways: fee removal primarily reduced low-friction or congestion classifications, while primary capacity proxies had the strongest effect on capacity-sensitive regimes. This tests structural dependence, not predictive accuracy.

Live published-data diagnostics

What does the currently published window look like?

These numbers are calculated from the best currently available published Meta window and can change as new rows are published. They are separate from the dated Validation Report v1 results above.

Rows inspected

1,460

Across currently available published meta windows.

Transitions observed

499

Regime changes in the diagnostic windows.

Good-confidence share

53%

Average share of rows with confidence ≥ 0.70.

Status legend

How to read the chain status labels.

These labels are caution flags for using the diagnostic sample. They do not say whether a chain is good or bad; they say how carefully the current published window should be interpreted.

Usable diagnostic sample

Enough observations, acceptable regime variation and at least half of rows carrying Good confidence.

Confidence-limited

There may be enough observations and variation, but fewer than half of rows pass the Good-confidence gate.

Low variation

One regime dominates at least 90% of rows, so segmentation may say more about the dominant state than about regime differences.

Too sparse

Fewer than 30 observations are available in the best published window, so the diagnostic sample is not yet strong.

Current published-data diagnostics

Can this chain be segmented?

A customer should quickly see whether a chain has enough observations, enough regime diversity and enough confidence to support downstream reporting or model diagnostics.

Open Analyst Kit

BTC

Bitcoin

2025-08-30 → 2026-08-29

Status

Confidence-limited

Obs.

365

last365d

Dominant

Stable

44% of rows

Transitions

41.1

per 100 obs.

Median run

1.0

observations

Regime distribution

Entropy 0.89
Stable161
Heating42
Congested35
Cheap76
Unknown51

Confidence coverage

Avg. 57%
Good58
Caution256
Degraded51
Missing0

ETH

Ethereum

2025-08-29 → 2026-08-28

Status

Usable diagnostic sample

Obs.

365

last365d

Dominant

Stable

42% of rows

Transitions

41.9

per 100 obs.

Median run

1.0

observations

Regime distribution

Entropy 0.83
Stable153
Heating50
Congested55
Cheap103
Unknown4

Confidence coverage

Avg. 76%
Good229
Caution132
Degraded4
Missing0

ARB

Arbitrum

2025-08-24 → 2026-08-23

Status

Usable diagnostic sample

Obs.

365

last365d

Dominant

Stable

58% of rows

Transitions

29.0

per 100 obs.

Median run

2.0

observations

Regime distribution

Entropy 0.71
Stable211
Heating49
Congested40
Cheap65
Unknown0

Confidence coverage

Avg. 79%
Good241
Caution124
Degraded0
Missing0

BASE

Base

2025-08-24 → 2026-08-23

Status

Usable diagnostic sample

Obs.

365

last365d

Dominant

Stable

54% of rows

Transitions

24.7

per 100 obs.

Median run

2.0

observations

Regime distribution

Entropy 0.71
Stable198
Heating41
Congested32
Cheap94
Unknown0

Confidence coverage

Avg. 79%
Good252
Caution113
Degraded0
Missing0

Class balance

A useful regime feature needs enough variation to segment analysis. Dominant-class share and entropy show whether a chain is informative or mostly constant.

Transition stability

A state layer should not flip randomly, but it also cannot be static. Transitions per 100 observations and median run length make that tradeoff visible.

Confidence coverage

Confidence should be used as a quality gate. This page shows how much of each chain has Good, Caution, Degraded or missing confidence.

Operational usefulness

The practical question is whether the feature helps explain daily app, protocol, fee, support or usage metrics more cleanly than an internal one-off rule.

Point-in-time discipline

Observation date, publication date and available-at timing must stay separate so downstream analysis does not accidentally use unavailable context.

Explicit limitations

Validation should state where the data is sparse, stale, low-confidence or too imbalanced to support a strong conclusion.

Minimum proof standard

A buyer needs proof of usefulness, not proof of elegance.

The validation layer should answer whether Urd Atlas is materially better than a customer building a small internal rule set. The answer may be stronger stability, more transparent confidence handling, reproducibility, lower maintenance cost or a measurable workflow improvement.

Do not overclaim

Validation is not outcome marketing.

This page should not claim that regimes determine external outcomes. Its job is to make the data product credible: where it varies, when it is reliable, what it can segment and where customers should not use it.

Next proof layer

Turn diagnostics into reproducible examples.

The next version should add downloadable notebooks that join Urd Atlas to public chain-activity or protocol-activity datasets and show a real regime-conditioned analysis without changing the product boundary.