{"binning_method":"HEADLINE data_tier = KMeans(k=5) on the composite scalar, clusters relabeled so the highest-mean cluster = tier 1. data_tier_quantile = equal-frequency (positional quintile) binning, kept as a cross-check. K-means is headline because most countries have little censorship data and an equal-frequency split would force 20% into tier 1 regardless of composite.","composite":-0.650106,"composite_method":"per-feature rank-transform to [0,1] -> z-score -> equal-weight mean of 6 z-scores.","country":"KP","country_name":"North Korea","data_tier":5,"data_tier_quantile":4,"disagreement":4,"disagreement_magnitude":4,"disagreement_quantile":3,"feature_keys":["confirmed_censorship_incident_count","mean_block_rate","n_distinct_domains_blocked","n_blocking_methods_observed","forecast_mean_risk","dbscan_anomaly_frequency"],"feature_z":{"confirmed_censorship_incident_count":-0.7093,"dbscan_anomaly_frequency":-0.9979,"forecast_mean_risk":0.0,"mean_block_rate":-0.9193,"n_blocking_methods_observed":-0.6404,"n_distinct_domains_blocked":-0.6337},"features":{"confirmed_censorship_incident_count":0.0,"dbscan_anomaly_frequency":0.0,"forecast_mean_risk":0.0,"mean_block_rate":0.0,"n_blocking_methods_observed":0.0,"n_distinct_domains_blocked":0.0},"forecast_imputed":false,"generated_at":"2026-10-04T05:15:45.922934Z","hand_set_tier":1,"honest_caveats":["PROPOSAL ONLY \u2014 the hand-set country_geography.risk_tier is NOT modified by this script and stays authoritative until a human reviews the diff.","Data-derived tier reflects the trailing 365 days only. The hand-set tier may encode a longer political history (e.g. a decade of shutdowns) that recent data does not show. This is why North Korea / Turkmenistan / Eritrea \u2014 hand-set tier 1 \u2014 land in a low data tier: they have near-total censorship but almost no OONI probe coverage, so recent data shows little.","Tier boundaries are arbitrary at the margin \u2014 a country one rank/cluster-edge either side of a boundary moves a whole tier. data_tier (k-means) and data_tier_quantile are both reported precisely so boundary cases are visible; treat any single-tier disagreement as soft.","A country can legitimately be a low data tier by recent measurement and a high hand-set tier by political context the data does not capture (e.g. a repressive regime in a quiet measurement month, or a country with no probes at all).","forecast_mean_risk is only computed from the live model for the 21 countries with a recent forecast feature row; the rest are mean-imputed, so that feature is weaker for non-watched countries.","Measurement volume is very uneven \u2014 CN/RU/IR have thousands of non-IODA rows, most countries have under 100. mean_block_rate is empirical-Bayes-shrunk toward the global pooled rate so a 20-row country is not scored 100%-blocked, but DBSCAN-frequency and domain counts are still noisier for sparse countries.","The six features are correlated (block rate, blocked domains and incident count all move together), so the equal-weight composite is not an information-theoretic average."],"in_country_geography":true,"proposal_disclaimer":"PROPOSAL ONLY. The hand-set country_geography.risk_tier stays authoritative; this data-derived tier is advisory until a human reviews it. It is NOT consumed by any model.","schema":"voidly-data-driven-risk-tiers/v1","tier_convention":"tier 1 = HIGHEST censorship risk, tier 5 = lowest. Matches country_geography.risk_tier. disagreement = data_tier - hand_set_tier, so a NEGATIVE value means the data places the country at higher risk than the hand-set tier."}
