Comparison sensitivity
This page records the two Monte Carlo stability tests behind the comparison map: resampling the axis weights across their whole space with the sub-scores frozen, and perturbing every sub-score by up to a point at the published weights. It holds the designs and seeds, the rank intervals for all twenty-two products, the head-to-head crossing statistics, one withdrawn claim, and a script that reproduces the record.
The first version of the methodology page tested exactly one alternative weighting and reported one crossover. The peer review did what we should have done: a Monte Carlo across the whole weight space and across sub-score uncertainty. We adopted the design and re-ran it with our own code rather than citing the reviewer's numbers. This page publishes ours; the full record of runs and the agreement check against the reviewer's figures is in review/index/SCORES.md, and every sub-score the tests draw on is on the scores page.
The two tests
Each test runs N = 200,000 samples per axis:
- Weights sampled, scores frozen. Non-negative weights drawn uniformly over each axis's weight simplex (normalized Exp(1) draws); every published sub-score held fixed. This asks: how much of the ordering is our weighting judgement?
- Weights frozen, scores perturbed. The published weights kept; independent uniform ±1 noise added to every sub-score of every product, clipped to 0–10. This asks: how much depends on integer judgement calls being exactly right?
A rank here is 1 + the number of products scoring strictly higher. Primary run: seed 20260807. Independent re-implementation (different code and generator, seed 271828). Recomputed 2026-08-07 over the 22-product field when Obscura was scored: rank is field-relative, so the intervals below were re-run with Obscura's row in rather than left stale, and the working record in review/index/SCORES.md preserves the 21-product runs as the reviewer-reproduction record. Recomputed again the same day when URnetwork was rescored on the default-on ruling (the methodology ledger): the tables below are the post-rescore record, re-run under both implementations and both seeds, and the pre-rescore runs stay in the working record's runs table. On the 22-product field the two implementations agree on every interval and median except five weights-test cells at near-tied adjacent pairs (0.05 apart in point score), noted in the working record; the same five cells differ before and after the rescore. Every score-noise interval, the kind the overview page publishes, is identical across both runs, as are URnetwork's and Obscura's rows in full.
URnetwork's results
Obscura's entry moved the X rank statistics: weights median 3→4, noise interval 2–3 → 2–4. The default-on rescore moved them again: weights median 4→3, weights interval 2–5 → 1–4. No noise-test rank statistic and no Y statistic moved either time.
| Statistic | Test 1, weights sampled | Test 2, ±1 score noise |
|---|---|---|
| X rank, median (5th–95th) | 3 (1–4) | 3 (2–4) |
| X score, 5th–95th | (weights vary) | 6.793–7.813 |
| Ahead of Mullvad on X | 74.0% (exactly 71/96) | 99.9% |
| Ahead of Obscura on X | 100% (degenerate; see below) | 93.9% |
| Ahead of NymVPN on X | 67.0% (exactly 343/512) | 30.6% |
| Y rank, median (5th–95th) | 3 (1–15) | 7 (4–11) |
| Y score, 5th–95th | (weights vary) | 5.747–6.653 |
| Y below 5.00 | never, by construction (see the withdrawn claim) | 0 of 200,000 (about 4.4σ under this noise model: rare, not impossible) |
The two axes do not carry equal confidence. URnetwork's X rank stays inside 1–4 across every weighting sampled: near the top of the X axis, which judgement you make decides the order, but the index itself is comparatively weight-stable. Its Y rank runs 1–15 across the same weight space, and even at the published weights the ±1-noise median is 7th, not the tied-5th its point score suggests. The Y index is far more weight-sensitive than X, and every Y placement on the map should be read with correspondingly less confidence.
The head-to-heads
Mullvad. The two Mullvad numbers remain the most informative sentence available about this chart, and the default-on rescore changed what they say. Before it, URnetwork's X lead over Mullvad was 55.55% under weight uncertainty, a coin flip that depended almost entirely on ranking structure (separation plus end-to-end, 0.60 combined) above verification (0.30). After the rescore the sub-score difference is (+5, +3, −3, −1): URnetwork leads both structure parts, Mullvad leads verification and, narrowly, post-quantum, and the same weight sampling gives URnetwork the lead 74.0% of the time (the exact simplex volume is 71/96 ≈ 73.96%). Under ±1 score noise at the published weights it is 99.9%. The lead is no longer a coin flip, but the residual weight sensitivity survives in scoped form: a reader who weights verification at about half the axis still gets Mullvad ahead, and that remains a reasonable reading, not a mistake. The worked example is in the re-weighting section below.
Obscura. The comparison inverted with the rescore. The two now tie on structure (E+S 16–16, both content legs default-on, Obscura's by construction) and tie on V, so the entire 0.70 gap is the post-quantum part Obscura lacks. A pure single-part difference, (0, 0, 0, +7), makes the weights test degenerate: any weighting with a nonzero post-quantum weight preserves the order, so "100% of sampled weightings" describes the difference's shape, not robustness. The informative test is score noise, under which URnetwork leads 93.9% of the time: outside a single point's reach, since the largest one-point, one-part swing is 0.35 against a 0.70 gap. The pre-rescore statistic was 69.5%, which was not outside that reach; at the pre-rescore scores the difference was (−1, 0, 0, +5) with exact simplex share 5/6 = 83.33% (sampled 83.4%). What would genuinely close the pair is Obscura shipping post-quantum.
NymVPN. The rescore bought a new adjacency in the other direction: NymVPN, at 7.50, now sits 0.20 above URnetwork's 7.30, closer than Obscura is below it. At the published weights NymVPN leads. Across sampled weightings URnetwork leads 67.0%: the difference is (−1, −1, −1, +7), so the order turns entirely on whether the post-quantum weight exceeds an eighth, and the exact share is (7/8)³ = 343/512. Under score noise NymVPN keeps its lead in 69.4% of draws (URnetwork ahead 30.6%). That pair is printed as adjacent and unsettled, exactly as Obscura and URnetwork were before the rescore.
Per-product rank intervals
5th–95th percentile with the median in parentheses, 22-product field post-rescore, re-implementation run, seed 271828. URnetwork's and Obscura's rows are identical under the primary run. The rescore changed six cells, all in the X-weights column: URnetwork 2–5 (4) → 1–4 (3), Tor 1–4 → 1–5, Obscura 3–10 → 4–10, Mullvad median 3 → 4, IVPN 2–8 → 3–8, ExpressVPN 2–12 → 3–12.
| Product | X rank, weights | X rank, score noise | Y rank, weights | Y rank, score noise |
|---|---|---|---|---|
| Tor | 1–5 (1) | 1–1 (1) | 2–22 (19) | 21–22 (22) |
| NymVPN | 2–9 (3) | 2–3 (2) | 18–21 (20) | 20–21 (20) |
| URnetwork | 1–4 (3) | 2–4 (3) | 1–15 (3) | 4–11 (7) |
| Obscura | 4–10 (6) | 3–5 (4) | 9–19 (15) | 12–19 (16) |
| Mullvad | 1–7 (4) | 4–6 (5) | 4–12 (7) | 4–11 (7) |
| IVPN | 3–8 (5) | 5–7 (6) | 6–16 (11) | 6–13 (10) |
| Apple Private Relay | 5–15 (9) | 6–8 (7) | 13–22 (20) | 13–19 (17) |
| Windscribe | 7–14 (11) | 7–10 (8) | 1–5 (2) | 1–5 (3) |
| PIA | 7–15 (12) | 8–13 (10) | 9–15 (11) | 7–14 (11) |
| ExpressVPN | 3–12 (8) | 8–13 (10) | 4–10 (6) | 4–11 (7) |
| Proton VPN | 9–16 (14) | 9–15 (12) | 1–5 (3) | 1–6 (3) |
| NordVPN | 6–13 (10) | 9–15 (12) | 4–10 (6) | 4–11 (7) |
| Tailscale | 6–18 (13) | 10–16 (13) | 1–8 (3) | 1–3 (1) |
| Cloudflare WARP | 7–16 (11) | 11–16 (14) | 1–16 (5) | 1–6 (3) |
| Surfshark | 7–16 (12) | 11–17 (15) | 8–13 (9) | 5–12 (8) |
| Orchid | 9–17 (16) | 12–17 (15) | 19–22 (21) | 20–22 (21) |
| Mysterium | 16–18 (18) | 16–20 (18) | 7–20 (18) | 14–19 (17) |
| TunnelBear | 15–19 (17) | 16–20 (18) | 8–16 (13) | 9–16 (13) |
| Sentinel | 17–20 (20) | 17–20 (19) | 9–19 (16) | 15–19 (18) |
| IPVanish | 19–20 (19) | 17–21 (19) | 12–19 (16) | 9–16 (13) |
| Hotspot Shield | 21–21 (21) | 20–21 (21) | 8–18 (12) | 9–16 (13) |
| Hola | 22–22 (22) | 22–22 (22) | 6–20 (17) | 14–19 (17) |
The intervals overstate nothing by accident of the models' symmetry. The ±1 noise model is deliberately simple and is not an empirical error distribution; its lesson is that rank confidence requires uncertainty in the assigned scores, not only a second set of weights.
Did we reproduce the reviewer?
Yes, on the 21-product field the reviewer measured, before Obscura was scored. The reviewer reported (seed 20260807, 200,000 samples, their own code): 55.64% ahead of Mullvad; X score 6.341–7.359 and Y score 5.749–6.652 under noise; rank intervals 2–5, 1–15, 2–3 and 4–11 with Y median 7. Our runs on that field returned the identical rank intervals and medians, score bounds within 0.005, and 55.55% on the crossing; the 21-product record is preserved in review/index/SCORES.md, and the pairwise Mullvad statistics carry over to the 22-product re-run unchanged within sampling error. The crossing had an exact answer to check everyone against: at the scores the reviewer measured, the URnetwork−Mullvad sub-score difference was (+4, +3, −3, −3), and the fraction of the weight simplex on URnetwork's side works out to 109/196 ≈ 55.61%. Every Monte Carlo estimate, ours and the reviewer's, sits within sampling error (±0.22 points at 2σ for this sample size) of the exact value. An outside result we could not reproduce would have been a finding; one we can reproduce is too. The reviewer's figures are the pre-rescore record: URnetwork's default-on rescore later the same day changed its X row, and the current exact crossing value is 71/96, derived and sampled in the tables above. The 109/196 check still reproduces on the pre-rescore scores, which is what the reviewer measured.
A withdrawn claim, recorded
Until 2026-08-07 the methodology page claimed URnetwork's Y placement was weight-robust: that no re-weighting of the quality parts could put its Y below 5.00. The reviewer's verdict, "mathematically true but evidentially empty", is accepted in full. Every URnetwork Y sub-score was assigned ≥ 5, so every convex weighted average of them is ≥ 5 by construction: we published a property of our own inputs as though it were evidence of stability, and the claim would have remained "true" no matter how wrong those inputs were. The same objection retires the equivalent X-axis claim before anyone makes it. URnetwork's X sub-scores are also all ≥ 5, so "no re-weighting takes X below 5.00" is equally guaranteed and equally empty. Neither statement appears in this corpus as evidence again. What can be said from the record is in the tables above: under re-weighting the informative output is rank, not line-crossing, and the informative test of the lines is score uncertainty, under which URnetwork's 5th-percentile Y is 5.747, a statement about a simple noise model, not about the world. What could genuinely move Y below the line is evidence: an outside measurement scoring latency or throughput beneath the scale midpoint. None exists in either direction, and the corpus's no-invented-numbers rule is why the two performance parts sit at the structural-class anchor of 5. The claim is withdrawn rather than deleted because this record's value is that it keeps its own errors.
The placements inside noise
Single-point sensitivity is not Monte Carlo, but it answers the same question at the boundaries, so it is recorded here with the rest.
- Apple Private Relay sits within noise of both lines at once, +0.10 on X and −0.05 on Y, alone at the axes' crossing, so it should be read as on the lines, not as a resident of any quadrant, including the Academic / Research cell its strict signs land in. Its judgement calls cut both ways: scoring E or S one point lower puts X at 4.85 or 4.75, back left of the line, and scoring S at 8 puts it clearly right at 5.45. That sensitivity is why the map annotates it rather than claiming it.
- Obscura sits within noise of the Y line from above (+0.10). Its right-of-the-X-line membership is confident at +1.60 and noise-robust: 0 of 200,000 ±1-noise samples put it left of the line, with a 5th-percentile X of 6.12. It is weight-sensitive, though: 32.6% of uniformly sampled weightings put it left, driven by PQ 0 and V 6 when the verification and post-quantum weights are drawn high. The Y side is a genuine coin flip and is claimed for nothing: Y falls below 5.00 in 36% of ±1-noise samples and 50.0% of sampled weightings, and the single judgement-labeled T part moves it either side (T 5 gives 4.85, T 7 gives 5.35).
- Mysterium (−0.10), Hola (−0.15) and Sentinel (−0.20) sit within noise below the Y line, and their above/below membership is not a confident call; their X deficits are the signal.
- From above, TunnelBear, IPVanish and Hotspot Shield clear the Y line by only 0.45–0.50 on datacenter-class L and T scores; a reader who scored their fleet consistency a point lower would put them on or under it.
- The X line is clear water for everyone else. Nothing other than Apple comes closer to it than 0.40 (Windscribe from the left; IVPN at 0.70 and Obscura at 1.60 from the right), so no single-point re-score moves any other product across the verifiable/trust-based split. Apple is the one product whose side of that line legitimately turns on a single judgement call.
Re-weightings that change the ordering
The fixed weights put structure (S + E = 0.60) above verification (V = 0.30). Re-weight toward verification hard enough, for example E 0.15 / S 0.25 / V 0.50 / PQ 0.10, and Mullvad (7.00) passes URnetwork (6.90) on the axis; push the verification weight higher still and IVPN eventually passes it too. Before the default-on rescore a milder example was enough (V at 0.45), and the whole-space number was a coin flip: across uniformly sampled weightings URnetwork then led Mullvad only 55.55% of the time. After the rescore it leads 74.0%, and the V 0.45 example no longer flips the order (URnetwork 7.00, Mullvad 6.70). The ordering at the top of the axis now rests on the structure sub-scores as well as the weights, but a reader who weighs audits and court evidence at about half the axis still gets Mullvad ahead and is not wrong. The per-product pages already tell that reader to pick Mullvad or IVPN until URnetwork closes its protocol-audit gap.
The same lever pulled the other way is the IVPN worked example: a structure-dominant E 0.35 / S 0.45 / V 0.10 / PQ 0.10 drops IVPN to 4.90, just left of the line, with Mullvad at exactly 5.00 on it. IVPN's membership in the top-right quadrant depends on the verification weight staying substantial. The frozen weights clear it, and the dependence is recorded here anyway.
On the quality axis, Eg at 0.20 is what lifts URnetwork's margin and holds WARP, Private Relay and Tor down. A reader who treats egress identity as niche compresses the Y spread, and URnetwork's Y rank runs 1st to 15th across the sampled weight space. An earlier version of the methodology page closed this analysis by claiming that no re-weighting could take URnetwork's Y below 5.00; that claim is withdrawn, and the section above records why.
What would move this record
Three things would move the results on this page rather than argue with them. Obscura shipping post-quantum would close the one degenerate pair (the head-to-heads above). An outside measurement of latency or throughput would replace the structural anchors the withdrawn-claim section names as the quality axis's soft ground. And any rescore under the methodology page's standing rule reruns both tests over the whole field, the way the Obscura addition and the default-on rescore did.
Reproduction script
The cross-check implementation, on the 22-product field with the post-rescore URnetwork row. It regenerates its row of the runs table in review/index/SCORES.md, the URnetwork rank statistics, and the per-product table above (Python 3, numpy ≥ 2). It asserts every published X and Y total recomputes exactly from the sub-score matrices before sampling, and draw order matters for exact reproduction.
import numpy as np
SEED, N = 271828, 200_000
# name, X parts [X1 e2e, X2 sep, X3 verif, X4 pq], Y parts [Y1 L, Y2 T, Y3 Eg, Y4 In, Y5 A]
# URnetwork X row is the post-rescore [8, 8, 6, 7]; pre-rescore was [7, 8, 6, 5]
PRODUCTS = [
("URnetwork", [8, 8, 6, 7], [5, 5, 8, 8, 7]),
("Tor", [10, 10, 10, 4], [2, 1, 1, 9, 9]),
("NymVPN", [9, 9, 7, 0], [5, 4, 2, 5, 4]),
("Obscura", [8, 8, 6, 0], [5, 6, 5, 5, 4]),
("Mullvad", [3, 5, 9, 8], [7, 8, 5, 5, 4]),
("IVPN", [3, 5, 8, 8], [7, 7, 5, 4, 4]),
("Apple Private Relay", [7, 7, 3, 0], [8, 7, 2, 1, 2]),
("Windscribe", [3, 5, 7, 0], [7, 7, 6, 7, 7]),
("PIA", [3, 3, 8, 0], [7, 7, 4, 5, 4]),
("ExpressVPN", [3, 3, 5, 9], [7, 8, 4, 6, 4]),
("Proton VPN", [3, 3, 7, 0], [7, 8, 4, 6, 8]),
("NordVPN", [3, 3, 5, 6], [7, 8, 4, 6, 4]),
("Tailscale", [6, 2, 5, 0], [9, 9, 5, 4, 6]),
("Cloudflare WARP", [3, 4, 3, 5], [9, 8, 2, 4, 8]),
("Surfshark", [3, 3, 4, 5], [7, 8, 4, 5, 4]),
("Orchid", [3, 5, 3, 0], [6, 3, 3, 3, 3]),
("Mysterium", [3, 3, 3, 0], [6, 4, 6, 3, 4]),
("TunnelBear", [3, 2, 4, 0], [7, 6, 3, 5, 5]),
("Sentinel", [3, 3, 2, 0], [6, 4, 4, 6, 4]),
("IPVanish", [3, 2, 3, 0], [7, 7, 3, 4, 4]),
("Hotspot Shield", [3, 1, 2, 0], [7, 7, 2, 5, 5]),
("Hola", [1, 0, 0, 0], [5, 4, 5, 3, 7]),
]
WX = np.array([0.25, 0.35, 0.30, 0.10])
WY = np.array([0.30, 0.25, 0.20, 0.10, 0.15])
names = [p[0] for p in PRODUCTS]
XS = np.array([p[1] for p in PRODUCTS], float)
YS = np.array([p[2] for p in PRODUCTS], float)
UR, MV = names.index("URnetwork"), names.index("Mullvad")
OB, NY = names.index("Obscura"), names.index("NymVPN")
# sanity: every published total must recompute exactly (SCORES.md tables, product order above)
EXP_X = [7.30, 9.40, 7.50, 6.60, 6.00, 5.70, 5.10, 4.60, 4.20, 4.20, 3.90,
3.90, 3.70, 3.55, 3.50, 3.40, 2.70, 2.65, 2.40, 2.35, 1.70, 0.25]
EXP_Y = [6.20, 3.30, 4.00, 5.10, 6.20, 5.85, 4.95, 6.80, 5.75, 6.10, 6.70,
6.10, 7.25, 6.70, 6.00, 3.90, 4.90, 5.45, 4.80, 5.45, 5.50, 4.85]
assert np.allclose(XS @ WX, EXP_X) and np.allclose(YS @ WY, EXP_Y)
def ranks(scores): # rank = 1 + #{strictly greater}
return 1 + (scores[:, None, :] > scores[:, :, None]).sum(axis=2)
def q(a, p):
return np.quantile(a, p, method="nearest")
def sim_A(rng, subs, w, chunk=50_000):
R, S = [], []
for start in range(0, N, chunk):
m = min(chunk, N - start)
sc = rng.dirichlet(np.ones(len(w)), size=m) @ subs.T
R.append(ranks(sc)); S.append(sc)
return np.vstack(R), np.vstack(S)
def sim_B(rng, subs, w, chunk=20_000):
R, S = [], []
for start in range(0, N, chunk):
m = min(chunk, N - start)
noisy = np.clip(subs[None] + rng.uniform(-1, 1, (m,) + subs.shape), 0, 10)
sc = noisy @ w
R.append(ranks(sc)); S.append(sc)
return np.vstack(R), np.vstack(S)
rng = np.random.default_rng(SEED) # draw order matters for exact reproduction:
rAX, sAX = sim_A(rng, XS, WX) # 1) test 1, X axis
rAY, sAY = sim_A(rng, YS, WY) # 2) test 1, Y axis
rBX, sBX = sim_B(rng, XS, WX) # 3) test 2, X axis
rBY, sBY = sim_B(rng, YS, WY) # 4) test 2, Y axis
print("exact P(UR>MV | test 1) = 71/96 =", 71 / 96) # pre-rescore: 109/196
print("test 1: P(UR>MV on X) =", (sAX[:, UR] > sAX[:, MV]).mean())
print("test 2: P(UR>MV on X) =", (sBX[:, UR] > sBX[:, MV]).mean())
print("exact P(UR>Obscura | test 1) = 1 (pure X4 difference; pre-rescore: 5/6)")
print("test 1: P(UR>Obscura on X) =", (sAX[:, UR] > sAX[:, OB]).mean())
print("test 2: P(UR>Obscura on X) =", (sBX[:, UR] > sBX[:, OB]).mean())
print("exact P(UR>NymVPN | test 1) = 343/512 =", 343 / 512)
print("test 1: P(UR>NymVPN on X) =", (sAX[:, UR] > sAX[:, NY]).mean())
print("test 2: P(UR>NymVPN on X) =", (sBX[:, UR] > sBX[:, NY]).mean())
print("test 2: UR X score 5-95%:", q(sBX[:, UR], 0.05), q(sBX[:, UR], 0.95))
print("test 2: UR Y score 5-95%:", q(sBY[:, UR], 0.05), q(sBY[:, UR], 0.95))
print("test 2: UR Y < 5.00 in", int((sBY[:, UR] < 5).sum()), "of", N)
for label, r in (("AX", rAX), ("BX", rBX), ("AY", rAY), ("BY", rBY)):
print(label, "UR rank p5/med/p95:", int(q(r[:, UR], .05)), int(q(r[:, UR], .5)), int(q(r[:, UR], .95)))
print("| Product | X rank, weights | X rank, score noise | Y rank, weights | Y rank, score noise |")
print("|---|---|---|---|---|")
f = lambda r, j: f"{int(q(r[:, j], .05))}–{int(q(r[:, j], .95))} ({int(q(r[:, j], .5))})"
for j in np.argsort(-(XS @ WX), kind="stable"):
print(f"| {names[j]} | {f(rAX, j)} | {f(rBX, j)} | {f(rAY, j)} | {f(rBY, j)} |")