# Comparison sensitivity

This page records the two Monte Carlo stability tests behind the [comparison map](/docs/comparison): resampling the axis weights across their whole space with the sub-scores frozen, and perturbing every sub-score by up to a point at the published weights. It holds the designs and seeds, the rank intervals for all twenty-two products, the head-to-head crossing statistics, one withdrawn claim, and a script that reproduces the record.

The first version of the [methodology page](/docs/comparison-methodology)
tested exactly one alternative weighting and reported one crossover. The
peer review did what we should have done: a Monte Carlo across the whole
weight space and across sub-score uncertainty. We adopted the design and
re-ran it with our own code rather than citing the reviewer's numbers.
This page publishes ours; the full record of runs and the agreement check
against the reviewer's figures is in `review/index/SCORES.md`, and every
sub-score the tests draw on is on the
[scores page](/docs/comparison-scores).

## The two tests

Each test runs N = 200,000 samples per axis:

1. **Weights sampled, scores frozen.** Non-negative weights drawn
   uniformly over each axis's weight simplex (normalized Exp(1) draws);
   every published sub-score held fixed. This asks: how much of the
   ordering is our weighting judgement?
2. **Weights frozen, scores perturbed.** The published weights kept;
   independent uniform ±1 noise added to every sub-score of every
   product, clipped to 0–10. This asks: how much depends on integer
   judgement calls being exactly right?

A rank here is 1 + the number of products scoring strictly higher.
Primary run: seed 20260807. Independent re-implementation (different code
and generator, seed 271828). **Recomputed 2026-08-07 over the 22-product
field when Obscura was scored**: rank is field-relative, so the intervals
below were re-run with Obscura's row in rather than left stale, and the
working record in `review/index/SCORES.md` preserves the 21-product runs
as the reviewer-reproduction record. **Recomputed again the same day when
URnetwork was rescored on the default-on ruling** (the methodology
ledger): the tables below are the post-rescore record, re-run under both
implementations and both seeds, and the pre-rescore runs stay in the
working record's runs table. On the 22-product field the two
implementations agree on every interval and median except five
weights-test cells at near-tied adjacent pairs (0.05 apart in point
score), noted in the working record; the same five cells differ before
and after the rescore. Every score-noise interval, the kind the overview
page publishes, is identical across both runs, as are URnetwork's and
Obscura's rows in full.

## URnetwork's results

Obscura's entry moved the X rank statistics: weights median 3→4, noise
interval 2–3 → 2–4. The default-on rescore moved them again: weights
median 4→3, weights interval 2–5 → 1–4. No noise-test rank statistic and
no Y statistic moved either time.

| Statistic | Test 1, weights sampled | Test 2, ±1 score noise |
|---|---|---|
| X rank, median (5th–95th) | 3 (1–4) | 3 (2–4) |
| X score, 5th–95th | (weights vary) | 6.793–7.813 |
| Ahead of Mullvad on X | **74.0%** (exactly 71/96) | **99.9%** |
| Ahead of Obscura on X | **100%** (degenerate; see below) | **93.9%** |
| Ahead of NymVPN on X | **67.0%** (exactly 343/512) | **30.6%** |
| Y rank, median (5th–95th) | 3 (1–15) | **7 (4–11)** |
| Y score, 5th–95th | (weights vary) | 5.747–6.653 |
| Y below 5.00 | never, by construction (see the withdrawn claim) | 0 of 200,000 (about 4.4σ under this noise model: rare, not impossible) |

**The two axes do not carry equal confidence.** URnetwork's X rank stays
inside 1–4 across every weighting sampled: near the top of the X axis,
which judgement you make decides the order, but the index itself is
comparatively weight-stable. Its Y rank runs 1–15 across the same weight
space, and even at the published weights the ±1-noise median is 7th, not
the tied-5th its point score suggests. The Y index is far more
weight-sensitive than X, and every Y placement on the map should be read
with correspondingly less confidence.

## The head-to-heads

**Mullvad.** The two Mullvad numbers remain the most informative sentence
available about this chart, and the default-on rescore changed what they
say. Before it, URnetwork's X lead over Mullvad was 55.55% under weight
uncertainty, a coin flip that depended almost entirely on ranking
structure (separation plus end-to-end, 0.60 combined) above verification
(0.30). After the rescore the sub-score difference is (+5, +3, −3, −1):
URnetwork leads both structure parts, Mullvad leads verification and,
narrowly, post-quantum, and the same weight sampling gives URnetwork the
lead 74.0% of the time (the exact simplex volume is 71/96 ≈ 73.96%).
Under ±1 score noise at the published weights it is 99.9%. The lead is no
longer a coin flip, but the residual weight sensitivity survives in
scoped form: a reader who weights verification at about half the axis
still gets Mullvad ahead, and that remains a reasonable reading, not a
mistake. The worked example is in the re-weighting section below.

**Obscura.** The comparison inverted with the rescore. The two now tie on
structure (E+S 16–16, both content legs default-on, Obscura's by
construction) and tie on V, so the entire 0.70 gap is the post-quantum
part Obscura lacks. A pure single-part difference, (0, 0, 0, +7), makes
the weights test degenerate: any weighting with a nonzero post-quantum
weight preserves the order, so "100% of sampled weightings" describes the
difference's shape, not robustness. The informative test is score noise,
under which URnetwork leads 93.9% of the time: outside a single point's
reach, since the largest one-point, one-part swing is 0.35 against a 0.70
gap. The pre-rescore statistic was 69.5%, which was not outside that
reach; at the pre-rescore scores the difference was (−1, 0, 0, +5) with
exact simplex share 5/6 = 83.33% (sampled 83.4%). What would genuinely
close the pair is Obscura shipping post-quantum.

**NymVPN.** The rescore bought a new adjacency in the other direction:
NymVPN, at 7.50, now sits 0.20 above URnetwork's 7.30, closer than
Obscura is below it. At the published weights NymVPN leads. Across
sampled weightings URnetwork leads 67.0%: the difference is
(−1, −1, −1, +7), so the order turns entirely on whether the post-quantum
weight exceeds an eighth, and the exact share is (7/8)³ = 343/512.
Under score noise NymVPN keeps its lead in 69.4% of draws (URnetwork
ahead 30.6%). That pair is printed as adjacent and unsettled, exactly as
Obscura and URnetwork were before the rescore.

## Per-product rank intervals

5th–95th percentile with the median in parentheses, 22-product field
post-rescore, re-implementation run, seed 271828. URnetwork's and
Obscura's rows are identical under the primary run. The rescore changed
six cells, all in the X-weights column: URnetwork 2–5 (4) → 1–4 (3), Tor
1–4 → 1–5, Obscura 3–10 → 4–10, Mullvad median 3 → 4, IVPN 2–8 → 3–8,
ExpressVPN 2–12 → 3–12.

| Product | X rank, weights | X rank, score noise | Y rank, weights | Y rank, score noise |
|---|---|---|---|---|
| Tor | 1–5 (1) | 1–1 (1) | 2–22 (19) | 21–22 (22) |
| NymVPN | 2–9 (3) | 2–3 (2) | 18–21 (20) | 20–21 (20) |
| URnetwork | 1–4 (3) | 2–4 (3) | 1–15 (3) | 4–11 (7) |
| Obscura | 4–10 (6) | 3–5 (4) | 9–19 (15) | 12–19 (16) |
| Mullvad | 1–7 (4) | 4–6 (5) | 4–12 (7) | 4–11 (7) |
| IVPN | 3–8 (5) | 5–7 (6) | 6–16 (11) | 6–13 (10) |
| Apple Private Relay | 5–15 (9) | 6–8 (7) | 13–22 (20) | 13–19 (17) |
| Windscribe | 7–14 (11) | 7–10 (8) | 1–5 (2) | 1–5 (3) |
| PIA | 7–15 (12) | 8–13 (10) | 9–15 (11) | 7–14 (11) |
| ExpressVPN | 3–12 (8) | 8–13 (10) | 4–10 (6) | 4–11 (7) |
| Proton VPN | 9–16 (14) | 9–15 (12) | 1–5 (3) | 1–6 (3) |
| NordVPN | 6–13 (10) | 9–15 (12) | 4–10 (6) | 4–11 (7) |
| Tailscale | 6–18 (13) | 10–16 (13) | 1–8 (3) | 1–3 (1) |
| Cloudflare WARP | 7–16 (11) | 11–16 (14) | 1–16 (5) | 1–6 (3) |
| Surfshark | 7–16 (12) | 11–17 (15) | 8–13 (9) | 5–12 (8) |
| Orchid | 9–17 (16) | 12–17 (15) | 19–22 (21) | 20–22 (21) |
| Mysterium | 16–18 (18) | 16–20 (18) | 7–20 (18) | 14–19 (17) |
| TunnelBear | 15–19 (17) | 16–20 (18) | 8–16 (13) | 9–16 (13) |
| Sentinel | 17–20 (20) | 17–20 (19) | 9–19 (16) | 15–19 (18) |
| IPVanish | 19–20 (19) | 17–21 (19) | 12–19 (16) | 9–16 (13) |
| Hotspot Shield | 21–21 (21) | 20–21 (21) | 8–18 (12) | 9–16 (13) |
| Hola | 22–22 (22) | 22–22 (22) | 6–20 (17) | 14–19 (17) |

The intervals overstate nothing by accident of the models' symmetry. The
±1 noise model is deliberately simple and is not an empirical error
distribution; its lesson is that rank confidence requires uncertainty in
the assigned scores, not only a second set of weights.

## Did we reproduce the reviewer?

Yes, on the 21-product field the reviewer measured, before Obscura was
scored. The reviewer reported (seed 20260807, 200,000 samples, their own
code): 55.64% ahead of Mullvad; X score 6.341–7.359 and Y score
5.749–6.652 under noise; rank intervals 2–5, 1–15, 2–3 and 4–11 with Y
median 7. Our runs on that field returned the identical rank intervals
and medians, score bounds within 0.005, and 55.55% on the crossing; the
21-product record is preserved in `review/index/SCORES.md`, and the
pairwise Mullvad statistics carry over to the 22-product re-run unchanged
within sampling error. The crossing had an exact answer to check everyone
against: at the scores the reviewer measured, the URnetwork−Mullvad
sub-score difference was (+4, +3, −3, −3), and the fraction of the weight
simplex on URnetwork's side works out to 109/196 ≈ 55.61%. Every Monte
Carlo estimate, ours and the reviewer's, sits within sampling error
(±0.22 points at 2σ for this sample size) of the exact value. An outside
result we could not reproduce would have been a finding; one we can
reproduce is too. The reviewer's figures are the pre-rescore record:
URnetwork's default-on rescore later the same day changed its X row, and
the current exact crossing value is 71/96, derived and sampled in the
tables above. The 109/196 check still reproduces on the pre-rescore
scores, which is what the reviewer measured.

## A withdrawn claim, recorded

Until 2026-08-07 the methodology page claimed URnetwork's Y placement was
weight-robust: that no re-weighting of the quality parts could put its Y
below 5.00. The reviewer's verdict, **"mathematically true but
evidentially empty"**, is accepted in full. Every URnetwork Y sub-score
was assigned ≥ 5, so every convex weighted average of them is ≥ 5 **by
construction**: we published a property of our own inputs as though it
were evidence of stability, and the claim would have remained "true" no
matter how wrong those inputs were. The same objection retires the
equivalent X-axis claim before anyone makes it. URnetwork's X sub-scores
are also all ≥ 5, so "no re-weighting takes X below 5.00" is equally
guaranteed and equally empty. Neither statement appears in this corpus as
evidence again. What can be said from the record is in the tables above:
under re-weighting the informative output is rank, not line-crossing, and
the informative test of the lines is score uncertainty, under which
URnetwork's 5th-percentile Y is 5.747, a statement about a simple noise
model, not about the world. What could genuinely move Y below the line is
evidence: an outside measurement scoring latency or throughput beneath
the scale midpoint. None exists in either direction, and the corpus's
no-invented-numbers rule is why the two performance parts sit at the
structural-class anchor of 5. The claim is withdrawn rather than deleted
because this record's value is that it keeps its own errors.

## The placements inside noise

Single-point sensitivity is not Monte Carlo, but it answers the same
question at the boundaries, so it is recorded here with the rest.

- **Apple Private Relay** sits within noise of both lines at once, +0.10
  on X and −0.05 on Y, alone at the axes' crossing, so it should be read
  as on the lines, not as a resident of any quadrant, including the
  Academic / Research cell its strict signs land in. Its judgement calls
  cut both ways: scoring E or S one point lower puts X at 4.85 or 4.75,
  back left of the line, and scoring S at 8 puts it clearly right at
  5.45. That sensitivity is why the map annotates it rather than
  claiming it.
- **Obscura** sits within noise of the Y line from above (+0.10). Its
  right-of-the-X-line membership is confident at +1.60 and noise-robust:
  0 of 200,000 ±1-noise samples put it left of the line, with a
  5th-percentile X of 6.12. It is weight-sensitive, though: 32.6% of
  uniformly sampled weightings put it left, driven by PQ 0 and V 6 when
  the verification and post-quantum weights are drawn high. The Y side
  is a genuine coin flip and is claimed for nothing: Y falls below 5.00
  in 36% of ±1-noise samples and 50.0% of sampled weightings, and the
  single judgement-labeled T part moves it either side (T 5 gives 4.85,
  T 7 gives 5.35).
- **Mysterium** (−0.10), **Hola** (−0.15) and **Sentinel** (−0.20) sit
  within noise below the Y line, and their above/below membership is not
  a confident call; their X deficits are the signal.
- From above, **TunnelBear**, **IPVanish** and **Hotspot Shield** clear
  the Y line by only 0.45–0.50 on datacenter-class L and T scores; a
  reader who scored their fleet consistency a point lower would put them
  on or under it.
- The X line is clear water for everyone else. Nothing other than Apple
  comes closer to it than 0.40 (Windscribe from the left; IVPN at 0.70
  and Obscura at 1.60 from the right), so no single-point re-score moves
  any other product across the verifiable/trust-based split. Apple is
  the one product whose side of that line legitimately turns on a single
  judgement call.

## Re-weightings that change the ordering

The fixed weights put structure (S + E = 0.60) above verification
(V = 0.30). Re-weight toward verification hard enough, for example
E 0.15 / S 0.25 / V 0.50 / PQ 0.10, and **Mullvad (7.00) passes
URnetwork (6.90)** on the axis; push the verification weight higher still
and IVPN eventually passes it too. Before the default-on rescore a milder
example was enough (V at 0.45), and the whole-space number was a coin
flip: across uniformly sampled weightings URnetwork then led Mullvad only
55.55% of the time. After the rescore it leads 74.0%, and the V 0.45
example no longer flips the order (URnetwork 7.00, Mullvad 6.70). The
ordering at the top of the axis now rests on the structure sub-scores as
well as the weights, but a reader who weighs audits and court evidence at
about half the axis still gets Mullvad ahead and is not wrong. The
per-product pages already tell that reader to pick Mullvad or IVPN until
URnetwork closes its protocol-audit gap.

The same lever pulled the other way is the IVPN worked example: a
structure-dominant E 0.35 / S 0.45 / V 0.10 / PQ 0.10 drops IVPN to 4.90,
just left of the line, with Mullvad at exactly 5.00 on it. IVPN's
membership in the top-right quadrant depends on the verification weight
staying substantial. The frozen weights clear it, and the dependence is
recorded here anyway.

On the quality axis, Eg at 0.20 is what lifts URnetwork's margin and
holds WARP, Private Relay and Tor down. A reader who treats egress
identity as niche compresses the Y spread, and URnetwork's Y rank runs
1st to 15th across the sampled weight space. An earlier version of the
methodology page closed this analysis by claiming that no re-weighting
could take URnetwork's Y below 5.00; that claim is withdrawn, and the
section above records why.

## What would move this record

Three things would move the results on this page rather than argue with
them. Obscura shipping post-quantum would close the one degenerate pair
(the head-to-heads above). An outside measurement of latency or
throughput would replace the structural anchors the withdrawn-claim
section names as the quality axis's soft ground. And any rescore under
the [methodology page](/docs/comparison-methodology)'s standing rule
reruns both tests over the whole field, the way the Obscura addition and
the default-on rescore did.

## Reproduction script

The cross-check implementation, on the 22-product field with the
post-rescore URnetwork row. It regenerates its row of the runs table in
`review/index/SCORES.md`, the URnetwork rank statistics, and the
per-product table above (Python 3, numpy ≥ 2). It asserts every published
X and Y total recomputes exactly from the sub-score matrices before
sampling, and draw order matters for exact reproduction.

```python
import numpy as np

SEED, N = 271828, 200_000

# name, X parts [X1 e2e, X2 sep, X3 verif, X4 pq], Y parts [Y1 L, Y2 T, Y3 Eg, Y4 In, Y5 A]
# URnetwork X row is the post-rescore [8, 8, 6, 7]; pre-rescore was [7, 8, 6, 5]
PRODUCTS = [
    ("URnetwork",           [8, 8, 6, 7],    [5, 5, 8, 8, 7]),
    ("Tor",                 [10, 10, 10, 4], [2, 1, 1, 9, 9]),
    ("NymVPN",              [9, 9, 7, 0],    [5, 4, 2, 5, 4]),
    ("Obscura",             [8, 8, 6, 0],    [5, 6, 5, 5, 4]),
    ("Mullvad",             [3, 5, 9, 8],    [7, 8, 5, 5, 4]),
    ("IVPN",                [3, 5, 8, 8],    [7, 7, 5, 4, 4]),
    ("Apple Private Relay", [7, 7, 3, 0],    [8, 7, 2, 1, 2]),
    ("Windscribe",          [3, 5, 7, 0],    [7, 7, 6, 7, 7]),
    ("PIA",                 [3, 3, 8, 0],    [7, 7, 4, 5, 4]),
    ("ExpressVPN",          [3, 3, 5, 9],    [7, 8, 4, 6, 4]),
    ("Proton VPN",          [3, 3, 7, 0],    [7, 8, 4, 6, 8]),
    ("NordVPN",             [3, 3, 5, 6],    [7, 8, 4, 6, 4]),
    ("Tailscale",           [6, 2, 5, 0],    [9, 9, 5, 4, 6]),
    ("Cloudflare WARP",     [3, 4, 3, 5],    [9, 8, 2, 4, 8]),
    ("Surfshark",           [3, 3, 4, 5],    [7, 8, 4, 5, 4]),
    ("Orchid",              [3, 5, 3, 0],    [6, 3, 3, 3, 3]),
    ("Mysterium",           [3, 3, 3, 0],    [6, 4, 6, 3, 4]),
    ("TunnelBear",          [3, 2, 4, 0],    [7, 6, 3, 5, 5]),
    ("Sentinel",            [3, 3, 2, 0],    [6, 4, 4, 6, 4]),
    ("IPVanish",            [3, 2, 3, 0],    [7, 7, 3, 4, 4]),
    ("Hotspot Shield",      [3, 1, 2, 0],    [7, 7, 2, 5, 5]),
    ("Hola",                [1, 0, 0, 0],    [5, 4, 5, 3, 7]),
]
WX = np.array([0.25, 0.35, 0.30, 0.10])
WY = np.array([0.30, 0.25, 0.20, 0.10, 0.15])
names = [p[0] for p in PRODUCTS]
XS = np.array([p[1] for p in PRODUCTS], float)
YS = np.array([p[2] for p in PRODUCTS], float)
UR, MV = names.index("URnetwork"), names.index("Mullvad")
OB, NY = names.index("Obscura"), names.index("NymVPN")

# sanity: every published total must recompute exactly (SCORES.md tables, product order above)
EXP_X = [7.30, 9.40, 7.50, 6.60, 6.00, 5.70, 5.10, 4.60, 4.20, 4.20, 3.90,
         3.90, 3.70, 3.55, 3.50, 3.40, 2.70, 2.65, 2.40, 2.35, 1.70, 0.25]
EXP_Y = [6.20, 3.30, 4.00, 5.10, 6.20, 5.85, 4.95, 6.80, 5.75, 6.10, 6.70,
         6.10, 7.25, 6.70, 6.00, 3.90, 4.90, 5.45, 4.80, 5.45, 5.50, 4.85]
assert np.allclose(XS @ WX, EXP_X) and np.allclose(YS @ WY, EXP_Y)

def ranks(scores):  # rank = 1 + #{strictly greater}
    return 1 + (scores[:, None, :] > scores[:, :, None]).sum(axis=2)

def q(a, p):
    return np.quantile(a, p, method="nearest")

def sim_A(rng, subs, w, chunk=50_000):
    R, S = [], []
    for start in range(0, N, chunk):
        m = min(chunk, N - start)
        sc = rng.dirichlet(np.ones(len(w)), size=m) @ subs.T
        R.append(ranks(sc)); S.append(sc)
    return np.vstack(R), np.vstack(S)

def sim_B(rng, subs, w, chunk=20_000):
    R, S = [], []
    for start in range(0, N, chunk):
        m = min(chunk, N - start)
        noisy = np.clip(subs[None] + rng.uniform(-1, 1, (m,) + subs.shape), 0, 10)
        sc = noisy @ w
        R.append(ranks(sc)); S.append(sc)
    return np.vstack(R), np.vstack(S)

rng = np.random.default_rng(SEED)   # draw order matters for exact reproduction:
rAX, sAX = sim_A(rng, XS, WX)       # 1) test 1, X axis
rAY, sAY = sim_A(rng, YS, WY)       # 2) test 1, Y axis
rBX, sBX = sim_B(rng, XS, WX)       # 3) test 2, X axis
rBY, sBY = sim_B(rng, YS, WY)       # 4) test 2, Y axis

print("exact P(UR>MV | test 1) = 71/96 =", 71 / 96)   # pre-rescore: 109/196
print("test 1: P(UR>MV on X) =", (sAX[:, UR] > sAX[:, MV]).mean())
print("test 2: P(UR>MV on X) =", (sBX[:, UR] > sBX[:, MV]).mean())
print("exact P(UR>Obscura | test 1) = 1 (pure X4 difference; pre-rescore: 5/6)")
print("test 1: P(UR>Obscura on X) =", (sAX[:, UR] > sAX[:, OB]).mean())
print("test 2: P(UR>Obscura on X) =", (sBX[:, UR] > sBX[:, OB]).mean())
print("exact P(UR>NymVPN | test 1) = 343/512 =", 343 / 512)
print("test 1: P(UR>NymVPN on X) =", (sAX[:, UR] > sAX[:, NY]).mean())
print("test 2: P(UR>NymVPN on X) =", (sBX[:, UR] > sBX[:, NY]).mean())
print("test 2: UR X score 5-95%:", q(sBX[:, UR], 0.05), q(sBX[:, UR], 0.95))
print("test 2: UR Y score 5-95%:", q(sBY[:, UR], 0.05), q(sBY[:, UR], 0.95))
print("test 2: UR Y < 5.00 in", int((sBY[:, UR] < 5).sum()), "of", N)
for label, r in (("AX", rAX), ("BX", rBX), ("AY", rAY), ("BY", rBY)):
    print(label, "UR rank p5/med/p95:", int(q(r[:, UR], .05)), int(q(r[:, UR], .5)), int(q(r[:, UR], .95)))
print("| Product | X rank, weights | X rank, score noise | Y rank, weights | Y rank, score noise |")
print("|---|---|---|---|---|")
f = lambda r, j: f"{int(q(r[:, j], .05))}–{int(q(r[:, j], .95))} ({int(q(r[:, j], .5))})"
for j in np.argsort(-(XS @ WX), kind="stable"):
    print(f"| {names[j]} | {f(rAX, j)} | {f(rBX, j)} | {f(rAY, j)} | {f(rBY, j)} |")
```

