How the comparison map is scored
The comparison map places twenty-two networks, URnetwork and every product in the twenty-one head-to-head pages, on two 0–10 indexes: verifiable anonymization across, connection quality up. This page is the method: what the two indexes measure, the rule for which products are on the map at all, the weights and why they are what they are, how the anchor ladders work, the dated record of every change since the rubric was frozen, and the limits a reader should hold against all of it. The intent is that a reader who doubts any coordinate can recompute it, and, where a score is a judgement call rather than a sourced fact, see it labeled as one.
The working-out lives in two appendices. The scores page holds the complete matrix: all 198 sub-scores, the anchor ladders they were assigned against, per-product reasoning for every number, and the disclosed rescores. The sensitivity page holds the Monte Carlo record: rank intervals under weight and score uncertainty, the head-to-head crossing statistics, a withdrawn claim, and a reproduction script. The audit trail behind all three sits in the documentation repository under review/index/: METHOD.md (the rubric, frozen before any product was scored), SCORES.md (the recomputable matrix), and NOTES.md (per-score reasoning with citations). The published pages and that record are kept in agreement; if they ever disagree, one of them has a bug, and the end of this page says where to report it.
The change ledger
A scored comparison that never changes is stale, and one that changes silently is worthless. This is the dated ledger of every change to the method or to its published claims since the rubric was frozen on 2026-08-07, each marked with the direction it cuts, because a reader auditing a self-scored index should be able to see at a glance whether the corrections keep landing in the scorer's favour.
| Date | Change | Direction |
|---|---|---|
| 2026-08-07 | IVPN's latency sub-score normalized 6→7 during scoring, to match every other single-hop datacenter VPN (disclosure 1 under "The freeze") | for IVPN; neutral to URnetwork |
| 2026-08-07 | Origin corrected by the product owner from field-centroid to the fixed scale center (disclosure 2) | neutral; moves the lines, not the dots |
| 2026-08-07 | Apple Private Relay rescored after independent peer review: E 5→7, S 6→7, X 4.25→5.10 (disclosure 3) | against our interest; corrects an error that had run in URnetwork's favour |
| 2026-08-07 | The claim that no re-weighting can put URnetwork's Y below 5.00 withdrawn as circular (recorded on the sensitivity page) | against our interest; a published stability claim about ourselves, retracted |
| 2026-08-07 | Monte Carlo rank intervals published; the single-alternative-weighting sensitivity check replaced | against our interest; prints a near-coin-flip X lead and an unstable Y rank |
| 2026-08-07 | URnetwork's S reasoning corrected: the stale "sign-up IP retained raw 180 days" penalty removed after verification against current server source; the score re-derived from the ladder and unchanged at 8 | in our favour by intent, no score effect; flagged because it is the flattering direction |
| 2026-08-07 | Comparator-selection rule published; two qualifying products we do not yet compare named | against our interest; records our own omissions |
| 2026-08-07 | The corpus's repeated "URnetwork has no independent audit" corrected: two third-party assessments exist, a third-party penetration test of the web application and API (April–May 2025) and a passed Leviathan Security Group MASA AL2 assessment of the Android app (May 2025). Neither covers the protocol, the connect engine or the operator's server code. V re-derived from the ladder and unchanged at 6 | in our favour; the documentation had understated us in about thirty places, and the product owner supplied the reports. Flagged because it is the flattering direction, and priced at zero for that reason |
| 2026-08-07 | Obscura scored and added, the twenty-first compared product, entering under selection bar (d) as the closest architectural competitor, the omission peer review used to prove the old ad-hoc selection failed. Scored strictly from the frozen ladders; no other product's sub-score touched. It lands at (6.60, 5.10): fourth on the X axis, 0.25 behind URnetwork as then scored, and ON the Y line (+0.10, inside noise). All rank intervals recomputed over the 22-product field the same day | against our interest; seats a competitor beside us in the top-right cell that outscores URnetwork on the structure parts as then scored (16 vs 15 across E+S; its always-on entry blindness beats our then-opt-in on E), ties us on V, and trails only on post-quantum; under ±1 score noise URnetwork keeps the X lead just 69.5% of the time, and all of that is printed rather than smoothed |
| 2026-08-07 | URnetwork rescored: the sealed session is on by default at launch. A product-owner ruling (recorded in the working record's process crib and verify/BEFORELAUNCH.md item 7) reversed the opt-in framing three of URnetwork's four X sub-scores were derived against. Re-derived from the unchanged frozen ladders: E 7→8 (default-on structural operator-blindness takes the band top, the same reading that placed Obscura's always-on 8; the browser extension still has no sealed session, and the seal's two standing pre-launch integrity defects, silent fail-open and an operator-substitutable provider key, are printed beside the score as the 9–10 bar, not waived), PQ 5→7 (the default-on-flagship rung, at its floor), S re-derived and unchanged at 8 (the 9–10 rung still fails on multi-party-by-construction and no-identifying-account), V untouched at 6. X 6.85→7.30; Y unaffected. Rank intervals and every head-to-head sensitivity recomputed over the field the same day | in our favour; the third in-our-favour correction processed on 2026-08-07 and the first to move a number. The earlier two (the stale raw-IP retention penalty struck from S, the audit-record correction on V) went through the same ladder re-derivation and correctly yielded nothing, which is the evidence the discipline bites rather than decorates. This one moves because it is a genuine capability change ruled by the product owner, not a stale penalty, and the integrity caveats that temper it are printed with the score |
What has not changed: no weight, no anchor, and no sub-score other than the two disclosed rescores, Apple Private Relay's and URnetwork's own when its sealed session went default-on.
The standing verdict this ledger operates under is the peer review's: the index is "transparent but not validated". Transparency is what this page has always sold: every weight, anchor and sub-score published so a disagreeing reader can recompute. And the first serious reader who did recompute found a circular robustness claim and a classification-material factual error. Transparency is not validity. The index remains a single-scorer exercise with a declared conflict of interest, aggregating ordinal judgements as if they were interval measurements, validated against no outside measurement of URnetwork. No independent audit covers URnetwork's protocol, its connect engine, or the operator's server code, which is the layer the whole argument rests on. (Two third-party assessments do exist, of other surfaces: a third-party penetration test of the web application and API in April–May 2025, and a Leviathan Security Group assessment of the Android app under Google's App Defense Alliance MASA programme at assurance level AL2, completed May 2025, which it passed and which Leviathan itself scopes as not a holistic security evaluation. Neither examined logging, retention, or the data path.) What would raise the index's standing is known and none of it is writing: independent scorers with published inter-rater agreement, an outside measurement campaign, an audit of the protocol and the server. Until then, read the map as our argument made checkable, not as a measurement.
What the two indexes measure
X, verifiable anonymization. The first question of the whole corpus: what can the service see about you, and how would you know? The axis rewards two things, deliberately in this order. Structure: a design in which no single party can hold your identity and your activity together, or read your traffic in transit. And verification: the evidence that lets an outsider check the claim, meaning open source, independent audits, court and raid outcomes, and outside measurement. The quadrant names on the map state the same distinction: right of the line the see-nothing claim can be checked; left of it, however good the operator's conduct, the claim rests on trust.
Y, connection quality. The second question: is the connection good enough to live on? Latency and throughput class, whether the exit actually delivers the location it claims, how many ways in exist when networks are hostile, and whether you can get and keep the service at all.
Why these parts and not others. The X parts are the four properties that bear directly on "who can see, and how would you know"; anything else would dilute the axis with proxies. Jurisdiction and ownership, the usual VPN-marketing shorthand, are excluded on purpose: they are priors about future behavior, not verifiable properties of the system, and they resist 0–10 anchoring (the per-product pages carry them as prose, with dates and owners named). RAM-only fleets, kill switches, and app polish are likewise real but belong to product reviews, not to an anonymization index. On Y, price and device limits are excluded because they are purchasing factors, not properties of the connection; torrenting policy and port forwarding appear only insofar as they are egress restrictions, and the first of those costs URnetwork part of a point on the scores page. Two axes cannot hold everything, and these two do not try.
Which products are on the map: the selection rule
The original twenty comparators were chosen ad hoc, market share plus a wish to cover every architecture class, and no inclusion rule was published. Independent peer review called that out and supplied the proof of why it matters: Obscura, the closest architectural competitor to URnetwork, was missing until a reviewer named it. "We noticed it" is not a method. The rule below exists so the next omission is a rule-failure anyone can point to, not a judgement nobody can inspect.
A product qualifies if it meets all three conditions:
- It is an operated intermediary service. A shipping service an ordinary consumer can obtain today, whose core function is to place an operated intermediary between the user and the destination for privacy, anonymity, or location change: commercial VPNs, platform relays, onion and mix networks, dVPNs, proxy networks. Self-hosted software where the user is the operator (a bare WireGuard server, Outline) is out of scope. There is no second party whose visibility is in question, which is the question both axes score.
- It clears at least one of four reach-or-relevance bars: - (a) adoption: top tier of consumer VPN adoption by installs or market share in current, dated public market data; - (b) archetype flagship: the flagship or most-adopted implementation of a distinct architecture class (onion routing: Tor; mixnet: NymVPN; platform relay: Apple Private Relay; edge relay: Cloudflare WARP; mesh overlay: Tailscale; dVPN marketplace: Orchid, Mysterium, Sentinel; residential P2P proxy: Hola); - (c) evaluator recommendation: a current recommendation of a major independent privacy evaluator. This is the bar IVPN enters under: its market share is niche, but independent evaluators single it out for its verification record, which is exactly what the X axis scores; - (d) architectural competitor: a split-trust or no-single-party claim comparable to URnetwork's own headline claim. This is the bar Obscura enters under, and the bar every reviewer or reader nomination is tested against first.
- It is scorable. Enough public evidence exists, in documentation, audits, measurements, and policies, to assign all nine sub-scores against the frozen ladders. A product that qualifies on reach but cannot be scored from that evidence is listed as pending, not scored on impression.
Disqualifiers: discontinued or unobtainable products (Google One VPN); enterprise-only offerings an individual cannot buy; white-label rebrands of an already-included stack (Betternet is not scored because Hotspot Shield already represents the Hydra stack); and self-hosted tools per condition 1.
How the set is revisited. The qualification checklist is re-run against dated market data at every corpus revision and at least every six months. Out-of-cycle triggers, any of which forces an immediate check: a product ships a no-single-party design (bar d); a product crosses an adoption bar; a reviewer or reader names a candidate. The nomination, date, and outcome get recorded whether or not the candidate passes, which is how Obscura's entry is recorded. Products that stop shipping are archived with their last-scored date, never silently removed.
The rule applied to the current set (2026-08-07). All twenty-two members qualify: Windscribe, PIA, ExpressVPN, Proton VPN, NordVPN, Surfshark, TunnelBear, IPVanish, Hotspot Shield and Hola under (a); Tor, NymVPN, Apple Private Relay, Cloudflare WARP, Tailscale, Orchid, Mysterium and Sentinel under (b); IVPN under (c); Mullvad under (a) and (c); URnetwork and Obscura, the latter scored 2026-08-07, the same day its comparison page was published, under (d). One borderline is named rather than glossed: Orchid is the member closest to the discontinued disqualifier (its mobile apps are years stale), stays while its shipped clients still function, and is the first row the revisit rule would archive.
What the rule says we are missing. A selection rule that never names a gap is decoration. Applying the same bars today names two products this corpus does not compare:
- CyberGhost clears (a) outright: it is among the largest consumer VPN brands by claimed user count, and our own evidence base already measures it (the IPinfo location study behind the egress scores includes CyberGhost, and the scoring notes cite its result). There is no principled reason it is absent while every sibling-scale major is in.
- Psiphon clears (b) as the most widely deployed dedicated censorship-circumvention service outside Tor, with documented mass adoption during national blocking events. The availability and anti-censorship part of the Y axis is its home ground, which makes the omission worse, not better.
Both are recorded as selection debt: they qualify, and until their pages and scores exist, this map's field is incomplete by its own published rule.
The weights, and why
Each index is a weighted sum of its parts; weights sum to 1.00 per axis.
X: verifiable anonymization index
| Part | Weight | Why this weight |
|---|---|---|
| S, separation of identity from activity | 0.35 | The heart of anonymization: does any single party hold both who you are and what you do, including the account and payment identity it takes to use the service? If one party holds both, everything else is mitigation, so this is the heaviest part. |
| V, "no logs" verifiable rather than promised | 0.30 | The axis is named verifiable. Open source, independent audits, court and raid evidence, and outside measurement move this score; a bare promise sits near zero however sincerely it is made. |
| E, end-to-end encryption through the service | 0.25 | Whether the operator is structurally unable to read your traffic, or merely promises not to. Encryption to a company that decrypts everything is not encryption through it. Second only to separation, because it is the other structural half of "cannot see". |
| PQ, post-quantum encryption | 0.10 | Whether the key exchange resists harvest-now-decrypt-later attacks. A real, forward-looking differentiator, but not yet decisive for present-day anonymity, hence the smallest weight. |
The deliberate stance inside these numbers: structure (S + E = 0.60) outweighs verification (V = 0.30), on the reasoning that structure is the thing verification exists to check. An audited promise is still a promise, while a design that removes the ability to see needs less trusting. A reader who weighs the evidence record above the architecture will reorder the top of the axis, and the sensitivity page works that alternative through rather than hiding it.
Y: connection quality index
| Part | Weight | Why this weight |
|---|---|---|
| L, latency | 0.30 | The most user-perceptible property of a connection, and the one most determined by architecture: hop count, path length, edge proximity. Scored as a structural class; no product here has published latency figures the corpus accepts, so there are no milliseconds anywhere in this scoring. |
| T, speed / throughput | 0.25 | Second most perceptible. Also structural (datacenter single-hop, residential-bounded, relay chain) plus each vendor's own published positioning. The one throughput number in the whole corpus is URnetwork's own, attributed to URnetwork: 40 Mbps+ average streaming speed. |
| Eg, egress quality | 0.20 | Whether the exit delivers what it claims: measured location truth (the IPinfo study), hosting-flagged versus residential address space, and how finely you can choose where you appear. Weighted third because a fast tunnel to an exit that is blocked, or somewhere other than claimed, is not a working connection for the person who chose it. |
| A, availability / anti-censorship | 0.15 | Whether you can get it and keep it working: free access, platform breadth, behavior where networks are hostile, regional withdrawal. |
| In, ingress options | 0.10 | The ways in: protocols, obfuscated transports, volunteer entry paths. Lowest weight because it matters enormously to some users and not at all to most. |
The freeze
The weights and the anchor ladders were written down and frozen on 2026-08-07, before any product was scored, and were not adjusted afterward. That is the first thing a skeptic should test, and the frozen rubric text with its date sits in the repository at review/index/METHOD.md. The rubric also pre-committed, in writing, to the outcomes most tempting to soften: URnetwork's verifiability score must be penalized for having no third-party audit; its sealed session must be scored as the opt-in it then was; and Mullvad, IVPN, Tor and NymVPN score whatever the evidence says, even where that places them above URnetwork on a part or, for Mullvad and IVPN, in the same quadrant. The frozen wording on the first of those is left as written, and it was factually wrong in our own disfavour: two third-party assessments existed at the time (the ledger row above, and the V reasoning on the scores page). The penalty it demanded still applies, because neither assessment touches the layer the axis measures, and the score was re-derived rather than adjusted. The second pre-commitment ended by ruling rather than by error: on 2026-08-07 the product owner ruled the sealed session on by default at launch, so the opt-in premise describes a state the product left. The affected sub-scores were re-derived from the unchanged anchors rather than adjusted; the ledger carries the change with its direction, and the frozen text again stands as written.
Four things did change after scoring began, and all are disclosed rather than smoothed over:
- One sub-score was normalized during scoring. IVPN's latency part was first scored 6 on account of its smaller fleet, then normalized to 7 to match every other single-hop datacenter VPN. The corpus records no per-vendor latency deficits, and docking one vendor for fleet size without measurements would have been invented precision. The working record keeps the original value, and IVPN's entry on the scores page prints the effect either way.
- The origin definition was corrected after scoring, by the product owner, and it is the only post-scoring change to the method. The next section explains it.
- Apple Private Relay was rescored after independent peer review, on the same day as first scoring: the evidence base itself, and every page built on it, had misdescribed Apple's split as held by contract when Apple's own overview describes it as enforced by layered encryption. E 5→7 and S 6→7 against the unchanged frozen anchors, X 4.25→5.10, which moves Apple from the Legacy quadrant to the X line itself. The error had run in URnetwork's favour, which is the direction we most need to catch; no other product's scores were touched. The full re-derivation is in Apple's entry on the scores page.
- URnetwork was rescored after the product owner ruled its sealed session on by default at launch: the second post-scoring rescore, and the one that moves our own dot. E 7→8 and PQ 5→7 against the unchanged frozen anchors, S re-derived and unchanged at 8, V untouched at 6; X 6.85→7.30. The re-derivation in URnetwork's entry on the scores page includes the explicit ruling that the encryption ladder's higher rung requires the default but not yet full integrity, and prints the seal's standing fail-open and key-substitution defects beside the score rather than waiving them.
These four are the score- and origin-level changes. The complete dated ledger, including the post-review method changes with the direction each cuts, is near the top of this page.
Why the axes cross at (5, 5)
The axes cross at (5.00, 5.00), the fixed center of the 0–10 scale. A position is therefore absolute against the rubric's anchors: crossing an axis line means crossing the scale midpoint, whatever the rest of the field looks like.
The first draft of the rubric defined the origin differently, as the centroid of the middle of the field (the mean of the middle eleven products per axis, which computes to (3.81, 5.66) on these scores). That definition was replaced, after scoring, because a centroid origin is relative to whoever happens to be on the chart: add three weak products and every other product moves rightward across the quadrant lines without anything about them changing. A fixed scale center is stable under any change to the field and lets a reader reason about a position without knowing who else was scored. The correction changed only where the axis lines cross, and therefore which quadrant name some products carry; no weight, anchor, or sub-score was changed with it, and nothing was re-scored afterward. Both origin definitions and the correction date are preserved in the working record.
One reading rule follows from the arithmetic. Because sub-scores are integers, a one-point change on a mid-weight part moves an index by about 0.25, which is why, throughout these pages, a product within ±0.25 of an axis line is read as sitting on the line rather than confidently on either side of it. Which products sit inside that band is tabulated with the coordinates on the scores page, and what single-point re-scores would move is worked through on the sensitivity page.
The scale: how the anchor ladders work
Every sub-score is an integer, 0–10, assigned against an anchor ladder written before any product was scored: a table of rungs that names the class of design or evidence each score band means. An integer that falls between two rungs means the product sits between those classes. As calibration across all nine parts: a score around 2 means the property is essentially absent, or exists as words only; around 5, present but partial, bounded, or held by contract and policy rather than by structure; around 9, present by construction, on by default, with the strongest evidence in the field. The nine ladders are printed in full with the matrix on the scores page, and the frozen originals sit in review/index/METHOD.md.
The limits
This is the part of the page that earns the rest, so none of it is softened.
- These are scores of published evidence, not measurements. The inputs are the twenty-one comparison pages, vendors' own disclosure pages, audit and court records, and the IPinfo measurement study. Where outside measurement exists it moved scores; the egress column leans directly on IPinfo's results. Every latency and throughput score on the map is a structural class judgement, because the corpus forbids invented performance figures. A real measurement campaign could reorder the middle of the quality axis, and nothing on this page should be mistaken for one.
- Nobody has measured URnetwork from outside. IPinfo measured Mullvad, IVPN and Windscribe at 0% location mismatch; no equivalent study has been run on URnetwork's exits, latency, or throughput, in either direction. Its Eg 8 rests on a mechanism you can check in code and on the structural properties of residential address space, not on a third-party result, and its L 5 and T 5 are unmeasured structural anchors. The asymmetry is stated here because it cuts against the product this documentation belongs to.
- No independent audit covers URnetwork's protocol, connect engine or server code, and the axis is named verifiable. That absence is priced in, V 6 against Mullvad's 9, IVPN's 8 and PIA's 8, and it is the single largest criticism of URnetwork this corpus records: full-stack open source is continuous verifiability of design, but no independent check of the layer the claim rests on, no court test and no outside measurement is a real hole in a claim of verifiability, and it costs real points on the axis that carries the product's quadrant name. The two third-party assessments that do exist cover other surfaces and moved nothing; the ledger above and the V reasoning on the scores page carry them. Until 2026-08-07 this corpus said flatly that there were none; that was wrong, and correcting it changed the sentence, not the score.
- One rule for defaults versus capabilities, applied both ways, with its inputs updated once by ruling. The rule: a product is scored on its shipping default, and an opt-in headline capability is credited as conditional structure, not as the default. Under it, URnetwork's then-opt-in sealed session scored E 7 while NymVPN's opt-in mixnet held its E at 9 rather than 10. When the product owner ruled URnetwork's session on by default at launch (2026-08-07, disclosed in the ledger), the same rule moved URnetwork to E 8. The facts changed, not the rule, while Nym's mixnet remains credited as the opt-in capability it still is. A reader who scores shipped defaults only should still lower Nym's E; one who scores capabilities only should raise it. What is not defensible is doing one without the other.
- Some placements sit inside noise, and some orderings are weight judgements. Apple Private Relay sits within noise of both axis lines at once; Obscura sits on the quality line; Mysterium, Hola and Sentinel sit within noise below it; and defensible re-weightings reorder the top of the X axis, including against URnetwork. The named list, the single-point arithmetic, and the worked re-weightings are on the sensitivity page, beside the Monte Carlo record they belong to.
- The scorer has a conflict of interest. URnetwork scored this field, and URnetwork is on the map. The mitigations are the ones on this page and its appendices: weights and anchors frozen before scoring, every sub-score published with its reasoning, judgement calls labeled, the re-weighting that demotes URnetwork printed rather than buried. None of them substitutes for an outside scorer. The matrix is deliberately complete enough that you can be that scorer: change the weights, or any sub-score you can argue from evidence, and the map is yours to redraw.
If the picture changes
If new evidence lands, an independent audit of URnetwork's protocol or the operator's server code, an outside measurement of its network, or a competitor shipping verifiable structural separation, the standing rule is to rescore and republish, in that order, and the ledger above gains a row with its direction marked. If a number on this page or its appendices fails to recompute, or a page drifts from the working record in review/index/, report it like any other documentation bug: issues on the repositories (github.com/urnetwork), or feedback in the apps.