How every index in the report is actually computed.
The combined report carries two headline composites. CXI — Cyber Exposure Index — rolls up the five exposure indices (CES, IMI, RPS, BLV, BIR): how exposed you are, lower is better. MPI — Market Position Index — rolls up the four market indices (WTI, SVI, BZV, SRS): how you are doing versus the market, higher is better. The competitive diagnostics (SOV, CTB, FLR) are reported beside MPI against a named peer set. Every index has a published formula, an enumerated source whitelist, weighted inputs, an uncertainty band, and a change log. This page is the contractual definition. The internal collector implementation lives behind it; the math does not.
What the methodology commits to before any individual index.
Quantification before collection
We instrument fewer sources than generalist threat-intel vendors. Every collected signal must feed a published index. Volume is not the product.
Comparable across clients
Indices are normalized so peer benchmarks are meaningful at N ≥ 5. Two firms with the same CES represent comparable credential exposure, controlled for workforce size and sector.
Uncertainty bands, always
Every reported value ships with an uncertainty band and a named method (currently a heuristic driven by source coverage and recency). We label it honestly — it is not a calibrated confidence interval, and we say so.
Public + licensed sources only
Collection sources are explicitly enumerated. No invite-only forums, no authentication bypass, no interaction with sellers. The source whitelist is part of the contract.
Not measured is never zero
When every source behind an index is unavailable, the index reports a data gap — "not measured" — and is excluded from its composite, which re-balances over what was measured. A source outage can never read as a low score.
Identifiers hashed at ingest
Emails, usernames, and any customer or employee identifier are SHA-256 hashed with a rotating salt before they hit our storage. Dereferencing requires verified domain ownership and a documented lawful basis.
Version-anchored, audit-replayable
Every index output carries the methodology version (currently v0.7). Auditors can independently recompute any historical value from preserved raw signals.
02·Exposure layer — the CXI composite, then its five indices
The exposure headline, then each index it rolls up. 0–100, higher is worse: Low / Elevated / High / Critical.
Σ
CXI
Cyber Exposure Index
Headline composite
A single 0–100 board-level number (higher is worse; bands Low / Elevated / High / Critical): a client-weighted roll-up of the five indices below. The CXI introduces no new collection — it is strictly a function of CES, IMI, RPS, BLV, and BIR. Internal collection sub-signals are absorbed into their parent index before the roll-up; none are reported as standalone indices. An index whose sources were all unavailable is reported as a data gap and excluded — the weights re-balance over what was measured.
Each index is normalized to a common 0–100 scale, weighted, and summed. Weights sum to 1.0 and re-balance over the measured (non-data-gapped) inputs.
Default weights
CES 0.30 · RPS 0.30 · IMI 0.19 · BIR 0.11 · BLV 0.10
Per-client weights are set at onboarding to match the buyer's risk priorities and audited quarterly. The applied vector ships with every report. Re-weighting is a client-level configuration, not a methodology version change.
Caveats
·A high reading on one index is surfaced as a standalone alert, never averaged away by a calm composite.
·Peer percentile on the CXI requires a cohort of N ≥ 5; below that we report the absolute value and trend only.
·The composite's uncertainty band is propagated from its components (root-sum-of-squares of weighted half-widths) — it never reads tighter than its widest weighted input.
Volume and severity of credentials tied to the client's domains — infostealer infections and breach-tied addresses — observed over a rolling window.
Decision it enables
Force proactive password resets by cohort. Reduce account-takeover incidents in the first quarter.
Reporting
0–100 with an uncertainty band and confidence label. High-severity fraction, driver counts, and Δ since the last period. Multi-source corroboration tightens the band; single-source is flagged.
A severity-weighted, log-compressed blend of infostealer hits (high severity) and breach-tied addresses, mapped through an anchored sigmoid so a measured zero reads ≈0 (not a sigmoid floor) and the top of the curve is 100. A measured zero is a real 'no exposure detected'; if every source is unavailable it is a data gap instead.
Inputs and source weights
Hudson Rock Cavalier (free, observe-only)
high
Infostealer infections by domain — the high-severity signal
Have I Been Pwned domain search (optional, verified-owner)
medium
Breach-tied addresses; only when HIBP_API_KEY is set
DeHashed (optional, pay-as-you-go)
low
Corroboration; not wired by default
Aggregation
Identifiers are SHA-256 hashed at ingest (per-tenant salt). Per-identifier dereferencing requires verified domain ownership and a documented lawful basis.
Caveats
·A hit in two sources scores more confidently than a single-source claim — single-source runs carry the single-source tag.
·Store only hashes of any enumerated identifiers; plaintext dereference requires verified ownership.
·Roadmap (not yet contracted): licensed stealer-log feeds — SpyCloud, Constella, Flare — would add breadth and family attribution.
Frequency and pricing of initial-access-broker listings referencing the client across enumerated public and low-cost sources — a best-effort, public-only signal today.
Decision it enables
Notify and harden the perimeter before brokered access is sold to a ransomware affiliate or downstream operator.
Reporting
Listing count, average price, per-source status, and Δ since last period, with an uncertainty band. Public-only runs are declared as likely undercounts.
Manually-corroborated public IAB mentions drive it; low-cost/public sources (Intelligence X, public Telegram previews, OTX pulses) are log-damped corroboration; a listed access price adds a small factor. OTX alone is capped — a raw pulse count is not a confirmed listing.
Inputs and source weights
Manually-corroborated public IAB mentions
high
Analyst-verified, observe-only
Intelligence X (optional, ~€7/mo)
medium
Paste/forum/leak index selector search; counts only
Exposure to active ransomware crews from externally-observable signals — exploited-vulnerability exposure, sector victim similarity, and infrastructure reputation.
Decision it enables
Quantify ransomware exposure for cyber-insurance pricing, vendor reviews, and board-level risk reporting.
Reporting
0–100 score with an uncertainty band and confidence label. Exposed CVEs, victim-similarity context, effective weights_used, and Δ since last period. Perimeter CVEs come from the EASI feeder (run first).
A weighted blend of exposed CISA-KEV CVEs on the perimeter, EPSS-weighted exploit probability, sector victim-similarity, and ASN infrastructure reputation, through an anchored sigmoid so a clean host reads ≈0. The four implemented terms are renormalized (÷0.80) because the designed FamilyActivity term is not yet wired.
Inputs and source weights
CISA KEV catalog (free)
0.30
Known-exploited CVEs present on the client perimeter
FIRST EPSS (free)
0.25
Exploit-probability weighting of those CVEs (capped)
ransomware.live recent victims (free)
0.15
Sector victim-similarity proxy
Infra reputation (Spamhaus/URLhaus/AbuseIPDB counts)
0.10
Client-supplied ASN hit count
abuse.ch ThreatFox (FamilyActivity)
TODO
Designed 0.20 term — NOT wired; excluded and the four terms above renormalized
Aggregation
FamilyActivity is excluded (not wired) and the four measured terms are renormalized to sum to 1.0; victim-similarity is a sector-share proxy. Both are stated in the finding note.
Caveats
·We do not interact with ransomware operators. No negotiation, no validation purchases, no contact.
·FamilyActivity (ThreatFox) is not wired — the published formula reflects the four implemented terms, renormalized.
·VictimSimilarity is a coarse sector-share proxy today; refine with revenue band + geo Jaccard.
Rate at which brand assets and data appear across public leak channels — ransomware-leak-site victim mentions and brand/domain phishing URLs — per trailing week.
Decision it enables
Prioritize takedown queues and legal escalation when leakage accelerates beyond the established baseline.
Reporting
Weekly signal with an uncertainty band and Δ since last period. True week-over-week slope reports once ≥3 cycles are stored.
A weekly leak-velocity signal: ransomware-leak-site victim mentions (weighted 3×, the heavier signal) plus brand/domain phishing URLs, through an anchored sigmoid so a quiet period reads ≈0.
Inputs and source weights
ransomware.live recent victims (free)
high
Client named as a victim (leak-site mention)
OpenPhish / PhishStats (free)
medium
Brand/domain phishing URLs
Public paste sites
roadmap
Best-effort paste monitoring — not yet wired
Aggregation
An ambiguous brand token is matched on the domain only, to avoid common-word false positives (e.g. 'bang' matching 'bangkok').
Caveats
·Velocity (the slope) matters more than the absolute count; it needs several stored cycles to stabilize.
·When both feeds are unavailable this is a data gap, never 'no leaks'.
·Roadmap (not yet contracted): licensed leak-channel coverage for broader reach.
Audience-weighted reach of typosquat and impersonation domains imitating the client — count × audience × persistence across impersonation surfaces.
Decision it enables
Direct defensive registrar spend, platform reporting, and takedown allocation against impersonators.
Reporting
Weighted reach with an uncertainty band, top impersonations, excluded legitimate neighbors, and Δ since last period. Without urlscan, the score is a raw (unweighted) count and tagged as such.
Sum over impersonation surfaces of count × audience weight (urlscan scans) × persistence (first→last sighting span). Established independent domains (long urlscan history) are excluded as typo-neighbors, not impersonations. Anchored so zero reach reads ≈0.
Inputs and source weights
dnstwist (local CLI)
high
Permutations that actually resolve (A record)
crt.sh / certspotter (free CT logs)
high
Exact brand-label certificate domains
urlscan.io (free key)
medium
Audience (scans) + persistence weighting; neutral weight if absent
Aggregation
Cert-only sightings carry a lower fixed weight (a cert string is the weakest reach surface). Enrichment is capped per run for free-tier courtesy.
Caveats
·Without a urlscan key the count is unweighted and legit typo-neighbors are not filtered — the finding carries the unweighted-count tag.
·Persistence is an observed-span proxy; extend with WHOIS age and true days-live-until-takedown.
·Extend to mobile / social / extension surfaces for fuller coverage.
02b·Market & competitive layers — the MPI composite
The market headline, then the peer-relative diagnostics. 0–100, higher is better: Weak / Moderate / Strong / Dominant.
Σ
MPI
Market Position Index
Headline composite
A single 0–100 board-level number (higher is better; bands Weak / Moderate / Strong / Dominant) answering “how are we doing vs. the market?” — a weighted roll-up of the four market indices. Sources are public or licensed, ToS-respecting, and brand/market-aggregate only: never individual-person profiling. Data-gapped inputs are excluded and the weights re-balance; the exclusions are listed in the report.
Default weights: WTI 0.30 · SVI 0.25 · BZV 0.25 · SRS 0.20 — set at onboarding, audited quarterly, shipped with every report. The exposure indices feed CXI, never MPI; the competitive diagnostics are reported beside MPI, never folded into it.
WTI
Web Traffic & Growth
Traffic strength and growth from public domain-rank data (licensed visit estimates optional).
Marketing ROI and channel health — is the audience growing?
SVI
Search & SEO Visibility
Search interest and lookup-demand trend across public search and encyclopedia signals.
Demand trend — is the top of the funnel warming or cooling?
BZV
Buzz Volume
Public mention volume across news, social, and forum sources over a trailing window.
Attention and campaign lift; doubles as incident early-warning.
SRS
Sentiment & Reputation
Tone of public coverage and mentions; 50 is neutral; no coverage is a data gap, never a score.
Brand health and PR risk, trended period over period.
SOV
Share of Voice · competitive
The brand's share of the peer set's total public mention volume.
Who is winning attention in the category.
CTB
Competitive Traffic Benchmark · competitive
Position in the peer-set traffic leaderboard from the same public rank data as WTI.
Competitive position and growth gaps vs. named rivals.
FLR
Feature & Launch Radar · competitive
Peer site/app changes and launch press, detected by public diffing and news signals.
Product gaps and rival shipping velocity.
Published formulas — market & competitive
WTI
WTI = 100 · (1 − log₁₀(rank) / log₁₀(10M))
Public domain-rank (Tranco) on a log scale; growth = the change over the trailing window. Unranked = data gap, never 0.
SVI
SVI = mean(GoogleTrends₀₋₁₀₀, WikipediaPageviews→0₋₁₀₀)
Mean of whatever public-interest signals returned (Google Trends already 0–100; Wikipedia pageviews log-scaled). Data gap only when both are unavailable.
Competitor homepage hash-diffs (baseline on first run) + GDELT launch/announcement press (14d). A feed outage is a data gap, not 'no activity'.
Why two band scales
The two headline numbers point in opposite directions. Exposure uses risk bands that get scarier as the number climbs (Low → Critical); Market uses strength bands that get stronger (Weak → Dominant). Using one scale for both would invert a score's meaning.
Caveats
·Peer percentile on MPI requires a cohort of N ≥ 5; below that we report the absolute value and trend only.
·Market and competitive collection is brand / site / category level only — never individual-person profiling.
·SOV, CTB, and FLR explain the score; they never move it. They are excluded from the MPI roll-up by construction.
03·Source whitelist
Enumerated. Contractual.
The sources we collect from are part of the contract. Adding a source requires a methodology version bump and a notice to existing clients. Removing a source likewise. Source weights are reviewed quarterly.
Licensed dark-web data partners
CES+IMI+RPS · roadmap
DarkOwl, SpyCloud, Constella, Flare — licensed-partner roadmap; not yet contracted. Current runs use the free/low-cost sources enumerated per index above.
Public stealer-log markets
CES+IMI · medium-high
Observed via archive mirrors; no purchases
Public ransomware leak sites
RPS+BLV · high
Tor mirror observation; victim postings, data-claim metadata
Tranco, Cloudflare Radar; Similarweb when licensed
Public search-trend + encyclopedia signals
SVI · high
Google Trends, Wikipedia pageviews; Semrush/Ahrefs when licensed
Public news / web mention indexes
BZV+SRS+SOV+FLR · high
GDELT, Google News, industry press; brand-aggregate only
Public social / forum APIs (keyed, ToS-respecting)
BZV+SOV · medium
Public Reddit API; no auth bypass, no scraping behind login
04·Identifier handling
Hashed at ingest. Dereferencing requires verified ownership.
The minimum-necessary principle is enforced at the collector boundary. Plaintext identifiers are normalized, hashed with SHA-256 + a salt rotated every 90 days, and discarded. Only the hash + metadata (source, observation time, severity flags) is persisted.
×Personally-identifying content from leaked documents
×Any signal touching CSAM (see red lines)
Verified-ownership dereferencing — for instance, a CISO needing to notify affected employees — is bound to a documented lawful basis (GDPR Art. 6(1)(f) + 34) and is logged for audit. The dereferencing endpoint is rate-limited and tenant-scoped.
05·Peer-cohort anonymization
N ≥ 5. No exceptions.
Peer benchmarks (the P<value> percentile shown in the report) are computed across a bucket of firms with comparable characteristics: sector, workforce-size band, and revenue band. A benchmark is suppressed if the bucket size falls below five firms.
·Bucket definitions are deterministic and published with the methodology version.
·Per-client inputs to peer aggregates are hash-blinded; we cannot retrieve which raw value came from which client when computing the bucket statistic.
·Outputs to client A are statistically constructed so client A's own contribution cannot be re-identified from the published percentile.
·Cross-client raw data is never accessible to any individual client, employee, or analyst.
06·Uncertainty bands
A heuristic band. Labeled honestly — not a calibrated CI.
Every index output ships with an uncertainty band. We are precise about what it is: a heuristic width (method heuristic-v1) driven by how many sources answered and how fresh the freshest observation is — thin or stale evidence widens the band. It is not a distributional confidence interval and we do not claim a coverage guarantee behind it. Calling it a calibrated confidence interval would overstate the math, so we don't.
For a composite (CXI, MPI) the band is propagated from its components, not recomputed from the number of sources — the composite half-width is the root-sum-of-squares of each component's weighted half-width, so a composite can never read tighter than its widest weighted input. As real observation history accumulates, a calibrated interval will replace the heuristic and the change will be versioned.
Reported as
RPS = 41.0 [33 – 49] ← 0–100 scale · heuristic-v1 uncertainty band
↑ ↑ ↑
central lower upper
Caveat tags on a finding
Individual findings carry caveat tags — the guardrails that separate a defensible number from a raw feed. Four appear across the report:
single-source
Only one source answered this period, so the value rests on a single feed — the uncertainty band is wider and confidence lower until a second source corroborates.
unverified-candidate
The finding is an unconfirmed match (e.g. a public-code secret candidate, or a probabilistic attribution) — it is capped below the top band until independently corroborated.
capped
A weak, corroboration-only signal (e.g. OTX pulses) whose contribution is deliberately limited so it can never headline a score on its own.
audience-weighted
Reach is weighted by observed audience and persistence (e.g. BIR with urlscan), not a raw count — a pile of parked, never-visited squats scores lower than a few actively-hit ones.
07·Versioning & change log
Every output carries the version it was computed with.
Methodology versions are immutable once shipped. Historical index values are recomputed-only on explicit request and the new values carry both the original version and the recompute version. Clients are notified at least 14 days before any minor version bump and 30 days before any major bump.
v0.7
2026-06
Combined-report extension. Added the Market layer (WTI, SVI, BZV, SRS) rolled into a second headline composite — MPI, Market Position Index — and the Competitive diagnostics (SOV, CTB, FLR), reported beside MPI but never folded into it. Adopted the 0–100 publication scale for all indices and composites (previously 0–10). Two band scales formalized: risk bands for the exposure layer (Low / Elevated / High / Critical, higher is worse) and strength bands for the market layer (Weak / Moderate / Strong / Dominant, higher is better). Data-gap rule published: an index whose sources were all unavailable reports 'not measured' and is excluded from its composite, which re-balances over the measured inputs.
v0.6
2026-05
Shipped the free CES teaser: a verified-owner, single-index snapshot of one's own domain. Benchmarks against a documented modeled cohort prior (sector × workforce-size band) until a bucket reaches the N ≥ 5 threshold, at which point an observed peer percentile supersedes the prior. Modeled vs. observed basis is labeled on every output — a prior is never presented as an observed percentile.
v0.5
2026-05
Published the CXI headline composite — a client-weighted roll-up of the five indices into a single 0–10 board number. Default weights documented; per-client weights are set at onboarding and audited quarterly. Internal collection sub-signals are absorbed into their parent index upstream and are never reported as standalone indices.
v0.4
2026-05
Added supplier-graph overlap signal to RPS. Refined β weighting for CES high-severity records. Cross-listing identity attribution now reported with confidence bands.
v0.3
2026-03
Introduced peer-cohort benchmarks (N ≥ 5). Added uncertainty-band reporting standard across all five indices. Cross-vendor deduplication agreement check formalized.
Initial release with five indices: CES, IMI, RPS, BLV, BIR. Methodology baseline.
08·Red lines
What this methodology will never compute.
The collection boundary is part of the math. Outputs are only as trustworthy as the inputs that produced them — these rules bound the inputs.
×No authentication bypass. No invite-only forums. No vouched-access communities.
×No buying stolen accounts or leaked data 'to validate.' We observe; we do not interact with sellers.
×No plaintext PII to clients. Hashes + metadata only.
×No cross-client data leakage. Peer benchmarks require N ≥ 5 in a bucket.
×No doxxing. No offensive OSINT. No HUMINT. No undercover personas.
×Zero tolerance for CSAM. Automated detection → immediate NCMEC report → zero retention.
Questions on the math?
We respond to methodology questions in writing.
CISO, head of risk, or anyone on the legal side — send the specific question and we'll respond inside 48h. First-cohort tier conversations get a longer-form methodology brief on request.
Every index output is anchored to (a) the methodology version, (b) the inputs at the time of computation, and (c) a SHA-256 hash of the evidence bundle. An auditor with read access can independently replay any historical value. Methodology audits are part of the first-cohort onboarding.