Methodology · v0.7 · June 2026 · CISM-led

How every index in the report
is actually computed.

The combined report carries two headline composites. CXI — Cyber Exposure Index — rolls up the five exposure indices (CES, IMI, RPS, BLV, BIR): how exposed you are, lower is better. MPI — Market Position Index — rolls up the four market indices (WTI, SVI, BZV, SRS): how you are doing versus the market, higher is better. The competitive diagnostics (SOV, CTB, FLR) are reported beside MPI against a named peer set. Every index has a published formula, an enumerated source whitelist, weighted inputs, an uncertainty band, and a change log. This page is the contractual definition. The internal collector implementation lives behind it; the math does not.

01·Principles

What the methodology commits to before any individual index.

Quantification before collection

We instrument fewer sources than generalist threat-intel vendors. Every collected signal must feed a published index. Volume is not the product.

Comparable across clients

Indices are normalized so peer benchmarks are meaningful at N ≥ 5. Two firms with the same CES represent comparable credential exposure, controlled for workforce size and sector.

Uncertainty bands, always

Every reported value ships with an uncertainty band and a named method (currently a heuristic driven by source coverage and recency). We label it honestly — it is not a calibrated confidence interval, and we say so.

Public + licensed sources only

Collection sources are explicitly enumerated. No invite-only forums, no authentication bypass, no interaction with sellers. The source whitelist is part of the contract.

Not measured is never zero

When every source behind an index is unavailable, the index reports a data gap — "not measured" — and is excluded from its composite, which re-balances over what was measured. A source outage can never read as a low score.

Identifiers hashed at ingest

Emails, usernames, and any customer or employee identifier are SHA-256 hashed with a rotating salt before they hit our storage. Dereferencing requires verified domain ownership and a documented lawful basis.

Version-anchored, audit-replayable

Every index output carries the methodology version (currently v0.7). Auditors can independently recompute any historical value from preserved raw signals.

02·Exposure layer — the CXI composite, then its five indices

The exposure headline, then each index it rolls up. 0–100, higher is worse: Low / Elevated / High / Critical.

Σ
CXI

Cyber Exposure Index

Headline composite

A single 0–100 board-level number (higher is worse; bands Low / Elevated / High / Critical): a client-weighted roll-up of the five indices below. The CXI introduces no new collection — it is strictly a function of CES, IMI, RPS, BLV, and BIR. Internal collection sub-signals are absorbed into their parent index before the roll-up; none are reported as standalone indices. An index whose sources were all unavailable is reported as a data gap and excluded — the weights re-balance over what was measured.

Formula
CXI = Σ_i  w_i · normalize(index_i)     i ∈ {CES, IMI, RPS, BLV, BIR}

Each index is normalized to a common 0–100 scale, weighted, and summed. Weights sum to 1.0 and re-balance over the measured (non-data-gapped) inputs.

Default weights

CES 0.30 · RPS 0.30 · IMI 0.19 · BIR 0.11 · BLV 0.10

Per-client weights are set at onboarding to match the buyer's risk priorities and audited quarterly. The applied vector ships with every report. Re-weighting is a client-level configuration, not a methodology version change.

Caveats
  • ·A high reading on one index is surfaced as a standalone alert, never averaged away by a calm composite.
  • ·Peer percentile on the CXI requires a cohort of N ≥ 5; below that we report the absolute value and trend only.
  • ·The composite's uncertainty band is propagated from its components (root-sum-of-squares of weighted half-widths) — it never reads tighter than its widest weighted input.
01
CES

Credential Exposure Score

# permalink

Volume and severity of credentials tied to the client's domains — infostealer infections and breach-tied addresses — observed over a rolling window.

Decision it enables

Force proactive password resets by cohort. Reduce account-takeover incidents in the first quarter.

Reporting

0–100 with an uncertainty band and confidence label. High-severity fraction, driver counts, and Δ since the last period. Multi-source corroboration tightens the band; single-source is flagged.

Formula
CES = 100 · anchoredσ( (0.7·ln(1+stealer_hits) + 0.3·ln(1+breach_addrs)) / 2.2 − 1.0 )

A severity-weighted, log-compressed blend of infostealer hits (high severity) and breach-tied addresses, mapped through an anchored sigmoid so a measured zero reads ≈0 (not a sigmoid floor) and the top of the curve is 100. A measured zero is a real 'no exposure detected'; if every source is unavailable it is a data gap instead.

Inputs and source weights
Hudson Rock Cavalier (free, observe-only)
high
Infostealer infections by domain — the high-severity signal
Have I Been Pwned domain search (optional, verified-owner)
medium
Breach-tied addresses; only when HIBP_API_KEY is set
DeHashed (optional, pay-as-you-go)
low
Corroboration; not wired by default
Aggregation

Identifiers are SHA-256 hashed at ingest (per-tenant salt). Per-identifier dereferencing requires verified domain ownership and a documented lawful basis.

Caveats
  • ·A hit in two sources scores more confidently than a single-source claim — single-source runs carry the single-source tag.
  • ·Store only hashes of any enumerated identifiers; plaintext dereference requires verified ownership.
  • ·Roadmap (not yet contracted): licensed stealer-log feeds — SpyCloud, Constella, Flare — would add breadth and family attribution.
02
IMI

IAB Mention Index

# permalink

Frequency and pricing of initial-access-broker listings referencing the client across enumerated public and low-cost sources — a best-effort, public-only signal today.

Decision it enables

Notify and harden the perimeter before brokered access is sold to a ransomware affiliate or downstream operator.

Reporting

Listing count, average price, per-source status, and Δ since last period, with an uncertainty band. Public-only runs are declared as likely undercounts.

Formula
IMI = 100 · anchoredσ( ln(1 + manual + ln(1 + intelx + telegram + otx_capped))·1.4 + price − 1.0 )

Manually-corroborated public IAB mentions drive it; low-cost/public sources (Intelligence X, public Telegram previews, OTX pulses) are log-damped corroboration; a listed access price adds a small factor. OTX alone is capped — a raw pulse count is not a confirmed listing.

Inputs and source weights
Manually-corroborated public IAB mentions
high
Analyst-verified, observe-only
Intelligence X (optional, ~€7/mo)
medium
Paste/forum/leak index selector search; counts only
Public Telegram preview pages (t.me/s/<channel>)
medium
Read-only public web; no login, no joining
LevelBlue OTX pulses (free key)
low
Weak corroboration, capped; ambiguous selectors skipped
Aggregation

We observe only — no interaction with sellers, no test purchases, no personas. Counts only; no raw records are stored.

Caveats
  • ·Public-only / best-effort without a licensed forum feed — likely undercounts. This is stated in the report's source-status section.
  • ·OTX pulses are capped and, when the sole source, cannot headline a Critical.
  • ·Roadmap (not yet contracted): licensed dark-web feeds — DarkOwl, Flare — for high-fidelity forum coverage.
03
RPS

Ransomware Proximity Score

# permalink

Exposure to active ransomware crews from externally-observable signals — exploited-vulnerability exposure, sector victim similarity, and infrastructure reputation.

Decision it enables

Quantify ransomware exposure for cyber-insurance pricing, vendor reviews, and board-level risk reporting.

Reporting

0–100 score with an uncertainty band and confidence label. Exposed CVEs, victim-similarity context, effective weights_used, and Δ since last period. Perimeter CVEs come from the EASI feeder (run first).

Formula
RPS = 100 · anchoredσ( (0.30·KEV + 0.25·EPSS + 0.15·VictimSim + 0.10·Infra) / 0.80 / 6 − 1.5 )

A weighted blend of exposed CISA-KEV CVEs on the perimeter, EPSS-weighted exploit probability, sector victim-similarity, and ASN infrastructure reputation, through an anchored sigmoid so a clean host reads ≈0. The four implemented terms are renormalized (÷0.80) because the designed FamilyActivity term is not yet wired.

Inputs and source weights
CISA KEV catalog (free)
0.30
Known-exploited CVEs present on the client perimeter
FIRST EPSS (free)
0.25
Exploit-probability weighting of those CVEs (capped)
ransomware.live recent victims (free)
0.15
Sector victim-similarity proxy
Infra reputation (Spamhaus/URLhaus/AbuseIPDB counts)
0.10
Client-supplied ASN hit count
abuse.ch ThreatFox (FamilyActivity)
TODO
Designed 0.20 term — NOT wired; excluded and the four terms above renormalized
Aggregation

FamilyActivity is excluded (not wired) and the four measured terms are renormalized to sum to 1.0; victim-similarity is a sector-share proxy. Both are stated in the finding note.

Caveats
  • ·We do not interact with ransomware operators. No negotiation, no validation purchases, no contact.
  • ·FamilyActivity (ThreatFox) is not wired — the published formula reflects the four implemented terms, renormalized.
  • ·VictimSimilarity is a coarse sector-share proxy today; refine with revenue band + geo Jaccard.
04
BLV

Brand Leak Velocity

# permalink

Rate at which brand assets and data appear across public leak channels — ransomware-leak-site victim mentions and brand/domain phishing URLs — per trailing week.

Decision it enables

Prioritize takedown queues and legal escalation when leakage accelerates beyond the established baseline.

Reporting

Weekly signal with an uncertainty band and Δ since last period. True week-over-week slope reports once ≥3 cycles are stored.

Formula
BLV = 100 · anchoredσ( ln(1 + 3·leak_site_mentions + phishing_urls) / 1.4 − 0.6 )

A weekly leak-velocity signal: ransomware-leak-site victim mentions (weighted 3×, the heavier signal) plus brand/domain phishing URLs, through an anchored sigmoid so a quiet period reads ≈0.

Inputs and source weights
ransomware.live recent victims (free)
high
Client named as a victim (leak-site mention)
OpenPhish / PhishStats (free)
medium
Brand/domain phishing URLs
Public paste sites
roadmap
Best-effort paste monitoring — not yet wired
Aggregation

An ambiguous brand token is matched on the domain only, to avoid common-word false positives (e.g. 'bang' matching 'bangkok').

Caveats
  • ·Velocity (the slope) matters more than the absolute count; it needs several stored cycles to stabilize.
  • ·When both feeds are unavailable this is a data gap, never 'no leaks'.
  • ·Roadmap (not yet contracted): licensed leak-channel coverage for broader reach.
05
BIR

Brand Impersonation Reach

# permalink

Audience-weighted reach of typosquat and impersonation domains imitating the client — count × audience × persistence across impersonation surfaces.

Decision it enables

Direct defensive registrar spend, platform reporting, and takedown allocation against impersonators.

Reporting

Weighted reach with an uncertainty band, top impersonations, excluded legitimate neighbors, and Δ since last period. Without urlscan, the score is a raw (unweighted) count and tagged as such.

Formula
BIR = Σ_surface (count × audience_weight × persistence);   score = 100 · anchoredσ( ln(1 + weighted_reach) / 1.6 − 0.8 )

Sum over impersonation surfaces of count × audience weight (urlscan scans) × persistence (first→last sighting span). Established independent domains (long urlscan history) are excluded as typo-neighbors, not impersonations. Anchored so zero reach reads ≈0.

Inputs and source weights
dnstwist (local CLI)
high
Permutations that actually resolve (A record)
crt.sh / certspotter (free CT logs)
high
Exact brand-label certificate domains
urlscan.io (free key)
medium
Audience (scans) + persistence weighting; neutral weight if absent
Aggregation

Cert-only sightings carry a lower fixed weight (a cert string is the weakest reach surface). Enrichment is capped per run for free-tier courtesy.

Caveats
  • ·Without a urlscan key the count is unweighted and legit typo-neighbors are not filtered — the finding carries the unweighted-count tag.
  • ·Persistence is an observed-span proxy; extend with WHOIS age and true days-live-until-takedown.
  • ·Extend to mobile / social / extension surfaces for fuller coverage.
02b·Market & competitive layers — the MPI composite

The market headline, then the peer-relative diagnostics. 0–100, higher is better: Weak / Moderate / Strong / Dominant.

Σ
MPI

Market Position Index

Headline composite

A single 0–100 board-level number (higher is better; bands Weak / Moderate / Strong / Dominant) answering “how are we doing vs. the market?” — a weighted roll-up of the four market indices. Sources are public or licensed, ToS-respecting, and brand/market-aggregate only: never individual-person profiling. Data-gapped inputs are excluded and the weights re-balance; the exclusions are listed in the report.

Formula
MPI = Σ_i  w_i · normalize(index_i)     i ∈ {WTI, SVI, BZV, SRS}

Default weights: WTI 0.30 · SVI 0.25 · BZV 0.25 · SRS 0.20 — set at onboarding, audited quarterly, shipped with every report. The exposure indices feed CXI, never MPI; the competitive diagnostics are reported beside MPI, never folded into it.

WTI
Web Traffic & Growth
Traffic strength and growth from public domain-rank data (licensed visit estimates optional).
Marketing ROI and channel health — is the audience growing?
SVI
Search & SEO Visibility
Search interest and lookup-demand trend across public search and encyclopedia signals.
Demand trend — is the top of the funnel warming or cooling?
BZV
Buzz Volume
Public mention volume across news, social, and forum sources over a trailing window.
Attention and campaign lift; doubles as incident early-warning.
SRS
Sentiment & Reputation
Tone of public coverage and mentions; 50 is neutral; no coverage is a data gap, never a score.
Brand health and PR risk, trended period over period.
SOV
Share of Voice · competitive
The brand's share of the peer set's total public mention volume.
Who is winning attention in the category.
CTB
Competitive Traffic Benchmark · competitive
Position in the peer-set traffic leaderboard from the same public rank data as WTI.
Competitive position and growth gaps vs. named rivals.
FLR
Feature & Launch Radar · competitive
Peer site/app changes and launch press, detected by public diffing and news signals.
Product gaps and rival shipping velocity.
Published formulas — market & competitive
WTI
WTI = 100 · (1 − log₁₀(rank) / log₁₀(10M))

Public domain-rank (Tranco) on a log scale; growth = the change over the trailing window. Unranked = data gap, never 0.

SVI
SVI = mean(GoogleTrends₀₋₁₀₀, WikipediaPageviews→0₋₁₀₀)

Mean of whatever public-interest signals returned (Google Trends already 0–100; Wikipedia pageviews log-scaled). Data gap only when both are unavailable.

BZV
BZV = 100 · anchoredσ( ln(1 + weekly_mentions) / 1.5 − 1.0 )

Trailing-window public mention volume (GDELT, plus Reddit when keyed). The query is always disambiguated, never the bare brand word.

SRS
SRS = mean(GDELT_tone→0₋₁₀₀, NewsLexicon→0₋₁₀₀)

Mean of media tone (GDELT) and a news-headline sentiment lexicon, each mapped 0–100 around a neutral 50. No coverage is a data gap, not a score.

SOV
SOV = 100 · client_mentions / Σ peer-set mentions

The client's share of total peer-set buzz (7d). Every entity is searched with the same precision rule for parity; undefined total = data gap.

CTB
CTB = 100 · (n − position) / (n − 1)

Position in the peer traffic leaderboard from the same public rank data as WTI. An unranked client is a data gap, never last place.

FLR
FLR = 100 · anchoredσ( ln(1 + 2·homepage_changes + launch_press) / 1.4 − 0.6 )

Competitor homepage hash-diffs (baseline on first run) + GDELT launch/announcement press (14d). A feed outage is a data gap, not 'no activity'.

Why two band scales

The two headline numbers point in opposite directions. Exposure uses risk bands that get scarier as the number climbs (Low → Critical); Market uses strength bands that get stronger (Weak → Dominant). Using one scale for both would invert a score's meaning.

Caveats
  • ·Peer percentile on MPI requires a cohort of N ≥ 5; below that we report the absolute value and trend only.
  • ·Market and competitive collection is brand / site / category level only — never individual-person profiling.
  • ·SOV, CTB, and FLR explain the score; they never move it. They are excluded from the MPI roll-up by construction.
03·Source whitelist

Enumerated. Contractual.

The sources we collect from are part of the contract. Adding a source requires a methodology version bump and a notice to existing clients. Removing a source likewise. Source weights are reviewed quarterly.

Licensed dark-web data partners
CES+IMI+RPS · roadmap
DarkOwl, SpyCloud, Constella, Flare — licensed-partner roadmap; not yet contracted. Current runs use the free/low-cost sources enumerated per index above.
Public stealer-log markets
CES+IMI · medium-high
Observed via archive mirrors; no purchases
Public ransomware leak sites
RPS+BLV · high
Tor mirror observation; victim postings, data-claim metadata
Public paste sites + forum mirrors
BLV+IMI · medium-high
Pastebin-class, archive.org-indexed boards
Public Telegram channels
IMI+BLV · medium
MTProto public-channel API; no group infiltration
Certstream + WHOIS
BIR · high
Domain-registration monitoring, typosquat detection
Public incident disclosures
RPS · medium
SEC 8-K, regulator filings, news; cohort-clustering input
DNS / passive-DNS / ASN feeds
BIR · medium
Infrastructure-side correlation, mirror-operator clustering
Public domain-rank + traffic data
WTI+CTB · high
Tranco, Cloudflare Radar; Similarweb when licensed
Public search-trend + encyclopedia signals
SVI · high
Google Trends, Wikipedia pageviews; Semrush/Ahrefs when licensed
Public news / web mention indexes
BZV+SRS+SOV+FLR · high
GDELT, Google News, industry press; brand-aggregate only
Public social / forum APIs (keyed, ToS-respecting)
BZV+SOV · medium
Public Reddit API; no auth bypass, no scraping behind login
04·Identifier handling

Hashed at ingest. Dereferencing requires verified ownership.

The minimum-necessary principle is enforced at the collector boundary. Plaintext identifiers are normalized, hashed with SHA-256 + a salt rotated every 90 days, and discarded. Only the hash + metadata (source, observation time, severity flags) is persisted.

What we store

  • +SHA-256(identifier + current_salt)
  • +Source family + observation timestamp
  • +Severity flags (active session, captured cookies, etc.)
  • +Hash-of-hash for cross-client peer benchmarks

What we never store

  • ×Plaintext emails, usernames, or passwords
  • ×Payment-card fragments (dropped at parser)
  • ×Personally-identifying content from leaked documents
  • ×Any signal touching CSAM (see red lines)

Verified-ownership dereferencing — for instance, a CISO needing to notify affected employees — is bound to a documented lawful basis (GDPR Art. 6(1)(f) + 34) and is logged for audit. The dereferencing endpoint is rate-limited and tenant-scoped.

05·Peer-cohort anonymization

N ≥ 5. No exceptions.

Peer benchmarks (the P<value> percentile shown in the report) are computed across a bucket of firms with comparable characteristics: sector, workforce-size band, and revenue band. A benchmark is suppressed if the bucket size falls below five firms.

06·Uncertainty bands

A heuristic band. Labeled honestly — not a calibrated CI.

Every index output ships with an uncertainty band. We are precise about what it is: a heuristic width (method heuristic-v1) driven by how many sources answered and how fresh the freshest observation is — thin or stale evidence widens the band. It is not a distributional confidence interval and we do not claim a coverage guarantee behind it. Calling it a calibrated confidence interval would overstate the math, so we don't.

For a composite (CXI, MPI) the band is propagated from its components, not recomputed from the number of sources — the composite half-width is the root-sum-of-squares of each component's weighted half-width, so a composite can never read tighter than its widest weighted input. As real observation history accumulates, a calibrated interval will replace the heuristic and the change will be versioned.

Reported as
RPS = 41.0 [33 – 49]   ← 0–100 scale · heuristic-v1 uncertainty band
        ↑     ↑    ↑
    central lower upper
Caveat tags on a finding

Individual findings carry caveat tags — the guardrails that separate a defensible number from a raw feed. Four appear across the report:

single-source
Only one source answered this period, so the value rests on a single feed — the uncertainty band is wider and confidence lower until a second source corroborates.
unverified-candidate
The finding is an unconfirmed match (e.g. a public-code secret candidate, or a probabilistic attribution) — it is capped below the top band until independently corroborated.
capped
A weak, corroboration-only signal (e.g. OTX pulses) whose contribution is deliberately limited so it can never headline a score on its own.
audience-weighted
Reach is weighted by observed audience and persistence (e.g. BIR with urlscan), not a raw count — a pile of parked, never-visited squats scores lower than a few actively-hit ones.
07·Versioning & change log

Every output carries the version it was computed with.

Methodology versions are immutable once shipped. Historical index values are recomputed-only on explicit request and the new values carry both the original version and the recompute version. Clients are notified at least 14 days before any minor version bump and 30 days before any major bump.

v0.7
2026-06
Combined-report extension. Added the Market layer (WTI, SVI, BZV, SRS) rolled into a second headline composite — MPI, Market Position Index — and the Competitive diagnostics (SOV, CTB, FLR), reported beside MPI but never folded into it. Adopted the 0–100 publication scale for all indices and composites (previously 0–10). Two band scales formalized: risk bands for the exposure layer (Low / Elevated / High / Critical, higher is worse) and strength bands for the market layer (Weak / Moderate / Strong / Dominant, higher is better). Data-gap rule published: an index whose sources were all unavailable reports 'not measured' and is excluded from its composite, which re-balances over the measured inputs.
v0.6
2026-05
Shipped the free CES teaser: a verified-owner, single-index snapshot of one's own domain. Benchmarks against a documented modeled cohort prior (sector × workforce-size band) until a bucket reaches the N ≥ 5 threshold, at which point an observed peer percentile supersedes the prior. Modeled vs. observed basis is labeled on every output — a prior is never presented as an observed percentile.
v0.5
2026-05
Published the CXI headline composite — a client-weighted roll-up of the five indices into a single 0–10 board number. Default weights documented; per-client weights are set at onboarding and audited quarterly. Internal collection sub-signals are absorbed into their parent index upstream and are never reported as standalone indices.
v0.4
2026-05
Added supplier-graph overlap signal to RPS. Refined β weighting for CES high-severity records. Cross-listing identity attribution now reported with confidence bands.
v0.3
2026-03
Introduced peer-cohort benchmarks (N ≥ 5). Added uncertainty-band reporting standard across all five indices. Cross-vendor deduplication agreement check formalized.
v0.2
2026-01
Brand-leak fingerprinting added (document hashing + brand-string detection). Hash-rotation salt schedule formalized (90-day rotation).
v0.1
2025-11
Initial release with five indices: CES, IMI, RPS, BLV, BIR. Methodology baseline.
08·Red lines

What this methodology will never compute.

The collection boundary is part of the math. Outputs are only as trustworthy as the inputs that produced them — these rules bound the inputs.

Questions on the math?

We respond to methodology questions in writing.

CISO, head of risk, or anyone on the legal side — send the specific question and we'll respond inside 48h. First-cohort tier conversations get a longer-form methodology brief on request.

Request a methodology session
Independent audit

Built to be re-computed by a third party.

Every index output is anchored to (a) the methodology version, (b) the inputs at the time of computation, and (c) a SHA-256 hash of the evidence bundle. An auditor with read access can independently replay any historical value. Methodology audits are part of the first-cohort onboarding.