Research
Email authentication across EVM infrastructure
We measured eight properties of email authentication across 61 domains the EVM ecosystem depends on, using public DNS only. Overall DMARC enforcement is 85%, which is healthier than we expected.
The shortfall is not spread across the population. It is concentrated in the stratum that writes the node software everyone else runs, which sits at 54%. Every domain in the study publishing no DMARC record at all belongs to that stratum.
Correction, 2026-10-07, after first publication. The first version of this
page counted one entry as a domain with no DMARC record. That was wrong twice.
The entry is a subdomain, and under RFC 7489 a subdomain with no record of its own inherits
the organisational domain's policy; that parent publishes p=reject with no
sp override, so the subdomain was covered all along and never belonged in the
finding. The parent is also already in the population, so the same organisation was counted
twice.
We removed the row. Population 62 → 61. Overall enforcement 84% → 85%. Client teams 50% (7/14) → 54% (7/13). Fisher's exact p 0.00058 → 0.00183. Domains with no DMARC record at all, five → four. The separation between strata survives the correction and every number below is the corrected one. We found this ourselves while checking whether the domains in the study are in active use, and we are leaving the arithmetic of the change visible rather than quietly restating the totals.
Threat modelWhy this is worth measuring here
DMARC is the layer that tells a receiving mail server what to do with a message that claims to come from your domain and fails authentication. Without it, a receiver that gets a forgery has no instruction from the domain owner and falls back on its own reputation heuristics, which are opaque, inconsistent between providers, and outside your control.
Most organisations treat this as a deliverability problem. For one group in this population it is closer to a supply chain problem. Node operators learn about releases, deprecations and security advisories by email, and they act on that email by running software. A forged urgent patch available sent from a client team's own domain requires no vulnerability in any product. It requires a domain that never told receivers to reject forgeries, and a recipient whose job is to apply patches quickly.
That is why we split the population by role rather than by size. If the shortfall were uniform, the finding would be a generic adoption statistic. It is not uniform.
MethodWhat we queried, and what we deliberately did not
Four strata: RPC and node infrastructure providers, protocol foundations, client implementation teams, and the security firms that audit the ecosystem. Sixteen domains each, except client teams at thirteen, where the population of independent teams with their own domain is simply smaller.
For each domain we resolved, over DNS-over-HTTPS against a public resolver:
- the apex
TXTrecords, for an SPF policy, itsallqualifier, and itsinclude:targets - the full SPF evaluation tree, following every
include:andredirect=recursively, counting the DNS-querying mechanisms against the limit of ten that RFC 7208 imposes _dmarc, parsed forp,sp,pct,ruaand the two alignment modes, rather than for the policy word alone_mta-sts, for a published transport security policy_smtp._tls, for TLS reporting- the
DSrecord, for whether the zone is DNSSEC-signed MX, to establish whether the domain receives mail at all
Nothing was sent to any of these domains. No mail, no SMTP connection, no HTTP request. We did not fetch the MTA-STS policy file itself, which is served over HTTPS, so we can report that a domain publishes a policy but not whether that policy is in enforce or testing mode. That is a real gap in the measurement and we chose it on purpose, because it is the only way to keep the whole study to public DNS reads.
Result 1The gap is in one stratum, and it widened when we added data
| Stratum | n | Enforcement | reject | quarantine | p=none | no record |
|---|---|---|---|---|---|---|
| Security and audit firms | 16 | 100% | 10 | 6 | 0 | 0 |
| Protocol foundations | 16 | 94% | 12 | 3 | 1 | 0 |
| RPC / node infrastructure | 16 | 88% | 8 | 6 | 2 | 0 |
| Client implementation teams | 13 | 54% | 6 | 1 | 2 | 4 |
| All | 61 | 85% | 36 | 16 | 5 | 4 |
An earlier run of this study covered 46 domains and put client teams at 70%. Adding domains to each stratum moved that figure to 54%.
We expected the security firms to be the weak stratum. Auditing other people's systems and administering your own are different jobs, and the second one tends to lose. They are the only stratum with no exceptions: sixteen of sixteen at enforcement, ten of those at p=reject.
The direction of the change when we expanded the population is worth more than the figure itself. Enlarging a sample usually pulls an outlying stratum toward the mean. This one moved away from it, from 70% down to 54%, which is weak evidence that the first measurement was understating the gap rather than inventing it. It is weak evidence and not proof: the four domains we added to that stratum were not chosen at random, they were the next independent client teams we could identify with their own domain.
Does the gap survive a sample this small?
The whole claim of this study is a difference between groups, and the groups are small enough that one domain moves a stratum by six or seven points. That deserves a test rather than an assertion, so here are 95% Wilson score intervals on each stratum:
stratum k/n point 95% interval Security and audit firms 16/16 100% [ 81% , 100% ] Protocol foundations 15/16 94% [ 72% , 99% ] RPC / node infrastructure 14/16 88% [ 64% , 97% ] Client implementation teams 7/13 54% [ 29% , 77% ] All 52/61 85% [ 74% , 92% ]
The interval for client teams runs from 29% to 77%, which is wide, and the interval for security firms starts at 81%. They do not overlap. The individual point estimates are imprecise and the separation between the two ends of the population is not.
Comparing the client stratum against the other three combined gives 7 of 13 against 45 of 48. Fisher's exact test, two-tailed, returns p = 0.00183.
That number should be read for exactly what it is. It answers one question: if enforcement were independent of which stratum a domain belongs to, how often would a split this extreme arise by chance? Roughly twice in a thousand. It does not establish that the result generalises to every client team in existence, because this population was assembled by us and not drawn at random, and no test can repair a sampling frame. The honest reading is that within these 61 domains the association is real and not an artefact of small numbers, and that whether it holds beyond them is an open question this study does not answer.
Result 2Four domains publish no DMARC record at all, and all four are client teams
Nine of the sixty-one have no enforcement. Five of those publish p=none, which is monitor-only: receivers send reports and are told to take no action. That is the correct first step for a domain still discovering which of its own senders fail alignment, and it is a position to pass through rather than to settle in.
The other four publish no DMARC record whatsoever, and the concentration is total:
stratum p=none no record Security and audit firms 0 0 Protocol foundations 1 0 RPC / node infrastructure 2 0 Client implementation teams 2 4
Not one RPC provider, foundation or security firm in this population is missing the record. Every absence is in the stratum whose software the other three run.
Having no MX does not protect you
Two of those four domains have no MX record, so they receive no mail at all. It would be easy to read that as harmless, and the reasoning does not hold. SPF and DMARC govern what a receiver does with a message claiming to come from a domain. Nothing about that depends on whether the domain can receive a reply. A domain that accepts no mail can still be impersonated to anyone, indefinitely, and the impersonation is arguably easier to sustain because no bounce ever arrives to alert the owner.
One domain in the RPC stratum demonstrates the correct handling of exactly this case: no MX, and nonetheless -all on SPF with DMARC at p=reject. The owner understood that a domain which sends and receives nothing still needs to say so.
Strict SPF without DMARC is weaker than it looks
One of the four has a hard-fail SPF policy, -all, and no DMARC record. From a glance at the SPF record alone that domain looks well configured, and the protection is much thinner than it appears, for a reason worth stating precisely.
SPF authenticates the envelope sender, the address in the SMTP MAIL FROM command. It does not authenticate the From: header, which is the only address a person ever sees. An attacker can pass SPF cleanly for a domain they control while putting any From: address they like in the message. The mechanism that closes that gap is DMARC's alignment requirement, which insists the domain that passed SPF or DKIM match the From: domain. Without a DMARC record there is no alignment requirement, and -all is protecting a field the recipient never reads.
Result 3A policy word is not a policy
One domain in the study publishes p=reject with pct=70. The pct tag tells receivers to apply the policy to that percentage of failing messages and to fall back to the next-weaker treatment for the rest. Thirty percent of forged mail from that domain is therefore exempt from the reject it advertises.
Any survey that reads p= and stops will record that domain as fully enforcing, and ours would have, which is why we parsed the whole record. It is one case out of sixty-two, so it changes no aggregate here. We are reporting it because it is the kind of detail that separates a measurement from a count, and because the ratio is a deliberate rollout setting that someone meant to raise later and did not.
Result 4Four domains are within two lookups of silently losing SPF
SPF evaluation has a hard ceiling. RFC 7208 limits a receiver to ten DNS-querying mechanisms per evaluation, counting every include:, a, mx, ptr, exists: and redirect= encountered anywhere in the tree, including inside third-party records you do not control. Exceed it and the result is permerror, which most receivers treat as a failure of SPF for every message, not for some of them.
lookups domains
0 # 1
1 ############## 14
2 ####### 7
3 ############# 13
4 #### 4
5 ######## 8
6 #### 4
7 ## 2
8 ### 3
9 # 1
10 limit — permerror beyond this
Fifty-seven of the sixty-two publish an SPF record, and none of them exceeds the limit today. Four sit at eight or nine. For those four, one more marketing tool, applicant tracking system or support desk added to the record can push the evaluation past ten, and the failure mode is quiet: no error surfaces to the sender, mail simply begins failing SPF everywhere at once.
The count is not fully under their control either, because it includes mechanisms inside the records of third parties they have delegated to. A vendor restructuring their own SPF record can consume a domain's remaining headroom without telling anyone.
Result 5Most of these policies are published in an unsigned zone
| Property | Domains | Share | What its absence means |
|---|---|---|---|
DNSSEC (DS present) | 21 / 62 | 34% | DNS answers for the zone are not authenticated |
| MTA-STS | 9 / 62 | 15% | Inbound TLS is opportunistic and downgradable |
| TLS-RPT | 7 / 62 | 11% | TLS negotiation failures go unreported |
The DNSSEC figure deserves more than an adoption percentage, because SPF, DKIM and DMARC are all published as DNS records and inherit whatever integrity DNS has. An adversary positioned to tamper with DNS responses between a receiver and the authoritative server can strip a DMARC record from the answer. The receiver then behaves exactly as it would for a domain that never published a policy. In an unsigned zone, there is nothing in the protocol for the receiver to notice.
Thirty-two of the fifty-two domains at enforcement publish that policy in an unsigned zone. This is a conditional weakness and not a live one: it requires an attacker already able to manipulate DNS for the target resolver, which is a substantial position. It is worth naming because the usual mental model treats a DMARC record as a durable statement, and in an unsigned zone it is a statement that holds only as far as the resolution path does.
There is also a correlation in the data, which we report as a correlation. Domains with DNSSEC reach 95% enforcement; those without reach 78%. The plausible reading is not that signing causes policy but that both are produced by a team that administers its DNS deliberately. With 21 signed domains the difference is not strong enough to carry weight on its own.
Result 6Concentration, and the senders these domains have authorised
largest hosted provider 53 87% no MX 4 7% second hosted provider 2 3% self-hosted or other 1 2% privacy-focused provider 1 2%
Eighty-seven percent of the population receives mail through a single hosted provider. That is one point of correlated failure for the inbound mail of nearly every organisation in this study, and it is also why the SPF records look the way they do: that provider's own include appears in fifty-two of them, and brings its sending estate with it.
Below that first include, each SPF record is a list of third parties the domain has granted permission to send on its behalf. Across the population there are 49 distinct third-party senders, and the overlap is where it becomes interesting:
cloud transactional sender A 7 domains CRM platform 6 support desk platform 6 marketing automation A 6 cloud transactional sender B 6 marketing automation B 5 newsletter platform 5 cloud transactional sender C 4 recruiting platform (ATS) 4
Each line is an authorisation that outlives the campaign it was added for. A compromise or an abuse failure at one of these shared platforms does not reach one domain, it reaches every domain that still lists it, and the affected organisations have no shared visibility into that fact.
The recruiting platform appearing in four records is a specific case worth drawing out. Recruiting mail is an unusually effective phishing pretext against engineering organisations, it is expected to arrive from an unfamiliar sender, it invites the recipient to open an attachment or follow a link, and in four of these domains it is sending with the domain's own authorisation.
Negative resultsThree things we went looking for and did not find
We ran three further checks hoping to turn posture observations into something exploitable. All three came back empty, and they are here because a study that only reports what it found is reporting half of itself.
Dangling SPF includes: 0 of 49
Every include: in an SPF record delegates sending authority to a third party. If one of those targets stops resolving, or the domain behind it lapses and can be registered by someone else, the delegation becomes an avenue rather than a convenience. We resolved all 49 distinct include targets across the population. Every one returns a valid SPF record. No dead delegations, no registration candidates.
Dangerous SPF mechanisms: none
We checked every top-level record for +all, which authorises the entire internet to send as the domain, for ?all, which is neutral and therefore close to meaningless, and for the deprecated ptr mechanism. None of the 62 publishes any of them. One record terminates with no all mechanism at all, which is a defect of a different and much smaller kind.
Over-broad IP ranges: present, and not a misconfiguration
Expanding every SPF tree to its leaf IP ranges shows most of the population authorising a /16, which on its face looks alarming: roughly a hundred thousand addresses permitted to send. It is inherited, not chosen. The range belongs to the mail provider that 85% of this population uses, and it arrives through that provider's own include. Every domain using a large hosted mail provider authorises that provider's entire sending estate, and there is no version of using such a provider that does not.
What it does change is the reading of the case in Result 2 with hard-fail SPF and no DMARC. That record authorises on the order of 98,000 shared addresses, so its -all asserts that one of ninety-eight thousand shared machines sent the message. The mechanism that would convert that into a statement about the sender is DMARC alignment, and it is the one that is missing.
DataThe measurements, in full
All 61 rows with all eight properties: data.json. Every aggregate in this article recomputes from that file, and we checked that it does before publishing.
Identifiers in the file are pseudonymous and stable, following the same withholding rule applied to the article: the redaction covers every row rather than only the weak ones, so nothing is identifiable by elimination. The file will be republished with domain names once the affected parties have been notified.
LimitationsWhat this study does not establish
DKIM was not surveyed, and could not be. DKIM public keys live under a selector name that is not enumerable from DNS. A survey can confirm a specific selector if it guesses the name, and cannot establish absence. Any domain here may have working DKIM we did not see, which matters most for the p=none cases, where DKIM alignment may already be passing and the owner may simply not have raised the policy yet.
MTA-STS mode is unknown. We read the DNS record and did not fetch the policy file, so a domain counted as publishing MTA-STS may be in testing mode, which enforces nothing. The 15% figure is an upper bound on real transport enforcement.
A published record is not an observed behaviour. Everything here is what a domain owner asked receivers to do. We did not test what any receiver actually did, and large receivers apply their own reputation signals regardless of policy. A domain at p=none is not thereby spoofable in practice. It has declined to state what should happen, which is a smaller and different claim.
The population was assembled by us, and that is the weakest part of this study. We assembled the list by role; it is not drawn from a reproducible rule, so a different author would build a different list and get different percentages. The next version of this work should define the population by an external criterion anyone can re-derive, which is the single change that would move it from a measurement to a replicable one.
The strata are not equal. Sixty-one domains chosen by us to represent four roles, with the client stratum at thirteen because the set of independent teams with their own domain is genuinely smaller. This is not a census. A different list moves every percentage, and the per-stratum figures rest on sixteen observations or fewer, so a single domain moves a stratum by six points.
We did not establish intent or impact. Nothing here shows that any domain has been impersonated, that anyone attempted it, or that an operator was ever deceived. The threat model in the second section is an argument about why this population matters more than most, not a finding.
One point in time. Measured 2026-10-07. Any value here can change with one DNS edit, which is the useful part.
VerificationReproduce any row in under a minute
Every number comes from public DNS. The two queries behind the main table:
curl -s "https://dns.google/resolve?name=_dmarc.example.com&type=TXT" curl -s "https://dns.google/resolve?name=example.com&type=TXT"
And the three that most surveys leave out:
curl -s "https://dns.google/resolve?name=_mta-sts.example.com&type=TXT" curl -s "https://dns.google/resolve?name=_smtp._tls.example.com&type=TXT" curl -s "https://dns.google/resolve?name=example.com&type=DS"
The SPF lookup count is the one figure you cannot read off a single response, since it requires walking every include: and redirect= to its leaves and summing the DNS-querying mechanisms along the way. If a result you get disagrees with one here, our measurement is wrong and we would like to know: [email protected].
No organisation is named in this study. Not the domains behind any result, and not the third-party platforms in the authorisation graph above, which appear by service category. Those platforms are not at fault for being listed; the delegation is the customer's decision, and naming them would attribute something that is not theirs. The names are withheld until the affected parties have been notified, and the rule is applied to every finding rather than to some of them, so the omissions do not themselves identify anyone. Every underlying value is a public DNS record, and the queries above reproduce the full table in a few minutes.
RemediationIf one of these is yours
The sequence is well worn and the whole of it is DNS. Publish DMARC at p=none with a rua address and read the aggregate reports until every legitimate sender of yours passes alignment, which normally takes a few weeks and normally surfaces at least one forgotten system. Then move to p=quarantine, confirm nothing broke, then p=reject. Leave pct alone unless you are deliberately staging, and if you do stage it, put a date on raising it.
If your domain sends no mail and receives none, publish v=spf1 -all and DMARC at p=reject anyway. It is two records and it closes the domain to impersonation permanently.
Count your SPF lookups before adding the next vendor rather than after, and if you are near the ceiling, flattening or a dedicated subdomain for bulk senders buys the headroom back.