Table of Contents
- What Matters Most for Cold Email in 2026?
- Why Do Results Die in Month Two or Three?
- How Much Harder Is Microsoft Than Google Now?
- Does AI Personalisation Actually Work Yet?
- What Has Changed About How Long a Tactic Lasts?
- Conclusion
- Key Takeaways
- Key Terms Glossary
- Ready to Pressure-Test Your List Before Your Infrastructure?
- Related reading
For years the first thing we told anyone starting an outbound programme was to get the infrastructure right. That is no longer the advice. Cold email in 2026 is decided upstream of infrastructure, at the list, and the reason is mechanical rather than philosophical: build a list of people who can feel they should not be receiving this, and they mark it as spam, and the infrastructure you carefully built degrades anyway. Everything downstream inherits the quality of that first decision.
Danish Lead Co. sends more than half a million messages a month across client programmes, so what follows is drawn from live campaigns rather than from a survey. Four things genuinely changed this year, and one of them reverses advice we gave ourselves.
What Matters Most for Cold Email in 2026?
The list, and it is not close. The most reliable predictor we have found is uncomfortable to act on: the easier a list was to build, the worse it performs. A list pulled straight out of LinkedIn Sales Navigator, Apollo, or ZoomInfo with the obvious filters is, by definition, the same list every competitor pulled with the same obvious filters. The people on it are the most contacted people in the market. They are not ignoring you specifically, they are ignoring the fortieth version of the same message that quarter.
Three ways to build a list that was not trivially available:
- Scrape a different surface. Association directories, trade-show exhibitor lists, permit and tender registries, certification bodies, procurement portals. Slower to assemble, and not sitting in anyone else's export.
- Layer a timing signal. Recently hired or recently promoted is the single most transferable signal across the programmes we run. Someone new in role is measurably more willing to change a vendor or trial a system, and it takes roughly half as many messages to open a qualified conversation. The constraint is supply, only so many people change role in a given month, so it works as a high-intent layer rather than the whole list. The way we build and sequence those outbound systems treats it as one segment among several.
- Build the same list twice and compare. Pull the same definition from two or three databases and run them as separate segments. The overlap tells you who everyone is contacting, and the non-overlap is usually where the response is.
Why Do Results Die in Month Two or Three?
Almost always because the sending domains and mailboxes have decayed, not because the copy stopped working. This is the failure mode we are called in to diagnose most often, and it is invisible from the reply rate alone. If half the sending accounts have quietly degraded, half the intended volume never arrives, and the programme looks like a messaging problem while being an infrastructure problem.
The tell is timing. Copy fatigue degrades gradually and unevenly across segments. Infrastructure decay shows up as a step change roughly six to ten weeks after launch, across every segment at once, with open and reply rates falling together. Our walkthrough of diagnosing broken email deliverability sets out the checks in order, and sender reputation management covers the maintenance cadence that prevents it rather than diagnosing it after the fact.
How Much Harder Is Microsoft Than Google Now?
Substantially. Pulling a month of sends across client programmes recently, it took roughly five to ten times as many messages to produce a positive response from a Microsoft-hosted inbox as from a Google-hosted one. Above about 5,000 employees, where Microsoft hosting and layered filtering are close to standard, the gap widens further. Mimecast and Proofpoint we now exclude from target lists outright, because the cost per conversation behind those gateways does not justify the sending reputation it consumes.
| Receiving environment | Relative difficulty | How we handle it |
|---|---|---|
| Google Workspace | Baseline | Standard volume and sequencing |
| Microsoft 365, under 5,000 employees | Roughly 5x the messages per positive reply | Lower per-mailbox volume, longer sequences, tighter targeting |
| Microsoft 365, over 5,000 employees | Hardest mainstream environment | Multi-thread the account, treat email as one channel of several |
| Mimecast or Proofpoint gateway | Not economic | Excluded from the list at build time |
This is a targeting decision rather than a technical one. You cannot out-write a gateway. You can decide, before the first send, which environments your list is weighted toward, which is why gateway composition is now part of how we scope a programme at all. Our mailbox provider performance data breaks the same picture down by provider.
Does AI Personalisation Actually Work Yet?
It works when it produces something genuinely specific to the recipient, and it fails when it produces a fluent sentence that could have been written about anyone. The distinction is not the model, it is how much real research sits behind each message.
A working example from a client selling digital signage content: rather than referencing a hotel's recent award or its website copy, the system looks at each individual hotel and proposes three specific things that property could put on its screens. Getting the output accurate took about half a day of prompt work and validation. It is now the best performing campaign in that account. The same research, done by a team by hand, would have taken a week for a fraction of the volume.
The rule we work to:
- Research something that required looking. If the personalisation could have been generated from the company name alone, it adds nothing and the recipient can tell.
- Make it a proposal, not an observation. "I noticed you have 40 screens in the lobby" is a fact about them. "Here are three things you could put on them" is a reason to reply.
- Validate on a sample before scale. Read fifty generated messages yourself. Accuracy below about 95 percent is worse than no personalisation, because one wrong detail discredits the whole message.
- Budget the prompt time properly. Half a day of iteration is normal for a campaign that will run for a year. Treating it as a five-minute task is why most AI personalisation reads as filler.
What Has Changed About How Long a Tactic Lasts?
The window between a tactic working and a tactic being everywhere is now very short, because the tools that let us run something at scale are available to everyone at the same time. That does not make the channel worse, it changes what you are buying. You are no longer buying a play, you are buying the ability to find the next one quickly. A programme built around a single message that worked in January is a depreciating asset by June.
Practically, that means running a live test lane at all times rather than in occasional bursts, and measuring by segment so you can see which part of the market is decaying rather than concluding the whole channel is. Our 2026 state of B2B outbound covers the wider pattern, and our outbound benchmarks give the ranges to compare your own numbers against.
Conclusion
Cold email in 2026 still works, and it works for the same underlying reason it always did: a relevant message reaching a person who can act on it. What changed is where the leverage sits. Infrastructure is now a maintenance discipline you cannot skip rather than the thing that differentiates you, Microsoft-hosted enterprise inboxes have priced themselves into a different strategy, and AI personalisation pays only when it does research a human would have had to do. The list is the decision that determines all of it, and the harder that list was to build, the better it will perform.
Key Terms Glossary
Ready to Pressure-Test Your List Before Your Infrastructure?
If your programme has slowed and you are not sure whether the problem is the list, the messaging, or the sending infrastructure, book a call with Danish Lead Co. We will look at how your list was built, which receiving environments it is weighted toward, and what your reply pattern over time actually indicates, then tell you which of the three to fix first. You can see what a rebuilt programme looks like in practice in our manufacturing case study.