The best cold email template runs under 90 words and follows four beats: a verifiable observation about the prospect’s company, the operational problem that observation usually signals, one concrete outcome you have produced for a similar company, and a single low-commitment question. Template structure sets the floor. The data feeding it sets the ceiling.
The best cold email template runs under 90 words and follows four beats: a verifiable observation about the prospect’s company, the operational problem that observation usually signals, one concrete outcome you have produced for a similar company, and a single low-commitment question. Template structure sets the floor. The data feeding it sets the ceiling.
A cold email template is a reusable message skeleton with variable slots that get filled from your data, one message per prospect. The skeleton is cheap. The variables are where the money sits.

Buyer behavior explains why. Gartner’s research on the B2B buying journey found that buyers spend only about 17% of their total purchase time meeting with potential suppliers, and when they are comparing several vendors, any single sales rep may get around 5% of that time. Gartner has also reported that roughly three quarters of B2B buyers would prefer a purchase experience with no rep involved at all. Your email is competing for a sliver of attention from someone who would rather not talk to you yet.
That reframes the job of the template. It is a relevance test the reader runs in about three seconds: does this person know something specific and true about my situation? Every word that fails that test costs you. Long paragraphs, credential dumps, and “I wanted to reach out” openers all fail it. Our breakdown of AI email lead generation covers where automation helps with this and where it quietly makes it worse.
Four lines, in this order.

Here is the shape, filled in for an illustrative case. Say you sell a data quality platform to Series B companies, and your signal is a job posting for an analytics engineer.
Subject: your analytics engineer req
Body: Saw you opened an analytics engineer role last week, second data hire this quarter by my count. At that point most teams are still fixing pipeline breaks by hand, so the new hire spends their first two quarters on cleanup instead of modeling. We put automated checks on the ingestion layer for a 180-person fintech and cut their broken-model tickets by about 60% in six weeks. Worth a look before the req closes?
That is 78 words. The observation is checkable. The implication shows you understand what happens after the hire, which is the part the reader is actually worried about. The proof carries a number and a timeframe. The ask is a yes or no.
Three details matter more than they look. Use the reader’s vocabulary, not your category’s. Cut every sentence that starts with “I” or “we” until the proof line. And send plain text with no images, no tracking pixel where you can avoid it, and no attachment, all of which drag inbox placement down.
There is no single winner across segments. Pick the structure your data can actually support.
| Template type | Best for | Data you need | Honest tradeoff |
|---|---|---|---|
| Trigger-based | Mid-market and enterprise, where change is publicly visible | Job postings, funding events, tech install changes, leadership moves | Trigger volume caps your list. In any given month most of your ICP has no live trigger, so this cannot be your only motion. |
| Problem-first | Tight, homogeneous segments that share one workflow | Accurate firmographics plus real role mapping, not job title guessing | Reads generic the moment the segment is drawn too wide. Requires discipline about who you exclude. |
| Peer proof | Categories with recognizable reference customers | Named accounts you have permission to cite, matched to the prospect’s tier | Weak when you are early and have no comparable customers to name. Do not stretch it. |
| Insight teardown | High ACV, a target list in the low hundreds | Manual or analyst-grade research per account | Expensive per email. Only defensible for tier-one accounts where one meeting pays for the effort. |
Most teams that scale outbound well run two of these at once: a problem-first template as the always-on baseline, and a trigger-based template layered on top of it that fires when a signal appears. Intent data can feed that trigger layer, though the accuracy varies a lot by vendor and category, which we get into in our look at what buying intent data actually changes.
Two causes, and they need different fixes.

Deliverability decay. Since 2024, Google and Yahoo have enforced bulk sender requirements: authenticated sending with SPF, DKIM, and DMARC, easy one-click unsubscribe, and a spam complaint rate kept below 0.3%. Practically, you want complaints under 0.1% and hard bounces under 2%. Miss those and your best template lands in a folder nobody opens. Check inbox placement before you blame copy, because a template that “stopped converting” has usually stopped arriving.
Signal exhaustion. The other cause is that you have already emailed everyone who matched the signal. A trigger-based template running against a 4,000-account list burns through its qualified subset in about a quarter. The fix is a new signal, not new adjectives. Teams that keep outbound compounding tend to add one new data source per quarter rather than rewriting subject lines weekly.
Manual research produces the best emails and the worst economics. A good rep can research and write maybe 25 genuinely personalized emails a day. At Seed to Series B headcount, that math does not reach a pipeline number.
The workable middle is programmatic personalization: build a data layer that resolves each account’s signals into structured fields, then map those fields into template slots. Tools like Clay are built for exactly this, chaining enrichment providers and running per-row research so the observation line writes itself from real inputs. The honest tradeoff is cost and complexity. Credits add up fast, and a poorly designed table burns budget enriching records you were never going to contact. Sequence your providers cheapest-first, which is the logic behind the data enrichment waterfall, and gate every expensive step behind a qualification filter. If you want the workflow built and maintained rather than assembled internally, that is what our Clay implementation work covers.
Set expectations on the AI layer honestly. Generated first lines are reliable when they summarize a structured field you already trust and unreliable when they freestyle from a scraped homepage. We mapped where that line falls in our piece on what AI BDRs actually automate and where they break.
Reply rate alone will mislead you. Track these instead:
Run tests at the segment level, changing one variable at a time, and give each variant at least 400 contacts before you read anything into the result. Below that, you are reading noise.
Build in-house when you have someone who owns outbound systems full time and your motion is a single segment with a stable signal. That person can maintain a data layer, a sequencer, and a CRM sync without it becoming anyone’s side project.
Bring in outside help when you are running three or more segments, when your data sits across tools that do not talk to each other, or when the person currently maintaining the system is your VP Sales doing it at 11pm. The failure mode we see most often is a strong template attached to a data layer nobody owns, which quietly degrades until outbound gets written off as a channel. That systems layer is the substance of GTM engineering, and how it connects to the rest of the motion is covered in our overview of AI-assisted outbound.
Between 50 and 90 words for a first touch. Long enough to carry an observation, an implication, and a proof point, short enough to read fully on a phone without scrolling. Follow-ups should be shorter, often under 40 words.
For a well-targeted list with clean deliverability, 5% to 10% total reply rate and 1% to 3% positive reply rate is a reasonable band. Anything above that usually reflects an unusually tight segment or a strong existing brand. Anything below 1% positive points at targeting or inbox placement before copy.
Yes, when the AI is summarizing structured data you already trust into a sentence. It performs poorly when asked to invent relevance from a scraped website, because readers recognize generated flattery immediately. Use AI for the assembly step and human judgment for the targeting logic.
Three to four touches over two to three weeks covers most of the available response. Each follow-up should add new information, such as a different angle, a relevant resource, or a second signal you spotted. Repeating “just bumping this” burns the account and raises complaint risk.
Use one structure across segments and different variables inside it. The four-beat skeleton travels well. The observation and implication lines need to be rebuilt per segment, because what counts as a meaningful signal for a 40-person startup differs completely from what matters at a 2,000-person enterprise.