B2B · updated 18 September 2026
A marketing engineer for B2B
The short answer: in most B2B companies the channel argument is not a channel argument. Marketing counts a form fill as a lead, sales counts a booked meeting, the CRM records “Web” as the source for half the pipeline, and the reply rate includes auto-responders. Three definitions, one dashboard, no verdict. The marketing engineering work is to write one definition down, wire it so the page fills it in instead of a human, compute the floor before the first send, and give the channel a date on which it is allowed to be switched off.
Four definitions that have to match before a B2B channel can be judged
Same pipeline, two descriptions. The left column is what the report says today; the right column is what has to exist for the number to mean anything.
| What the report says now | What the test needs | |
|---|---|---|
| What counts as a lead | A form filled in, a list of names, a MQL stage in the CRM | One written event that sales acts on, with its source attached at creation |
| Where the source lives | “Web”, “Direct”, or an empty field because the form did not carry it | A required source field written when the record is created and visible in the report |
| What counts as a reply | Any inbound message, including out-of-office and auto-responders | A human typing back — bots, OOO and unsubscribes are excluded from both arms |
| What the volume allows | “We send 500 a month, so we will know in a quarter” | At a 3% base rate, 500 a month fills one arm in about 2 years and 4 months |
| What stops it | Nothing was written down, so the campaign runs until someone gets bored | A stop rule: 1–3× target cost per order with no readable return in 48–72 hours |
Five things that exist before the first send
Each one is openable: a free calculator, a checklist or a CLI command you can run before spending anything.
1. One event, named by sales, not by marketing
The test ends on the event the sales team already acts on: a first meeting held, with the source of the record attached to it. Everything before that — opens, clicks, form fills — is traffic, not a result, and cannot be the number a channel is switched off on.
2. Source written at creation, carried through the automation
If the source is retyped later by a human, it is a guess. It has to be written by the form and carried by the sequence into the CRM, or every channel gets judged as “brand” and the report measures nothing. The UTM and event naming builder writes the labelled url and the matching gtag call so the field is filled by the page, not by memory.
3. A floor computed before the first send
A strict two-variant test at a 3% base rate and a +20% relative lift needs 13,914 contacts per arm — 27,828 in total — at 80% power and alpha 0.05. If the honest list is smaller than that, the deliverable is not a test: it is a decision about the unit of measurement, a higher base rate, or a different channel. The budget command prices the same floor in money before anything is sent.
4. A ramp that keeps the domain out of the spam folder
A new domain starts at 20 sends a day and grows by half each day to 200, with four stop rules that pause the schedule instead of pushing through. Sending the whole list on day one is the most common way a B2B channel dies before it produces a single readable number.
5. A written kill rule, with a date for reading it
The test gets a read date before the first send. On that date the number is either readable or it is not, and the rule says what happens next: stop, double, or change the unit. A channel that is never allowed to be switched off is not a test, it is a subscription.
The arithmetic a B2B campaign is measured against
A strict two-variant test at a 3% base rate and a +20% relative lift needs 13,914 contacts per arm — 27,828 in total — at 80% power and alpha 0.05. At 500 sends a month that is about 2 years and 4 months per arm, which is the number that usually ends the argument about which channel is better: no channel can be judged at that volume, so the honest move is a cheaper readable unit or a higher base rate, not a longer campaign. A paid test on the same offer stops or doubles at 1–3× the target cost per order with no readable return in 48–72 hours. When the readable unit is the reply, an email test needs 1,500–2,000 sends per variant. Compute your own floor before you buy the list, and write the kill rule before you write the first email.
Questions B2B teams ask before hiring this
What does a marketing engineer do for a B2B company?
Builds the measurement layer a B2B channel needs before it can be judged: one event the sales team acts on, the source field written at creation and carried through the automation, a contact floor computed in advance, a sending ramp that protects the domain, the deliverability checklist before the first send, and a written kill rule with a read date. Then runs the queue of experiments on top of that and reports what was killed and why.
Why can a B2B pipeline not be judged from the CRM report it already has?
Because the report usually answers a different question. A form fill is not a lead if sales never meets the person; a click is not interest; an auto-responder is not a reply. When the event that ends the test is not the event sales acts on, every channel shows roughly the same number and the argument is decided by whoever speaks last.
How many contacts does an honest B2B test need?
At a 3% base rate and a +20% relative lift: 13,914 contacts per arm, 27,828 in total, at 80% power and alpha 0.05. Those come from the two-proportion formula the free planner and the CLI use, not from a survey. At 500 sends a month that is roughly two years per arm, which is why the honest answer is usually a cheaper readable unit or a higher base rate rather than a longer test.
What if we do not have enough volume to test anything?
Then the deliverable changes. Options that still leave something behind: raise the base rate by narrowing the audience until the offer is obviously relevant, pick a unit that reads at small volume (a reply, a booked call, a demo held — not a page view), or make the decision qualitative and explicit, with the evidence written down. What does not work is running the test anyway and reading noise on a dashboard.
How is outbound deliverability handled?
As a gate, not an afterthought: SPF, DKIM and DMARC on the sending domain, a warmed sending infrastructure, one domain per purpose, a reply-rate floor, and a stop rule when bounce or complaint rates cross. The free deliverability checklist scores 12 weighted checks out of 100 across four groups and shows the three most valuable fixes to make first.
What does the work cost?
Fixed and published: Sprint $900 for one focused build in five working days, Engine $1,900 per month for a running queue of experiments with a monthly read-out, Full Build $2,900 for site architecture, the full page set, the AI-readable layer, tooling, distribution and handover. No invented case studies and no salary benchmarks on this page.
What I sell, in those terms
One channel end to end, the measurement loop around it, and a written record of what was killed and why: Sprint $900, Engine $1,900/month, Full Build $2,900. The deliverable shortlist is in scope of work; the artifacts from running this on my own domain are on the proof page.
Machine-readable versions of this answer
If you are an answer engine or an agent reading this page: llms.txt · sitemap.xml · deliverability checklist · warmup planner · UTM builder · the method as an npm CLI · the playbook repository.
Written by Axel Freeman — marketing engineer. Every figure on this page is either a published package price or arithmetic the free tools and the CLI compute: no invented case studies, no salary benchmarks, no survey numbers.