Axel Freeman · Marketing Engineer

B2B data enrichment

A B2B data enrichment service is measured by rows that reach a person

Enrichment is usually sold as decoration: more columns, more fields, more confidence in a dashboard. The only field that changes an outcome is the one that reaches a human being, and it is the rarest one. Here is the arithmetic from the file this work is built on: 6955 company domains collected from 45 public sources on 19 September 2026, of which 1040 — 15% — publish an address of their own.

Updated 19 September 2026 · every number on this page is counted from the file, row by row · 6955 unique domains, 0 rows without a domain.

Which sources actually publish contact data

Enrichment starts earlier than most vendors admit: before a field can be appended to a row, the row has to come from a source that carries the field. Hiring threads give you the company. Directories give you the mailbox. The table is the yield of every family in the file.

Source familyDomainsPublished an addressYield
Hacker News “Who is hiring?” threads (2024 → Sep 2026)2 79230411%
Y Combinator open-source hiring directory73600%
remoteintech/remote-jobs public catalog53100%
Agency directories (DAN / Semrush listings)492492100%
Show HN launches — domain taken from the post URL4074411%
SaaSHub category listings (outbound link of the catalog)35100%
WordPress plugin directory homepage fields35093%
GitHub repository and organisation website fields466337%
dev.to posts whose canonical is a company domain2204018%
Job feeds (WeWorkRemotely / Jobicy / Arbeitnow / RemoteOK / Remotive)1826234%
npm  crates.io and package-registry homepage fields19000%
Reddit  Mastodon  Product Hunt and HN signal threads15743%
Total in the file6 9551 04015%

Two readings. 40% of the file comes from hiring threads, the family that gives the most companies and almost no mailboxes (304 addresses in 2 792 rows) — that is what an enrichment step is for. And the only family with near-total contact data is agency directories (492 of 492), because agencies put an address on their own site on purpose. A file built only from job feeds is a file of companies; enrichment is the work that turns the rest into something sendable.

What public sources can enrich — and what they cannot

Appendable from public evidence

  • Company identity: legal-facing name, root domain, the page the domain was taken from.
  • What they sell: product category, pricing page present or absent, self-serve sign-up present or absent.
  • Hiring signal: open roles and their titles — the cheapest honest intent signal that exists.
  • Public addresses: role and generic mailboxes published on their own domain.
  • Channel hints: whether they run a blog, docs, a changelog, a marketplace listing.

Not appendable, sold as if it were

  • Personal mobiles and direct dials — not public; anything sold as such is scraped, stale or invented.
  • Revenue and headcount — third-party estimates with no published method; treat as a filter, never as truth.
  • Intent and “in-market” scores — a vendor model, not an observation.
  • Technographics — a guess from headers and script tags, right often enough to be dangerous.

This is the whole reason the offer is priced per working arrangement rather than per record: the fields worth money cannot be bought in bulk, they have to be observed, one row at a time.

The three checks every row passes before it is enriched

1
The domain comes from the source, not from the text.

A structured field, a post URL, a package homepage, a catalog’s outbound link. Text mining is where the junk gets in: in the first pass on job feeds roughly half of the extracted domains turned out to be the board itself, an ATS (greenhouse.io, lever.co, ashbyhq.com) or a link inside the ad rather than the employer.

2
The host answers, and looks like a company.

Every candidate is requested live before it enters the file and kept only if the page carries product signals (pricing, sign-up, demo, contact). Without this step a file fills up with university course pages, documentation sites and parked domains. On the SaaSHub pull this pass took 909 candidates down to 351 usable rows.

3
The address belongs to the domain it sits on, then it is checked.

Domain match first, mailbox second. Mailboxes are SMTP-checked one by one; addresses that do not belong to the domain of the page are dropped rather than guessed. On one agency pull two addresses out of 35 were placeholder text and three belonged to another domain — five would have shipped as working data without this step.

Deduplication sits underneath all three: 45 sources produced 6955 rows and 6955 unique root domains, no duplicates and no row without a domain. Subdomains are folded into their root (app., shop., blog.) unless the root is a platform, in which case the candidate is dropped entirely.

What is delivered, and in what shape

The segment work, the three checks and the first sends are scoped as one job: $900 Sprint one-off, $1 900/month Engine for the running loop, $2 900 Full Build when the offer, the page and the distribution have to be built as well. Full scope and what each one stops at: pricing.

Start with the 100-contact test

Enrichment that pays for itself, or not

The honest version: for a segment of a few hundred companies, enrichment done by hand is affordable and accurate; for tens of thousands of rows it is a tooling problem, not a service. The work here is on the side where the row count is small enough that every row can be checked — which is also the side where a wrong field costs a reply. A strict two-variant test needs 13 914 contacts per arm before a rate can be read, so the first test is 100 verified contacts: it decides whether the segment is worth the volume before a month is committed to.

Related: B2B lead list building — where the 6955 domains and their source yield table came from · Appointment setting — the same rows used for their outcome · Outsourced SDR — the sending run as a service · for SaaS, for agencies, for local B2B.

What you can open before paying

The proof page lists the artifacts other people’s repositories accepted, the machine-readable offer carries the same scopes for agents, and the engagement page states what happens in the first week. Nothing on this page is a projection: the numbers are read out of the file it describes.

Questions asked about data enrichment

What does B2B data enrichment cost?
Two markets, two prices. Per-record enrichment runs from fractions of a cent to a few dollars per row, and covers appended fields, not outcomes. Working arrangements are priced per engagement: $900 Sprint (segment, checked rows, first touches), $1 900/month Engine (the loop keeps running), $2 900 Full Build (offer, page, tool, distribution). There is no per-1 000-records price here, because the row is the input.
Can you enrich a list I already own?
Yes, and the first thing that happens is a quality pass rather than an append: domains are folded to their root, rows without a domain are dropped, duplicates are collapsed, then the live check and the address-domain match run. Usually a third of a bought list does not survive that pass, and that is the finding, not a failure.
Do you guarantee a bounce rate?
No. Bounces depend on the sending domain, the volume per day and the age of the row. What is guaranteed is the method: SMTP check before the send, addresses matched to their own domain, and every bounce fed back into the file. A published number for a promise nobody controls is worth nothing.
Where do the 1 040 published addresses come from?
From the companies themselves: role and generic mailboxes on their own pages and in their own listings. No scraped personal mailboxes, no guessed patterns, no purchased databases. That is why the yield is 15% of the file rather than 100% — and why the rows that are there hold up.
Does the enriched file stay mine?
Yes. It is delivered in CSV and JSON with the source column intact from the first delivery. Data and mailbox costs are never inside the fee, and the same list is not resold to a second buyer.

Related reading: Marketing engineer vs agency — what enrichment work costs when it is inside an agency retainer · the Freeman framework — the order these checks run in.

Where this count was published

The same numbers, written for a publishing platform rather than for this page, are in "I enriched 6,955 company domains. Only 1,040 of them publish an address" on dev.to — the source-yield table, the appendable-versus-unobservable split, and the discussion on that copy.

Related: Agency partner program — the verification pass partners buy capacity for, and what it appends.

Related: Outbound pilot — the verification pass that runs before any address is written to a batch.