Axel Freeman · Marketing Engineer

Lead list verification · the reject file

Every row we throw away, and the reason it was thrown away

A list is only as good as the rows that did not survive. This page publishes both sides: the checks that run before a single row is written to a segment, and the yield each family of sources actually produces when it is measured instead of claimed. The numbers below are counted from the working file at the moment this page was built — 7,728 company domains from 49 public sources, 1,381 of them publishing an address of their own (18%), and 1,350 of the 1,358 domains behind those addresses answering an MX lookup right now.

Updated 19 September 2026 · the working file is the same one the $900 pilot hands over; the reject file is handed over with it, reasons included.

The six checks

Nothing here is sophisticated. It is simply done in this order, and every step is allowed to shrink the list.

check 1

The domain has to come from a field, not from prose

A row enters only when the domain sits in a field the owner filled in: a registry's website, a package's projectUrl or homepage_uri, a repository's homepage, a launch post's canonical URL. Domains dug out of a job description or a sentence in a post are the main reason bad rows appear at all, so they are not used.

check 2

Folded to the root, then de-duplicated

blog., app., shop. and demo. subdomains collapse into one company, otherwise a single vendor arrives as four rows. After folding, 7,728 rows carry 7,728 unique domains and 0 rows have no domain at all.

check 3

The host is fetched live

A candidate that does not answer 200 on HTTPS and then HTTP is dropped, parked domains and expired projects included. This is the single cheapest filter and the one that removes the largest block of dead weight.

check 4

The site has to behave like a product company

At least two signals out of four: a pricing or plans page, a sign-up or free trial, a “book a demo” or “contact sales” path, and integrations or API documentation. This is what keeps a blog, a documentation site and an academic page out of a B2B segment.

check 5

The address has to belong to the domain it sits on

An address is kept only when its domain matches the site it was found on, or a subdomain of it. Agency directory pages publish third-party addresses and placeholders like hello@world.com surprisingly often; a row with someone else's address is worse than no row, because it is a complaint waiting to happen.

check 6

The domain has to answer for its mail

Last check, and the only one that requires the address to be real: an MX lookup on the domain. Of the 1,358 domains behind the 1,381 published addresses, 1,350 answer right now and 8 do not — those rows are marked and left out of any batch rather than counted as contacts.

What each family of sources actually yields

The same rule applied to different ecosystems produces very different lists, and that difference is the reason a list cannot be priced per thousand rows. It is counted from the file, not estimated from a sample.

Family of sourcesRowsPublished an addressYield
Agency and SaaS directories (DAN, Semrush, SaaSHub)87652260%
Hiring threads and job feeds (HN Who is hiring, Arbeitnow, Jobicy, Remotive)4,01243211%
Company directories (YC public dump)54826548%
Package registries (NuGet projectUrl, RubyGems homepage_uri, npm, crates.io)3285015%
Repository homepage fields (GitHub org website, repo homepage, topic)5535911%
Launch feeds and product posts (dev.to canonical, Show HN, Product Hunt)7458812%

Two observations that only exist because the numbers are published. First, directories built for selling — agency and SaaS catalogues — are the only family where the majority of rows arrive with a contact already on the page (60%), and hiring feeds are the opposite (11%), because a job post is written to attract candidates, not to be harvested. Second, the field is not enough by itself: repository homepages come from a structured field and still yield 11%, while package registries, the same idea one ecosystem over, yield 15% and agency directories 60%. A public field is a guarantee about the domain, not about the contact.

The reject file

What fails is not deleted: it is written out with the reason attached, so the count can be audited instead of trusted. A batch is built from the file, and anything not in it is not in the batch — no purchased lists are mixed in.

1

Rows dropped at sourcing: the domain was not in a structured field, or the host never answered.

2

Rows dropped at the product test: the site behaves like a blog, docs site or directory rather than a vendor.

3

Rows dropped at the mail test: no MX, or an address belonging to a different domain than the site.

A vendor selling “100% verified” lists is describing a promise about the file they sold you. What is describable is the other number: how much was thrown out, why, and which check did it. That number is on this page because it is the honest half of the one on the pricing page.

Questions this page gets asked

What exactly is in the reject file?

Three lists with a reason column each: rows dropped at sourcing because the domain was not in a structured field or the host never answered; rows dropped at the product test because the site behaves like a blog, documentation site or directory; and rows dropped at the mail test because the domain has no MX record or the address found on the page belongs to a different domain.

Why do you not just scrape and clean up later?

Because cleaning cannot recover what was never there. A domain taken out of a job description cannot be matched to a company, and an address found on an agency catalogue page frequently belongs to someone else. Each of those becomes a bounce or a complaint later, so the row is refused at the door instead.

How many of the published addresses are usable?

Of the 1,358 domains behind the 1,381 addresses in the working file, 1,350 answer an MX lookup and 8 do not. Rows without MX are marked and kept out of batches; they stay in the file so the client can see the size of the gap.

Is a verified list a guarantee of replies?

No, and any number that claims to be is about the sender's motivation rather than the list. Verification removes bounces and mismatched addresses; whether a segment replies depends on the offer and the market, which is why a pilot ends with a read-out counting replies rather than a promise made before it.

Can the checks run on a list I already own?

Yes. The same six checks run against an existing file and produce the same three reject lists, so an old list can be measured before it is sent again rather than after.

Related: how each segment is built, the enrichment pass and the fourteen-day $900 pilot that hands the working file over. For agencies the same checks run white-label under the partner programme; for SaaS and local teams the segment is sourced per market (SaaS, agencies, local B2B). The SMTP check behind the address column is published as an API, and what the machine has already produced sits on proof.

Where this was written up

The reasoning behind the six checks and the three reject lists is published as an article: The reject file: every row my B2B list threw away, and the check that killed it — with the yield table and the numbers this page counts.

After the list, the send

The six checks decide what may be sent. What happens to that segment at the mail server — authentication, tracking domains, MX and send caps — is written out separately in the cold email deliverability audit.

What the addresses are worth

Verification decides which rows survive; the source decides how many survive. The yield of six source families, and the 300-row walk that turned “contactless” rows into 169 addresses with MX, is counted on the source-yield page.