Lead list building · updated 19 September 2026

B2B lead list building: what 6,032 companies from public sources taught me

The short answer: a B2B lead list is not judged by its size, it is judged by one number — records that reach a person, per hundred collected. Everything else (total rows, sources, “verified” in a vendor’s spreadsheet) is an input. Below is the list this page was built from: 6,032 company domains collected from 41 public sources in one working day, of which 998 published an address of their own. The table shows the yield of every source family, then the three rules that keep a public list honest, then the fixed-scope version of the same work: $900 Sprint, $1,900/month Engine, $2,900 Full Build — the first test is 100 verified contacts, so the list is judged before a month is committed to.

Where 6,032 company domains actually came from

Every row in the file carries the source it came from and the domain it was checked against. Nothing is estimated: the numbers below are counted from the file itself, one row at a time.

Source familyDomainsPublished an addressYield
Hacker News “Who is hiring?” threads (2024 → Sep 2026)2 79230411%
Y Combinator open-source hiring directory73600%
remoteintech/remote-jobs public catalog53100%
GitHub repository homepage fields (topic search)33000%
Agency directories (DAN  Semrush listings)52552299%
Show HN launches — domain taken from the post URL3434413%
dev.to posts whose canonical is a company domain22000%
npm and crates.io package homepage fields19000%
Job feeds (WeWorkRemotely  Jobicy  Arbeitnow  RemoteOK  Remotive  LaunchingNext)18200%
Reddit  Mastodon and HN signal threads14200%
Product Hunt launches3700%
Total in the file6 03299817%

Two readings matter more than the total. First, 66% of the domains come from one family (hiring threads), and that family publishes almost no mailboxes — 304 addresses out of 2,792 rows, because the post gives you the company, not the contact. Second, the only family with near-total published contact data is agency directories: 522 of 525 rows carry an address, because agencies put one there on purpose. A list built only from job feeds is a list of companies; a list built only from directories is a list of competitors.

Three rules that keep a public list honest

1. The domain comes from the post, not from the text. On job feeds the employer’s domain sits inside the body copy, and roughly half of the extracted rows turn out to be the board itself, an ATS (greenhouse.io, lever.co, ashbyhq.com), a link inside the ad, or the same company under two ccTLDs. When the domain is taken from a structured field or from the post URL — Show HN, package registries, dev.to canonicals — the reject rate drops from about 50% to 5-12%. That is why the file counts 41 sources and not 4,000 scraped pages.

2. A domain is not a company until it answers. Every host is requested live before it enters the file; only pages that look like a product (pricing, sign-up, demo, integrations) are kept. Without that step the file collects university course pages, documentation sites and personal blogs, which is what happened on the first pass — 21 candidates became 10 usable rows.

3. A published address is not a verified address. Of 6,032 rows, 998 expose an address; the rest are either SMTP-checked one by one or left out of the send queue. In the agency pull, two addresses out of 35 were placeholder text (hello@world.com) and three belonged to a different domain than the page they sat on — all five would have gone out as working data without a domain match.

What the list is for, and who buys it

The list is the cheap half of outbound. The expensive half is the sentence that earns a reply, and the reading of what came back. That is why the offer is scoped as work, not as a spreadsheet:

Three kinds of buyer get the most out of it: SaaS and seed-stage teams that have a product and no pipeline, agencies that want the list and the sending run white-label under their own brand, and local and regional B2B where the list is small enough to be checked by hand. The artifact trail for each is public: for SaaS, for agencies (white-label), for local B2B.

What you can open before paying

The list work is not the only thing this page can show. The proof page carries the third-party artifacts that were accepted by other people’s repositories, the machine-readable offer carries the same scopes for agents rather than people, and pricing keeps the three numbers in one place without a form. If the first 100 verified contacts are the only thing you want to buy, that is the Sprint.

See the engagement Write on Telegram

Where this count was published

The same numbers, written for a publishing platform rather than for this page, are in "I collected 6,032 B2B company domains in a day. Only 998 publish an address" on dev.to — the yield table and the three source rules, with the discussion comments on that copy.

Questions asked before buying a list

What does a B2B lead list cost?
Two prices exist in the market and they are not comparable: a price per record (a few cents to a few dollars, seller-side verification only), and a price per working arrangement. My own numbers are published: $900 for the Sprint that ends with a list and first touches, $1,900/month for the running loop, $2,900 for the full build. There is no per-1,000-records price, because the list is the input, not the product.
How is a built list different from a scraped email dump?
By where the domain came from and what was checked after. Four things are recorded for every row: the public source it came from, the live check of the domain, the domain match between the page and the address, and the SMTP check of the mailbox. A dump has none of the four.
How many records do I need before the list is readable?
A strict two-variant test needs 13,914 contacts per arm — 56 days at 500 contacts a day. Below that volume the honest output is not a conversion rate, it is a reading list: which segment answers, which sentence gets a reply, which objection comes back first. The 100-contact first test exists so that this is decided on evidence rather than on a spreadsheet.
What happens when a mailbox bounces?
It is removed from the queue and its domain is re-checked, not re-guessed: a dead mailbox on a live domain usually means a role address changed, not that the company is wrong. Bounce handling is part of the running loop, not a separate line item.
Who owns the list and the domains?
You do. The file is yours from the first delivery, in a plain format with the source column intact, so a different sender, a different tool or a different agency can pick it up. Data and mailbox costs are never inside my fee, and I do not resell the same list to a second buyer.

Related: Appointment setting — the same list used for its outcome: the four checks before a send, and how much volume a meeting number needs.

Related: B2B data enrichment — the enrichment pass in detail: what public sources can append, what they cannot, and the three checks each row survives.