Six million companies, one million websites: anatomy of a nationwide data enrichment run
From a registry of six million Italian companies to the right website for one million of them — with phones, emails, WhatsApp contacts and social profiles extracted in a ten-day run. The real case, and the three hard questions a job like this has to answer.
Italy's chamber-of-commerce registry says who companies are and where they're based. It doesn't say how to actually reach them: the website, the phone numbers that answer, the working emails, the social profiles. A client needed exactly that gap closed — not on a sample, but across six million companies.
The numbers, up front
- 6,000,000 Italian companies in, from the business registry.
- 1,000,000 company websites identified and matched to the right company.
- 10 days for the full run, start to finish.
- Extracted per site: phone numbers, email addresses, physical addresses, WhatsApp contacts, social profiles (LinkedIn and others).
The three hard questions
The value of a job like this isn't in "downloading pages" — it's in the three questions such a system has to answer well, millions of times in a row.
1. Which website is the right one?
A company name is not a domain. Namesakes, directories posing as the company's site, resellers ranking above the manufacturer: matching a company to itswebsite is the most underrated problem in all of data enrichment — and getting it wrong means attributing someone else's contacts. One million matches is the number of times that question got answered.
2. How do you extract from a million different websites?
There is no template: every site puts its contacts wherever it likes — in the footer, on a contact page, inside an image, behind a WhatsApp link. Extraction has to recognize the data for what it is, not for where it sits: a phone number, an email, a LinkedIn profile only count if the system can identify them in the middle of everything else.
3. How does it survive ten days?
A nationwide run isn't a script you launch and hope: it's a process that has to survive unreachable sites, broken pages and interruptions, and resume where it stood instead of starting over. It's the same difference — the one I describe under automations on documents and data — between a one-off script and a system.
What a normal company does with this
Six million was that client's scale. The structure is the same at any scale: a list of companies — your prospects, your industry, your province — that today holds only company names, and tomorrow holds verified websites, phones, emails and profiles. If your CRM is full of silent records, this is the work that makes them talk.
Frequently asked questions
Where did the data come from?
Public sources: the chamber-of-commerce business registry as the starting point, and the companies' own public websites as the source of contacts. The job wasn't finding secret data — it was correctly connecting, at national scale, two public things that don't talk to each other.
Can this be done for our industry, or a smaller list?
Yes, and that's the most common case: a list of a few thousand companies — one industry, one province, the accounts in your CRM — gets enriched with the same structure and much shorter timelines. Scale changes the numbers, not the method.
Were the extracted contacts reliable?
Extraction at this scale always produces noise: dead fax numbers, generic inboxes, abandoned social pages. That's why extraction is only half the job — the other half is normalizing and cleaning, because a wrong record in a CRM costs more than a missing one. Same principle as the other cases: the output gets verified, not hoped for.
What does a job like this cost?
It depends on the scale and on what needs extracting. The order of magnitude gets defined in a call after seeing the starting list and a sample of what you need — half an hour, and the first thing I'll tell you is whether it's worth doing at all.
