What data enrichment is
At Graph, we help complex-B2B sales teams turn incomplete CRM records into reliable account and contact data, so they can prioritise opportunities and follow up with the context they need.
It is the practice of capturing, matching, appending, verifying, activating and refreshing account and contact records. Done properly, a commercial or sales leader can trust what the CRM says about a company before acting on it.
That is a different claim from the one most vendors make. Data enrichment is not a list you buy, a connector you switch on, or a once-a-year cleanup before a board report.
It is an operating system for identity data: a defined set of stages, each with an owner and a control point.
A partial record, a name, a work email, maybe a company, becomes something a rep can act on, a marketer can segment on safely, and a scoring model can weight correctly.
For complex-B2B teams, manufacturers, industrial suppliers, financial-services firms and anyone selling into organisations with real hierarchy, this matters more than it does for a simple SaaS lead form.
A single company might be a holding group with a dozen sites, each running different applications, each with its own buying group. Person-and-company fields alone cannot describe that reality.
What follows is the field model, the pipeline and the governance boundary that make an incomplete record into evidence a rep can act on.
Data capture, cleansing, enrichment and verification are different
These four operations get treated as one job on most teams. That is where the trouble starts.
Data capture is collection: a form fill, a badge scan at a trade show, a sales call note, an imported list. Capture creates the first record, and it should also create the first piece of provenance, recording where the data came from.
Data cleansing standardises and de-duplicates what you already have. It fixes formatting, merges obvious duplicates and removes clearly invalid entries. It does not add anything you did not already know.
Data enrichment adds new, governed information to a matched record: firmographic, site, application or buying-group detail, appended from an external or internal source and recorded with its origin.
Verification checks that captured or appended data is still correct and still permitted to use. An email still deliverable, a job title still current, a consent state still valid. Verification is what keeps enrichment from decaying quietly into fiction.
Blur these four together and records end up both enriched and wrong at the same time. New fields get added without anyone checking whether the underlying identity match was right in the first place.
Separating them is what lets you diagnose which one is actually missing, before you buy a fix for the wrong problem. The next question is what, specifically, complex-B2B teams should be enriching.
What data complex-B2B teams should enrich
Generic enrichment guides describe person-and-company fields: name, title, company, employee count, industry. That is a reasonable start for a SaaS lead form.
It is not sufficient for an account with sites, facilities, applications and a buying group spread across functions. The fields that actually drive a complex-B2B decision live below the company level.
A concise data enrichment field map
Rather than a downloadable worksheet, here is the practical field set worth enriching for complex-B2B accounts, organised by what each field is actually for.
Account and hierarchy data
Account and hierarchy fields describe the corporate structure behind a contact: legal entity name and registration, parent, subsidiary and division structure, headquarters and operating geography, employee count and revenue band, industry and sub-sector classification, and the account's position in any group or franchise hierarchy.
This tells you whether you are looking at one buyer, or one node in a much larger structure.
Site, facility and application data
Site and facility fields describe where the work actually happens: individual site or facility identifiers, site function (manufacturing, distribution, R&D, administrative), what is actually produced or processed there, the equipment or application in use where relevant, and site-level headcount or capacity where it is obtainable.
Generic enrichment tools rarely model this layer at all. That is precisely why it needs deliberate treatment here.
Contacts, roles and buying groups
Contact fields describe the individual person and their place in the decision: name, verified work email, job title, function and seniority, reporting line where available, buying-group role (economic buyer, technical evaluator, user, influencer, blocker), and which site or facility the contact is actually attached to.
A contact record without a buying-group role and a site attachment is barely more useful than a business card.
Commercial context, provenance, confidence and freshness
Commercial context fields describe why a record matters right now and how much to trust it: trigger events (expansion, acquisition, leadership change, relevant signal), existing relationship or opportunity status, and, for every appended field, its source, the date it was captured or last verified, and a confidence rating.
Provenance and freshness are not optional extras. Without them, neither a scoring model nor a rep can tell whether a field is current fact or six-month-old guesswork.
That trust problem is exactly what the operating pipeline below is built to close.
How data enrichment works from capture to refresh
Enrichment runs as a pipeline with five distinct control points.
Skip any one of them, and that is usually where trust breaks down later.
Capture and normalise
Every record starts somewhere: a form, an event, a sales note, an import. At capture, normalise formats, company names, phone numbers, addresses, and record the source and method immediately.
Provenance recorded at the point of capture is far more reliable than provenance reconstructed after the fact.
Match identities and resolve duplicates
Before anything is appended, resolve the record to the correct existing account and contact, or confirm it is genuinely new.
This means matching on company domain, name variants, address and known aliases. Not just an exact-string match, which misses parent and subsidiary relationships and near-duplicate spellings constantly.
Weak matching here is the single most common cause of enrichment appending correct data to the wrong account.
Append from governed sources
Only once identity is resolved should new fields be appended, from a defined and governed set of sources: commercial data providers, verified public sources, internal systems, or sales-supplied intelligence.
Each append should be scoped to the fields that source is actually reliable for. Not treated as a universal top-up.
Verify and record confidence
Appended data gets checked: email deliverability, title currency, site-level facts against a second source where possible. Each field receives a confidence rating and a recorded verification date.
Low-confidence fields should be visibly flagged. Never silently blended in with verified ones.
Activate, monitor and refresh
Enriched, verified data is released to the systems that use it: CRM, scoring, segmentation, sales workflows. Then it is monitored for decay.
Job titles change. Companies get acquired. Sites close.
A refresh cadence, driven by field-level freshness rather than a blanket annual re-run, is what keeps the pipeline from quietly going stale months after launch. It also depends on where that data came from in the first place, which is the next problem to govern.
Why one data provider is rarely enough
No single data provider covers every field, every geography and every account type equally well.
One source might be strong on US firmographic data and weak on European site-level detail. Another might verify emails reliably but know nothing about facility function.
Relying on one provider means accepting its blind spots as your blind spots, without knowing where they are.
The practical response is to treat provider output as evidence to govern rather than truth to copy in. Record, at the field level, which source supplied a value, when, and with what confidence.
Use multiple sources where coverage or accuracy genuinely differs. Resolve conflicts by rule, most recent, most authoritative, most specific, rather than by whichever provider happened to answer first.
Platforms such as Clay have popularised orchestrating several providers behind one workflow. That is a reasonable implementation pattern.
But the mechanism matters more than the platform. No single vendor should be allowed to define what "complete" coverage looks like for your accounts, especially once that account turns out to have six sites behind it.
Data enrichment for manufacturers and complex account structures
The account, site and buying-group model above is not specific to any one industry. It earns its keep fastest in manufacturing, industrial and other complex-B2B settings, where a single account genuinely spans multiple physical locations and applications.
A manufacturer with six sites is six facilities, potentially six different applications for your product, and a buying group that may include a corporate procurement lead, a plant engineer at each site, and a technical evaluator who has never spoken to procurement.
Enrich only at the company level, and all of that flattens into a single generic record. That is exactly what leads sales teams to pitch the wrong application to the wrong contact at the wrong site.
The same model, account and hierarchy, site and facility, contacts and buying groups, provenance and confidence, applies whether the complexity comes from manufacturing sites, financial-services business units, or a distributed services organisation.
What changes is which fields carry the most weight. Not the underlying operating model, and not the governance boundary that comes next.
Public sources, privacy notices and the EU/UK boundary
Enrichment routinely draws on publicly available information: company websites, professional profiles, registries, press coverage.
That does not put it outside data-protection law. Treating it as though it does is one of the fastest ways to create a real compliance problem.
Publicly available does not mean unrestricted
Under EU and UK data-protection law, personal data still requires an applicable legal ground to process, whether you collected it directly or found it publicly. GDPR Article 21 also gives the individual an unconditional right to object to direct-marketing processing built on that data.
Guidance from the European Commission and the UK Information Commissioner's Office is consistent on this point. Adding a publicly findable business contact to your database still engages data-protection obligations. It is not exempt simply because the information was easy to find.
Where legitimate interest is the intended ground, EU guidance is explicit: it requires a specific, identifiable interest and a genuine necessity and balancing test. Not a default justification reached for by habit.
What your privacy notice may need to explain
When you collect data about someone indirectly, through enrichment rather than directly from them, transparency obligations generally still apply. GDPR Article 14 is the relevant EU provision: where personal data was not obtained from the individual, they must generally be told where it came from, including whether it was sourced from publicly accessible sources.
EU guidance points to explaining, typically at first contact or within a defined period, where the data came from, along with the purpose of processing and the person's rights.
In practice, this usually means your privacy notice should be able to explain, in plain language, that some contact and account data may be enriched from public or third-party sources, and why.
Treat this as operational guidance rather than legal advice. Wording should be checked against your specific jurisdictions, sectors and current legal review before publication. The rules here continue to evolve.
Manual research and automation at scale
A sales rep manually looking up a prospect's job title before a call, and a pipeline that automatically appends and infers data across an entire database, are not the same activity in the eyes of most regulators. Even though both start from public information.
Automating enrichment at scale tends to increase the practical weight given to purpose limitation, data minimisation, accuracy and a person's right to object. The scale and system nature of the processing is part of what regulators weigh.
Treat manual lookup as an illustration of business purpose. Not as a substitute for the checks that scaled automation requires.
Sales follow-up and newsletter marketing are different permissions
This is one of the most consistently mishandled boundaries in B2B enrichment. Requested follow-up, ordinary sales contact under an applicable legal basis, and ongoing marketing or newsletter communication are three different permission states.
A badge scan at an event, a business-card exchange or a LinkedIn connection records an interaction. It does not, by itself, prove consent to add someone to a recurring newsletter or marketing programme. A follow-up specifically requested in the conversation should be recorded separately from any ongoing marketing permission.
UK ICO guidance on business-to-business marketing explains that the applicable UK rules depend on the communication channel and subscriber type. Its electronic-mail marketing guidance is explicit that a public contact detail is not itself consent. US law operates differently: the FTC's CAN-SPAM compliance guide requires commercial email to use accurate sender information and provide a working opt-out mechanism, rather than treating a prior interaction as general marketing permission.
Keep these permission states separate in your CRM, rather than merged into a single "contacted" flag. Automatic enrichment and inference need their own limit too, which is the next boundary worth setting.
What should not be enriched or inferred automatically
Good governance is as much about restraint as it is about capability.
Some data should not be appended or inferred automatically, regardless of whether a source or model makes it technically possible.
Avoid automatically inferring or appending sensitive personal categories, health, political or similar special-category data, where they are not directly relevant to a legitimate commercial purpose.
Avoid low-confidence inferences presented with the same visual weight as verified fact. A guessed job function should never look identical to a verified one in the CRM.
Avoid appending data beyond what your stated purpose actually requires, even when a source happens to offer more.
Avoid carrying enriched data into a use it was not originally intended for. Data appended to support account routing should not silently become the basis for an unrelated marketing decision, without a fresh check that the purpose still fits.
None of this is a case for full automation running unsupervised. Automated matching and appending can do the volume work.
Confidence thresholds, sampling checks and a human review path for high-stakes or low-confidence inferences remain part of a properly governed pipeline. What that looks like in practice is easiest to see in one worked record.
Worked example: from trade-show capture to governed follow-up
A trade-show badge scan is a useful test of the whole model. It happens fast and under pressure, exactly when governance discipline is easiest to skip.
At capture, the record holds what was actually exchanged: name, the work email the visitor deliberately supplied, company, job title, the event, booth and date, the capturing rep, and the notice or consent version shown at the booth.
That capture record is a provenanced starting point, still to be enriched.
Enrichment then resolves the visitor to the correct account, matching company and domain, checking for an existing owner or open opportunity. It appends site, division and buying-group context where a governed source supports it, verifies the work email, and records a confidence and source for each appended field.
Permission is handled separately from enrichment. The badge scan or the conversation itself supports a relevant, requested follow-up.
It does not, on its own, support adding the person to an ongoing marketing newsletter. That requires its own, distinct permission, captured and recorded on its own field.
A full operational walkthrough of trade-show capture, booth workflow, tooling and follow-up cadence lives in our companion trade show playbook. This guide covers the enrichment and governance mechanics that playbook depends on, and the same mechanics are what decide whether downstream scoring can be trusted.
How enriched data supports account research, signals and account scoring
Enrichment is the evidence layer that everything downstream depends on.
Account Research and Signals draw on the same governed, provenanced records described above: current site and application data, verified contacts, confidence-rated fields.
When that evidence is thin, stale or unverified, anything built on top of it inherits the same weakness.
Account Scoring is the clearest example. A scoring model weighs fields like company size, buying-group completeness, site relevance and observed signals to prioritise accounts and contacts.
If those inputs are guessed, unverified or out of date, the score reflects the quality of the guess. Not the quality of the account.
That is why Data Enrichment sits underneath Account Scoring in Graph's model rather than beside it. Reliable scoring is only as good as the enrichment feeding it.
Graph's dedicated account scoring guide covers how the model itself weighs and combines these inputs.
Frequently asked questions
What is data enrichment and how does it work?
Data Enrichment is the practice of capturing, matching, appending, verifying, activating and refreshing account and contact data so it can be trusted for sales action and scoring. It works as a governed pipeline where each stage builds on the last: data is captured and normalised, identities are matched and duplicates resolved, new fields are appended from governed sources, this data is verified and confidence recorded, and finally, it's activated and continuously refreshed to maintain trust and accuracy.
What is the difference between data capture, data cleansing, enrichment and verification?
Capture collects a record from its original source. Cleansing standardises and de-duplicates what you already hold. Enrichment appends new, sourced information to a matched record. Verification checks that captured or appended data is still accurate and still permitted to use. Treating these as one job is the most common cause of enrichment that looks complete but is not trustworthy.
What account and contact data should complex-B2B teams enrich?
Beyond basic company and person fields, complex-B2B accounts need account and hierarchy data, site and facility data, application context, buying-group roles for each contact, and commercial context, with provenance, confidence and freshness recorded for every appended field.
Why is one data provider rarely enough?
Coverage, accuracy and freshness vary by provider, region and field type. No single provider is strong everywhere, so governed enrichment typically draws on multiple sources, records which source supplied each field, and resolves conflicts by rule rather than by accepting whichever answer arrived first.
Does publicly available data mean I can use it without restriction?
No. Under EU and UK data-protection law, personal data sourced from public places still requires an applicable legal ground and appropriate transparency to the person it describes. Publicly findable is not the same as unrestricted, and this guidance is operational rather than a legal ruling, so check current rules for your jurisdictions before publication.
Does a badge scan or LinkedIn connection give me permission to add someone to my newsletter?
No. A badge scan, business-card exchange or LinkedIn connection can support a relevant, requested follow-up, but it is a separate permission state from ongoing marketing or newsletter communication, which requires its own basis and its own record.
How does enriched data support account scoring?
Reliable scoring depends on the same governed evidence that enrichment produces: verified, current, confidence-rated account and contact data. When that evidence is unreliable, the resulting score reflects the quality of the underlying guess rather than the quality of the account. Graph's dedicated guide to that model covers how it weighs these inputs.
