Home/Thoughts
Thoughts

The Same Person Showed Up 900 Times

AI enrichment pipelines don't hallucinate - they obsess. And the fix is almost embarrassingly simple.

Pipeline Diagnostic
Is Your AI Enrichment Pipeline Silently Broken?
Answer 5 quick questions to find out if you have a data obsession problem - before it torches your domains.
1. How do you source decision-makers for your outreach list?
2. After enrichment, do you check whether the same contact name appears across multiple companies?
3. When your pipeline returns a result with no errors or flags, do you assume the data is correct?
4. If your reply rate dropped suddenly, what would you investigate first?
5. How large are the lists you run through AI enrichment before any sanity check?
0
Risk Breakdown
Silent data corruption risk
Domain / deliverability burn risk
Wasted enrichment spend
Your Priority Fixes

We Thought the Pipeline Was Working

I was on a coaching call with a guy I've been working with for a while. He's building a cold outreach system from scratch - scraping leads, enriching them through a Claygent workflow, finding emails, the whole stack. Solid operator. Puts in the hours. Actually implements what we talk about, which puts him in the top 5% of people I work with.

He's been running a Claygent-based enrichment setup to find decision-makers at companies he's scraping from BuiltWith. The goal is simple: take a domain, run a LinkedIn search on Google, identify the owner or decision-maker, pull their first name, pass it downstream to an email finder, and get a verified contact. Good workflow in theory. Clean logic. Should work.

Except when he pulled up the output sheet on our call, something was off.

The same name kept showing up. Over and over and over again.

Not a handful of times. Not a dozen. We're talking hundreds of rows returning the same exact person as the supposed decision-maker - for completely different companies, completely different industries, companies this person had never worked at in their life.

His name was something like Kevin Lee. A marketing professional. Totally real person, zero connection to 99% of the businesses in the list. But he ranks high on Google for LinkedIn searches. Really high. For some reason, when Claygent ran its site:linkedin.com/in [role] [domain]-style search queries across hundreds of different companies, Kevin kept floating to the top of the results.

So the agent, doing exactly what it was told - take the top result, confirm it looks like a decision-maker, return the name - returned Kevin Lee. Hundreds of times. Confidently. With no indication anything was wrong.

This Isn't Hallucination. It's Worse.

The AI community loves to talk about hallucination. Models making things up. Inventing citations, fabricating facts, producing confident nonsense from thin air. That's a real problem, but it gets a lot of attention. What doesn't get enough attention is the failure mode we ran into on that call - and it's far more common in outbound enrichment pipelines.

Call it obsession. The model isn't making anything up. Kevin Lee is a real person. He really does have a LinkedIn profile. He really does rank high on Google. The Claygent isn't hallucinating - it's doing exactly what you told it to do, and it's doing it perfectly, except that one result has colonized every output because it happens to dominate the search index for reasons that have nothing to do with your list.

The result looks clean. The spreadsheet looks populated. The workflow says it completed. No errors. No flags. Just a lead list where every fifth, fifteenth, and fiftieth row returns the same wrong person, at scale, ready to be emailed.

If you don't look - and most people don't look closely enough - you'd never know. You'd load that list into Smartlead or Instantly and start sending. You'd be cold emailing hundreds of companies addressed to a guy who has never worked there. Your reply rate would tank, you'd blame the script, and you'd waste weeks debugging the wrong thing.

Why Claygent Gets Obsessed

To understand why this happens, you need to understand what Claygent is actually doing under the hood. It's not querying some verified B2B database. It's running Google searches - basically the same search you'd run manually if you typed site:linkedin.com/in CEO companyname.com into Google and grabbed the first result.

That's powerful. But Google's ranking algorithm is built for SEO, not for your lead list hygiene. A person with a high-authority LinkedIn profile, a lot of backlinks, a popular blog, or just a common name that happens to pattern-match well across industries will rank disproportionately high for a huge variety of search queries. Claygent sees that person at the top. It confirms they look like a decision-maker. It returns them. Every time.

The person we found in the output had nothing wrong with them. They weren't a bot. They weren't a spam trap. They were just someone Google thought was relevant - for hundreds of search queries where they weren't.

One high-authority result. Hundreds of wrong outputs. Zero model errors.

Free Download: 7-Figure Offer Builder

Drop your email and get instant access.

By entering your email you agree to receive daily emails from Alex Berman and can unsubscribe at any time.

You're in! Here's your download:

Access Now →

The Part That Made It Worse

We dug in further on the call. He mentioned he'd been getting email find rates of around 20% using first name plus company domain - which is actually a normal rate for that approach. Not great, but not broken.

Then he told me he'd figured out that if he passed the actual LinkedIn URL instead of just the first name into the email finder, the match rate jumped to 40-50%. That's a real, meaningful improvement - potentially double the emails found from the same list. He was right to get excited about it.

But then we looked at where those LinkedIn URLs were coming from. Same problem. The Claygent was pulling URLs based on those same Google searches. Which meant when Kevin Lee was showing up as the "decision-maker" for a roofing company in Ohio, the LinkedIn URL being passed downstream was Kevin's profile - and the email finder was potentially surfacing Kevin's email and tagging it to the wrong company's row.

The enrichment looked richer. The email find rate looked better. But the data was still wrong. It was just wrong with more information attached to it now.

This is what makes the obsession failure mode so insidious. It doesn't break your pipeline. It corrupts your pipeline while making it look like it's working better than ever.

The Fix Is Embarrassingly Simple

We talked through a few technical angles - tweaking the search queries, trying different Claygent prompts, exploring whether there was a way to filter by actual employment data. Some of that is worth pursuing.

But the real fix - the one that would have caught this immediately - is a deduplication check. That's it.

Before you pass any enriched name downstream to the email finder, before you do anything else with it, add one rule: if the same name appears more than twice across your entire enrichment output, flag it.

If one name is showing up three times, five times, fifty times, nine hundred times across a list of thousands of companies - that is structurally impossible if the enrichment is working correctly. No individual person is the decision-maker at hundreds of different companies. That pattern doesn't exist in real life. When you see it in your data, you're not looking at a great result. You're looking at a broken input signal that the model is confidently repeating at scale.

You don't need a smarter AI model to catch this. You don't need better prompts. You don't need a more sophisticated Claygent. You need a COUNTIF in a Google Sheet. Literally. Count how many times each name appears in your output column. Flag anything over two. Review those rows. Done.

The AI isn't going to tell you it's stuck on someone. It has no concept of "I've returned this person too many times." It's stateless across rows - it processes each row independently and returns its best answer each time. The obsession only becomes visible at the list level, and the AI never sees the list level. You do. So you have to check it.

The Broader Lesson About AI Enrichment

I've been preaching volume for years. If you've read The Cold Email Manifesto, you know my position: the number of qualified contacts you can reach at scale beats the depth of enrichment on a small list almost every time. I've said this repeatedly in my emails, on calls, everywhere.

But volume only works if the data underneath it is structurally sound. Sending 100,000 emails addressed to Kevin Lee - a guy who has never worked at any of those companies - isn't volume. It's noise. Worse than noise, because it burns your domains, tanks your deliverability, and makes you think your script is broken when the script is actually fine.

This is why I keep saying that when you build an AI enrichment pipeline, you need at least one human-logic sanity layer baked in. Not AI logic - human logic. Stuff that's obvious to any person who thinks about it for five seconds: the same person shouldn't be the decision-maker at 900 different companies. If your output shows that, something is wrong upstream.

The smarter your pipeline looks, the more dangerous these silent failures become. A manual list built by a VA has obvious errors you catch immediately. An AI-enriched list that looks clean, exports beautifully, and fills every column with confident data - that's the one that can quietly destroy a campaign at scale before you even know what happened.

Need Targeted Leads?

Search unlimited B2B contacts by title, industry, location, and company size. Export to CSV instantly. $149/month, free to try.

Try the Lead Database →

What We're Doing Instead

After running through this on the call, we landed on a two-part approach. First, the deduplication check as described - count name occurrences before anything moves downstream. Flag and quarantine any name that appears more than twice.

Second, for the companies that don't return any usable LinkedIn result at all - which was a real chunk of this list, since a lot of the smaller companies being scraped from BuiltWith just don't have much of a Google footprint - we talked about adding an AI qualifier at the website-scraping step. When Claygent is already visiting the domain to confirm they have live chat or whatever the target technology is, it can also make a simple call: is this a legitimate business with a real web presence? Is there enough signal here to enrich further? If not, drop it before the enrichment credits run.

That's the logic: don't spend money enriching companies that aren't going to yield good data anyway. Qualify first, enrich only the ones that pass, deduplicate the output names before anything moves downstream.

It's not a complicated system. The whole thing can run in Clay or in an n8n workflow - and the deduplication check is literally a spreadsheet formula either way. What matters is that you build it in. It doesn't happen automatically.

One More Thing on Lead Sources

We also talked about the source list itself. He was pulling domains from BuiltWith filtered by live chat technology - a solid signal for the product being sold. But when he checked Apollo with the same technology filter applied, he found around 400,000 companies with live chat. The BuiltWith list was giving him a different, often non-overlapping set - which is exactly what you want.

When we cross-referenced the two, around 600,000 leads from Apollo's full database weren't showing up in the BuiltWith data at all. That's not a bug - those are leads that aren't in the standard databases, which means less competition reaching them. That's the whole point of scraping from BuiltWith and similar sources in the first place.

If you're building a similar setup and want a clean B2B lead database to cross-reference against or start from, ScraperCity's B2B email database is worth looking at - and if you're pulling from Apollo specifically, the Apollo scraper handles that without burning your Clay credits on something that doesn't need AI. For finding emails once you've got your names sorted, Findymail is what we use - it's fast, verified, and gets you to 40-50% match rate when you're passing LinkedIn URLs instead of just first names.

The overall lesson isn't that AI enrichment is broken. It's that AI enrichment is confidently stupid in specific, predictable ways - and those ways are easy to catch if you're looking for them. Build the sanity check in. Run the deduplication. Look at your output before you send.

Kevin Lee doesn't work at any of those companies. Check your data and make sure he's not in your list 900 times either.

If you want to go deeper on building lead lists that actually hold up, grab the Best Lead Strategy Guide - it covers the sourcing and verification layers in detail. And if you want to work through your own pipeline on a live call, that's what Galadon Gold is for.

Ready to Book More Meetings?

Get the exact scripts, templates, and frameworks Alex uses across all his companies.

By entering your email you agree to receive daily emails from Alex Berman and can unsubscribe at any time.

You're in! Here's your download:

Access Now →