Do You Have to Clean Up Your CRM Before AI Is Useful?

Omkar Pandharkame
Co-founder of Otto.
You do not need a clean CRM. You need four structural things agreed: account hierarchy, meaningful stages, a quote or RFQ field, and clear ownership. Cleaning first is a trap. Contact data decays around 22.5% a year, so a six-month cleanup loses roughly a tenth of its accuracy before it ships. Industrial teams are luckier than most: manufacturing decays at about 10 to 15% a year, the slowest sector tracked. Your CRM is more empty than wrong. Spreadsheets import as columns, not as history. Take two years of structured data, attach the free text, and stop there. The CRM stays the system of recor
No, and waiting until it is clean is how this project dies.
That said, there is a real dependency, and it is narrower than most people assume. AI does not need a tidy CRM. It needs a small number of structural things to be true, and it needs a source of new information that is not a rep typing. Those are different problems, and only one of them is a cleanup job.
This matters because "we need to sort our data out first" is the single most common reason a field sales project sits in a drawer for eighteen months. The data never gets sorted, because the reason it is messy has not changed.

The dependency is structure, not tidiness
There is a difference between data that is wrong and data that is unmapped.
Wrong data is a stale phone number, a contact who left in 2023, a deal value that was a guess. Messy, and mostly harmless to an AI system, because new information overwrites it.
Unmapped data is different. If nobody can say which field means "branch", or your stages are named after a process the team abandoned two years ago, or the same customer exists as four accounts with no parent, then no tool can write into it correctly. That is not a cleanliness problem. It is that the shape of the record does not match the shape of the business.
The short list of what actually has to be true before any AI writes to your CRM:
- One agreed customer record per customer. Multi-site industrial accounts need a parent and its sites, not four look-alike accounts nobody merged.
- Stages that mean something to the people selling. If reps cannot say what "qualified" means at your company, an AI cannot infer it either.
- A field for the thing you actually sell on. Quote number, RFQ reference, whatever your business runs on. If it lives in a free text note, it cannot be reported on by anyone, machine or human.
- A clear owner per account. Shared territories are fine. Ambiguous ownership is not, because every follow-up needs a person.
That is the list. It is a week of decisions, not a quarter of data entry. Notice that none of it requires the existing records to be accurate.
Why "clean it first" is a trap
Data does not hold still while you clean it.
Business contact data decays at roughly 2.1% a month, about 22.5% a year, per the Dun and Bradstreet B2B data benchmark. People change jobs, companies get acquired, sites close. A six-month cleanup project hands you a database that has already lost around a tenth of its accuracy by the time you finish, and you have spent six months not fixing the thing that caused the mess.
There is one piece of genuinely good news for industrial teams. Decay is not uniform by sector. Manufacturing sits at roughly 10 to 15% a year, the slowest of any industry tracked, well under the 25 to 35% seen in technology and the 30 to 40% in startups (B2B data decay benchmarks, 2026). Plant contacts stay put. A maintenance manager at a bearing plant is likely to still be there in three years.
So an industrial CRM is usually less rotten than it feels. What it is, is empty in the places that matter, because visits never got logged. That is a capture problem, not a data quality problem, and the two get confused constantly. We laid out how that emptiness produces a wrong forecast in why field sales pipelines are always wrong.

What happens to ten years of spreadsheets
Most industrial branches have a shadow system. A spreadsheet per rep, a shared workbook of quotes, a folder of visit notes. This is usually where the real history lives, and the question is whether any of it can come across.
Honestly: some of it, and less than you hope.
What imports well. Anything in consistent columns. Account names, sites, contacts, quote numbers with dates and values. If a rep kept a tidy quote log, that is genuinely useful history and it maps cleanly.
What imports badly. Free-text notes. A column called "Notes" holding six years of shorthand across four authors will import as text nobody reads. It is not lost, but it will not become structured history, and any vendor who tells you their AI will reliably turn a decade of abbreviations into clean records is overselling.
What should not import at all. Dead accounts, contacts who left, quotes that closed years ago. Bringing them across buys you a bigger version of the same problem. Take the last 24 months and archive the rest somewhere readable.
The practical sequence: agree the four structural things above, import the columnar data for the last two years, park the free text as attachments, and let new capture fill the gap forward. Within a quarter the recent record is better than the spreadsheet ever was, because it is being written continuously instead of when someone remembers.
Where the new information comes from
This is the part that decides whether any of it holds, and it is worth being precise about the loop.
An AI sales coordinator is a tool a field rep talks to rather than types into. The rep gets a briefing by phone before a visit covering the account's open quotes, recent orders and any service issues. After the visit, the rep gives a two-minute spoken debrief from the truck. From that debrief the coordinator writes the CRM update, drafts the follow-up email, flags the quote action and notifies the manager. The rep does not open the CRM at any point.
That is the mechanism. The reason it matters to a cleanup conversation is that it changes what your CRM is for. The CRM stays the system of record, and it should. It is where reporting, forecasting and integrations live, and replacing it is almost never the right project. What changes is who feeds it. The capture layer sits in front of the CRM and protects the investment already made in it, rather than asking a rep in a truck park to protect it by hand at nine in the evening. The wider case for that split is what an anti-CRM actually is.
Put simply: you are not cleaning the CRM so AI can read it. You are fixing the structure so AI can write to it, and the writing is what makes it clean over time.

Evaluation checks any vendor can fail
Use these on whoever you are talking to, and use them on the tool you already like. Several products fail at least one, and it is better to know which.
- Does it write into the fields your managers actually report on, or does it deposit a transcript and a summary somewhere nobody opens? A summary is not a CRM update.
- Can it handle a parent account with several sites without creating duplicates every time a rep says the customer's name slightly differently?
- What does it do when it is unsure? Silently guessing a stage is worse than flagging it. Ask to see the uncertain case, not the clean one.
- Can you correct it, and does the correction stick? Ask specifically whether fixing something also fixes the records it already got wrong, or only future ones. The answers differ by product.
- Does it work with your CRM at your tier? Native integration is frequently an enterprise-only feature across this category. Confirm what your actual contract includes rather than what the website implies.
- What is the failure mode when a rep skips three days? Every tool degrades. The ones worth buying degrade visibly.
If a vendor cannot answer the third and fourth without a follow-up call, that tells you something about how mature the write-back is.
So what do you actually do first
Do the week of structural decisions. Skip the quarter of data cleaning.
Get the account hierarchy agreed, get stages named in language your reps use, add the field for the thing you sell on, and assign owners. Import two years of columnar history. Then turn on capture and let the record rebuild itself going forward.
The teams that stall are the ones who treat this as a data project. It is a capture project with a small data prerequisite, and once new information starts arriving continuously, the old mess matters far less than it did when it was the only thing you had. Teams cutting field sales admin this way, including with Otto Sales (Otto), the AI sales coordinator for industrial field sales, generally start before the CRM is tidy, because the tidying is a side effect of capture rather than a condition for it. If you are weighing what this costs across a branch where several people only need to look at the data, how view only access and seat pricing work is the next thing to work out.

FAQ
Do I need to clean my CRM data before implementing AI? No. You need four structural things: one record per customer with sites under a parent, stages your reps understand, a field for quotes or RFQs, and a clear owner per account. Those are decisions, not data entry. Existing inaccurate records get overwritten as new capture arrives.
How fast does CRM data go stale? Business contact data decays at roughly 2.1% a month, about 22.5% a year, per the Dun and Bradstreet benchmark. Manufacturing is the slowest sector at roughly 10 to 15% a year, so an industrial CRM is usually emptier than it is wrong. The gap comes from visits that were never logged.
Can AI import my old spreadsheets? Columnar data imports well: accounts, sites, contacts, quote numbers with dates and values. Free-text note columns do not become structured history, whatever a vendor claims. Import the last two years, attach the rest as documents, and let new capture fill forward.
Does an AI sales coordinator replace my CRM? No. The CRM stays the system of record for reporting, forecasting and integrations. The coordinator changes who feeds it: the rep gives a spoken debrief and the coordinator writes the update, so field sales admin stops being an evening job. It is a capture layer in front of the CRM, not a replacement for it.
What if our stages are wrong? Fix them before anything writes into them, because an AI infers stage from your definitions. If your team cannot say out loud what "qualified" means at your company, no tool can apply it consistently. This is the one prerequisite genuinely worth doing first, and it takes a meeting.
TL;DR
- You do not need a clean CRM. You need four structural things agreed: account hierarchy, meaningful stages, a quote or RFQ field, and clear ownership.
- Cleaning first is a trap. Contact data decays around 22.5% a year, so a six-month cleanup loses roughly a tenth of its accuracy before it ships.
- Industrial teams are luckier than most: manufacturing decays at about 10 to 15% a year, the slowest sector tracked. Your CRM is more empty than wrong.
- Spreadsheets import as columns, not as history. Take two years of structured data, attach the free text, and stop there.
- The CRM stays the system of record. What changes is who feeds it, which is what makes field sales admin stop being a second shift.
Set aside a week for the structural decisions, not a quarter for the cleanup. The record repairs itself once something other than a tired rep is writing to it.
By Omkar Pandharkame, Co-founder of Otto.