A conversation that, in some variation, repeats itself in every manufacturing company:
“How long do you need to deliver the material?"
"Seven days."
"And when did it last arrive in exactly seven?"
"…”
Sound familiar?
Someone calculated that number, or eyeballed it, somewhere around 2019. Ever since, it has gone into every plan, every quote, and every promise to a customer. And nobody has ever compared it against reality. Meanwhile, circumstances change.
And now we get to what this text is really about.
The problem isn’t that the number is wrong. The problem is that it’s one number.
You can fix it, refresh it, calculate it more precisely. It will still wreck your plan, because the reality it describes is not a single quantity.
How common is this
It’s not an exception. A survey of manufacturing companies in the region (Qlector, 2025) shows that half still plan with Excel or paper despite having an ERP, and the hidden cost of inefficient planning is estimated at 5 to 10 percent of annual revenue. Only one in four companies rates its own planning as truly effective.
In broader supply chain surveys, the top of the problem list is held by lack of visibility (70 percent) and inconsistent or missing data (65 percent).
ERP consultants describe it in almost the same words: lead times entered at implementation grow stale as supplier performance changes, so the recommendation is to regularly compare actual receipt dates against planned ones.
One caveat, because I think it’s a fair one. Most of this research is funded by software vendors, the samples skew toward larger companies and Western markets, and figures like “5 to 10 percent of revenue” are estimates, not measurements. Use them as a direction, not as fact. The only numbers that truly matter for your company are your own — and, as you’ll see at the end of this text, getting them takes one afternoon.
How that one number costs you
The plan fails half the time. If you plan based on the average, half of the deliveries are, by definition, slower than it. That’s not bad luck, that’s arithmetic.
The error only goes one way. The material arrives three days early, but you don’t install early: you’re waiting for the crew and the glass. It arrives three days late and everything downstream shifts.
Why is that? Because no step in production starts when the average input arrives — it starts when all the inputs arrive. Assembly waits for the profile, and the glass, and the hardware, and a free crew. So you’re not waiting on the average of those branches, you’re waiting on the slowest one. And the slowest branch is, by definition, worse than the average.
That’s why the sum of average durations is almost never the actual duration of the job. And that’s why companies where everyone works to plan still run chronically late, and nobody can point a finger at exactly where it happened.
The buffer is in the wrong place. A supplier with a range of six to eight days and one with a range of four to twenty have an identical average. The first one doesn’t need a buffer at all. The second one sinks your quarter, usually in exactly the month it must not. If you plan by the average, you treat both the same, so you either smear the buffer everywhere a little (and pay for it) or have it nowhere (and pay for that).
Knowledge leaks out of the company. The planner sees the system is wrong, so they plan “from their head” and from their own spreadsheet. At first that looks like resourcefulness. But when they leave, the spreadsheet “leaves” too.
Customer trust leaks. Miss the deadline half the time and it gets booked nowhere, but it’s felt at the next price negotiation.
The first measurement: where your lead time actually lives
Before we get into what AI changes, it’s useful to see what the picture the average is hiding looks like. You don’t need any tool for this.
Take one supplier and pull thirty receipt dates from last year. For each, calculate the number of days from order to receipt. Then sort them from shortest to longest.
Now you’re looking at a distribution, not an average. And you can immediately read off what the trade calls a percentile:
- The fifteenth in the list is your median. That’s how long a typical delivery takes.
- The twenty-fourth is roughly the eightieth percentile (P80). In eight out of ten cases, it arrives by that deadline.
- The twenty-ninth is close to the ninety-fifth percentile (P95). That’s your bad scenario, the one that happens once or twice a year and wrecks your whole month.
What the trade calls the “tail” is nothing mystical. It’s the gap between the twenty-fourth and thirtieth rows in your table.
If that gap is small, you have a reliable supplier and the average hasn’t done you much harm. If the gap is large, you’ve just seen how much risk you were carrying without measuring it.
How to become AI-fficient
When you see how much the average is costing you, the first thought is: we need a more precise average. You don’t. You can do that yourself, in Excel, in five minutes.
AI does two different things here. They’re worth separating, because they constantly get blended into one, which makes them look like magic.
First: an estimate that knows the context
Lead time isn’t one number at all. It differs by supplier, item, color, order size, by whether it’s December or May.
Nobody holds eight or ten such cases in their head, so in practice one number gets used for everything and that’s the end of the story.
A model can capture and account for all of it at once: “for this supplier, this profile, an atypical color, in December: the median is nine days, but the tail goes to twenty.”
The point isn’t precision. The point is that an estimate made in a richer context is richer, and therefore better.
But there’s a trap you should know about in advance. The finer you slice, the less data remains in each group. If you split by supplier, then by item, then by color, then by season, you easily end up with a group of three deliveries. A tail estimate on three records isn’t knowledge, it’s noise — and dangerous noise, because it looks precise.
A practical rule: for a tail estimate to mean anything, you need at least twenty or so deliveries in the group. When there aren’t enough, you don’t invent a number — you merge the group into a broader one: instead of “this profile in this color,” you look at “this group of items from this supplier.” Coarser, but honest.
A good system does that on its own and, more importantly, tells you how many records the estimate rests on. If your software gives you a number without that information, ask for it.
Second: reading the history you already have
The first capability is impossible without this one. Segmented estimation demands a pile of historical records, and those records, as a rule, don’t exist in a form you can open.
They do exist, however, on goods receipt notes. Even when those receipts are scans in a binder nobody has opened since it was filed.
Why wasn’t that usable until yesterday? Not because document-reading software didn’t exist. OCR (text recognition) has existed for decades. The problem is that OCR sees letters, not meaning.
To explain to it which of the six dates on a receipt note is actually the receipt date, you built a template. A separate one for each supplier, because everyone has their own field layout. Then you fixed it whenever someone changed their form. Add to that a stamp over the text, a hand-written quantity, a table that breaks across two pages, and a scan taken with a phone. A job that never paid off.
Language models (LLMs) don’t work with templates. They read a document roughly the way a new hire in procurement would: this is a receipt note, here’s the supplier, this date is when the goods arrived, regardless of where it sits on the paper. The same procedure works on one supplier’s form and another’s, even though they look nothing alike.
And, most important in practice: the model can say when it isn’t sure. Those documents go to a human, instead of quietly entering the data as an error.
How do you know it was read correctly? This is the question to ask anyone offering you this, including me.
The answer is measurement, not trust. You take fifty random documents, manually check the extracted fields, and get an accuracy percentage. On top of that there’s a cross-check: the sum of amounts from the processed invoices has to reconcile with the books. If the totals match, the extraction stands.
Accuracy below ninety-something percent on the key fields means something isn’t configured right. It’s measurable, and it should be treated that way.
So what actually changes in the company
Not the software. What changes is how you choose the date you’re going to promise.
And let me clear this up right away, because it’s the most common misunderstanding: the probabilities never have to leave the company. The customer still gets one date, same as before. The difference is that from now on you choose it, instead of guessing it.
Internally it looks like this: eighty percent chance we finish by March 14, ninety-five percent by March 19.
Then you decide. For a regular customer who can tolerate a slip and with whom you have a relationship, you probably go with March 14. For a job with contractual penalties, for a first job with a new customer, or for a project the whole quarter depends on, you promise March 19 and keep your word.
Before, you had one number and no choice. Now you have a choice, and you know exactly what you’re risking when you opt for the earlier date.
Along the way, you also get the price of your suppliers’ unreliability
Look again at that gap between March 14 and March 19. That’s five days you’re not spending on production, or transport, or installation. You’re spending them on uncertainty.
That is this supplier’s unreliability, expressed in days. And it can be converted into money: multiply those days by what a day of buffer costs you, whether through inventory sitting idle or through a job invoiced later.
Once you see that amount written down, the conversation with the supplier looks different. It no longer goes “you’re late,” which is an argument without end, but “your variability costs us this much per year — let’s see what we can do about it.” That’s a conversation with an outcome.
Sometimes the answer isn’t a better estimate but less variability
This is the part you rarely hear from someone selling a tool, so I’ll say it explicitly.
Everything above solves how to live with variability. But a buffer is a cost. If you optimize it, you’re still paying it, just more intelligently.
Sometimes it’s cheaper to attack the cause itself:
- A conversation with the supplier, but with numbers on the table instead of an impression.
- A second source for the few positions that vary the most, even at a somewhat higher price. A higher price is a known cost; a delay is not.
- A change in ordering rhythm, but only for the unreliable items, not for everything.
- Standardization where it’s possible. Atypical colors and atypical dimensions usually carry the longest tail.
The measurement this whole text is about is useful here too, because it shows you which few items carry most of the risk. There usually aren’t many. Once you see them, the answer is often a negotiation, not an algorithm.
Five questions you probably have, and answers you might not expect
“So we have to digitize everything?”
No. And please don’t, because that’s the surest way to turn this into a year-long project that never finishes.
It works the other way around: pose one question, then extract only the fields that answer it. Here you need four: supplier, item, order date, receipt date. Two years back. Enough.
When you answer that question and see the value, you pose the next one. Each subsequent one is cheaper, because the infrastructure is already there. But each one has to earn its keep on its own.
“But we don’t have that data.”
You do — on the receipt notes. Until yesterday you couldn’t read them because OCR saw letters, not meaning. Language models read a document like a person who works in procurement: wherever the date sits on the paper, they recognize it.
The data you need exists even when it doesn’t exist in the ERP. That’s a difference worth keeping in mind for all future questions, not just this one.
“That must be expensive.”
It used to be. Retyping ten thousand documents was one person’s job for several months, and that’s why it never got done.
Today it’s a job of a few days. And, more important for the decision: you see the result before you change anything about how you work. First you get the numbers, then you decide whether it’s worth going further. There’s no big upfront stake.
“Do we need a new ERP then?”
Probably not. A new system inherits the same inaccurate data that was wrecking the old one, because the planning engine doesn’t correct what you feed it. That’s why companies that go into a system replacement with untended parameters often end up, a year later, with a more expensive version of the same problem.
The order is reversed: first find out which of your numbers are actually accurate, and only then decide which system to put them in.
“Do we have to talk to the customer about probabilities and context?”
No. The customer still gets one date. The difference is that you now choose it better: eighty percent chance by March 14, ninety-five percent by March 19. For a regular customer you go with the 14th, for a job with penalties, the 19th.
Probability is your internal instrument. What goes out the door is a promise, same as before — except now you know what it’s actually worth.
What you can do today
Take one supplier, pull thirty receipt dates from last year, sort them by duration, and look at the three longest. Compare them with the number you use when you plan.
Then count in how many cases last year your plan would have failed by that number.
For this you need no software, no project, no budget. You need one afternoon and the willingness to see a number that probably won’t be pleasant.
The real question isn’t “how long does the supplier take to deliver on average?” but “how do (all) the characteristics of the supplier’s delivery time affect our plan?”