Every process in your organisation exists in two versions. There is the one in the procedure document, with its swim lanes and its approval steps. And there is the one that actually happened last Tuesday, to order 20471, when the credit check ran twice and the shipment went out on Friday instead of Wednesday.
Process mining reconstructs the second version. It does not interview anyone. It reads the traces that the systems left behind as the case passed through them, and rebuilds every path a case took: the designed one, and the fourteen others nobody documented.
That makes the log the whole game. A process-mining platform is only as good as the event log it is given, and most of the pilots that stall do so for reasons that were visible in the data before the platform was ever switched on. This article is about those reasons: what the log needs to contain, where it usually hides, how it goes wrong, and how to check your own before you commit to anything. First, though, why go to the trouble at all.
Why mine a process, and what you get
Most organisations already know their processes are slower and messier than the documentation says. What they do not have is an objective, shared picture of where and by how much, and so the conversation about fixing them runs on anecdotes: operations blames the credit team, the credit team blames the data, and the improvement budget goes to whoever argued best.
Process mining replaces the anecdotes with a measured record. Because it is built from every case rather than a sample, and from system timestamps rather than recollection, it produces things that interviews and dashboards cannot:
- The real process map, with every path cases actually take, and how many took each one.
- Where cases wait, loop and get reworked, quantified by product, team, channel and case type, so the biggest delay is a number rather than an opinion.
- Deviations from the designed process, including the hand-offs between systems and people where compliance quietly breaks, detected from the log rather than reported after the fact.
- One improvement roadmap the whole organisation can agree on, because everyone is looking at the same evidence.
Once the discovered process is kept live, the same event stream supports the next layer: models on the case attributes that explain why cases stall and score the open ones for the probability of breaching an SLA or a control, so teams intervene while the case is still open rather than in the monthly review.
All of that rests on the log. So: what has to be in it.
The three fields
Strip away the vocabulary and an event log is a table with three columns. For every step a case passed through, it records which case, what happened and when.
Case ID is the field that ties one case together. In an order-to-cash process it is the order number; in onboarding it is the application reference; in claims, the claim number; in service management, the ticket. The requirement is not that it be pretty, only that it be the same identifier on every event belonging to that case. A log where the same order appears as ORD-20471 in the sales system and SO/2026/20471 in the warehouse system has a problem, but a solvable one, which we come to below.
Activity is what happened at that step. The useful form is a status change: “credit check completed”, “document requested”, “approved”, “shipped”. The unhelpful form is a free-text note, a screen name, or a field called status that was overwritten each time so that only the latest value survives. Activities also need to be consistent: if one team records “Approved” and another records “APPR” and a third records “Stage 4 complete” for the same step, the model will show three activities where there is one. That is fixable in harmonisation, but only if somebody knows they are the same thing.
Timestamp is when the step happened. Not when the record was last modified, not when the nightly batch wrote it, not the date somebody typed into a form from memory. The distinction matters more than it sounds, and it gets its own section.
With those three fields, and nothing else, a platform can discover the process, count how many cases took each path, measure how long each step took, and show where cases waited, looped or deviated. Everything richer than that, such as who handled the step, which product it concerned, which channel it came through, is an optional attribute that makes the analysis sharper. But three fields is the entry ticket.
Where the log hides
The most common objection we hear is “our system doesn’t keep that”. It almost always does. Screens show the current state of a record; databases keep its history, because auditors, regulators and support teams have needed it for decades.
Where to look depends on the system:
- Workflow and case-management tools keep a task or stage history by design. It is usually the easiest export.
- ERP systems record change documents against orders, deliveries and invoices: every status change, with a user and a time. They are rarely exposed on a screen, and always in the database.
- CRM platforms keep a field-history or audit table for opportunity and case stages, often switched on by default for the fields that matter.
- Ticketing and service tools keep a full event stream; it is what their own SLA reports are built from.
- Core banking, policy and claims systems vary most, but the transaction and status logs that satisfy the regulator are usually the ones you need.
The practical step is to ask IT the right question. “Do we have an event log?” gets a no. “Does the system keep an audit or history table for status changes on this object, and can we export the last six months of it?” gets a table.
When the case ID changes between systems
A process that crosses systems usually crosses identifiers too. The sales order becomes a delivery number becomes an invoice number; the application reference becomes a customer ID becomes an account number. Each system is internally consistent and none of them agrees with the others.
There are three ways through, in order of preference:
- A carried key. Very often the downstream system stores the upstream identifier in a reference field, because someone once needed to trace it. If the delivery record holds the order number, the join is direct.
- A mapping table. Where no reference field exists, the integration layer that moved the record between systems usually kept a cross-reference, or can be made to. One table with two columns solves it.
- A secondary key. Failing both, cases can be stitched on a combination of attributes that is unique enough in practice: customer plus date plus amount, or product plus location plus week. It is less exact, and the analysis should say so, but it is far better than treating each system as a separate process.
What does not work is hoping the platform will figure it out. Case identity is a decision, and it has to be made before the first model is built.
Timestamps: the four ways they lie
Timestamps are where the confident pilot quietly goes wrong, because the data looks fine and the model looks plausible right up to the point where someone who knows the process says “that can’t be right”.
Batch updates. A system that writes status changes once a night will show every step of the day happening at 02:00. Durations inside the day vanish; everything looks like it took exactly one night. The fix is to find the real event time in the source, or to treat that system’s steps at day resolution and say so.
Time zones. A process that runs across New Jersey and Visakhapatnam, or across a regional operation and a shared-service centre, will produce timestamps in different zones. Left unharmonised, a step can appear to finish before it started. Every log needs to be normalised to one zone before anything is measured.
Last-modified overwriting history. Some tables keep only the latest change: a modified_at field that moves every time the record is touched. That is not an event log, it is a snapshot. If the history is not kept elsewhere, the process can only be reconstructed from the point at which logging is switched on, which is a reason to switch it on early.
Human-entered dates. A “completion date” typed into a form by the person completing the task is a recollection, not a record. Typed dates cluster on Fridays and month-ends, and they are often entered days after the fact. Where the system also stores when the form was saved, the saved time is the better timestamp.
None of these is fatal. All of them are common. The difference between a pilot that produces findings and one that produces arguments is whether they were found in the first two days or in the review meeting.
How much history is enough
Three to six months of events is the usual answer, and more is not automatically better.
Shorter than three months and rare paths, the exceptions and rework loops that are usually the point, do not appear often enough to be measured. Longer than about a year and the process itself has often changed: a system was replaced, a team was restructured, a control was added. Mixing the old process with the new one produces a model of neither.
The honest framing is that the window should cover at least one full cycle of whatever drives the process. Month-end for finance processes, a full season for retail replenishment, a renewal cycle for insurance. And it should stop at the last significant change to how the process is run.
If nobody in-house can build the log
A common reaction to all of the above is that it sounds like a data-engineering project before the real work can start, and that nobody in the organisation has the time to do it. That is a fair reading, and it is why building the log is part of what we do rather than a precondition for it.
In a two-week pilot, the extraction and harmonisation is the first half of the work: connecting to the systems the process runs through with default connectors for the common workflow, ERP, CRM, ticketing and core platforms; agreeing the case definition with the process owner in one hour; joining identifiers across systems; and normalising timestamps into one timeline. The first discovered model is on the table by day five. What we need from your side is an export or read access, and the hour with the process owner.
The harmonised log is also not a throwaway. It stays with you, and because the hard part was deciding what a case is and how the systems fit together, the second process onboards far faster than the first. That is the point at which process intelligence stops being a project and becomes a capability, which is what the Process Mining & Intelligence service is built around.
A five-minute check before you commit
You can find out how ready your process is before anyone extracts anything. Five questions decide most of it: how many cases a month it handles, how many systems a case passes through, whether those systems record status changes with a time, whether one identifier follows the case across them, and whether someone owns the outcome and can act on what is found.
We put those five questions into a readiness check. It scores the answers, tells you which of the points above applies to your process, and, if the answer is “not yet”, says what to fix first. If the answer is “yes”, the next step is a two-week pilot: one process, your real event data, a working model on day five and a quantified list of where it stalls on day ten.
Either way, the data your systems already write is the place to start. It has been recording the real process the whole time; process mining is just the first thing to read it.