Imagine an ordinary Monday. There is an email from a client asking to move a meeting and add one more person. Your CRM - the program where the team manages clients and sales - still shows the old time. In a spreadsheet, the manager has already written “agreed.” You show all three sources to an AI assistant, and it confidently suggests a fourth option: send a confirmation for an available slot.
It looks as if the AI made a mistake. In this hypothetical situation, one possible cause is conflicting records without a rule establishing which one takes precedence. That does not rule out a model error: this article examines one failure mode, not every reason AI automation fails. If you give it permission to act, it may do more than write an unfortunate text. It may change the state of a deal, create a task for the wrong person or leave the team with two different agreements.
That is why it is worth running a process audit before development - a brief review of how work actually moves from one person and system to another. This is not a check of whether your CRM is “modern enough.” It is a way to avoid handing a new tool a process whose rules still live in the team’s heads.
1. The model sees data, but not always its weight
AI can help read emails, find the request in them, explain the content briefly and suggest a next action, but its accuracy needs checking on your own examples. There is an important boundary between “read” and “change a working record.”
Schnittstelle is the German word for interface: here, the point through which two systems exchange data. Context can be lost at this boundary. One service sends a client’s name without a deal number. Another does not say that a record is already closed. A third accepts any update, although only the responsible manager should make it. AI cannot reliably infer an organisation’s unwritten rule.
The problem is often not the data exchange itself, but its meaning. For example, the “stage” field in a CRM may mean that a manager has prepared a proposal. The same word in a spreadsheet may mean that the client has already approved it. Technically, both records can be passed between systems. But automation will not distinguish a working note from a decision unless you define that distinction.
An older system may still be fit for purpose
This is not a problem limited to old software. A survey commissioned by Zapier reports integration difficulties, cost, vendor dependence and missing AI skills as barriers to adoption. Centiment surveyed 532 US executives, presidents, owners or partners at companies with 1,000 or more employees on September 19–23, 2025. These responses do not establish how common the barriers are among small B2B businesses. Our editorial recommendation for a smaller business is simpler: you do not need to declare your existing system unusable. Check one specific handoff - what comes in, what changes, who confirms it and what comes out.
For example, in a company providing Gebäudereinigung - professional building cleaning - an enquiry may arrive through a website form, email or phone call. If all three channels ultimately end up in the CRM, a first pilot could limit AI to reading new enquiries, extracting the address and property type, and preparing a draft card. But if a manager still checks the service area in a personal spreadsheet, the agent should not independently make a promise to the client. It can show the mismatch and ask for a decision.
Another example: if a client agrees to a price in an email but the CRM still contains a different amount, automatically sending a contract is a bad first step. A useful first run can instead collect such discrepancies into one list for the manager. You do not lose control, but you can see exactly where the process diverges: in correspondence, data entry or the rule by which the team considers an amount final.
2. Six questions to go through before any development
A process audit does not require a technical lecture. Take one recurring process: handling a new enquiry, approving a commercial proposal, sorting email or handing a project over for delivery. The six questions below are our editorial method for structuring that review, not a framework prescribed by the cited publishers.
It is useful to examine not an imagined ideal process, but one recent case. Open the email, the CRM card and the document where the team clarified something. You will then quickly see not only the official route, but also the detours: a message in a messenger, manually copying a number, the agreement that “I will update it later.” Do not try to describe the whole company at once or prepare a long presentation. Choose the point where the team regularly copies, asks again or searches for the latest version.
Then go through six questions in the order of the actual work.
- Where does the request start? Write down not “from a lead,” but specifically: from a form, email, call, messenger or a manually created card. This helps you avoid forgetting the channel that exists only because “it is more convenient this way.”
- Which source is the Quelle der Wahrheit - the record people trust when data conflicts? For price, this may be the approved proposal; for the deal status, the CRM; for the visit date, the calendar. If there is no answer, do not let AI change that state.
- Which state transitions exist? Name them simply: “new enquiry,” “clarification needed,” “proposal sent,” “approved,” “closed.” Next to each transition, write what must happen to move forward. Not “when everything is ready,” but “the client approved the amount in writing.”
- Who owns each transition? The owner is not the person who sometimes helps, but the person who has the right to say: “yes, the status has now changed.” If there is no owner, stop that transition for clarification rather than leave the automation to guess who is responsible.
- Where may AI read, and where may it write? Leserecht means permission only to view data; Schreibrecht means permission to change it. Start with only the read access the task needs, and narrowly limited write access or none. Check the actual connector permissions, not just the instruction given to the model. This lets you test the rule against real exceptions while limiting exposure and changes to important records.
- Which cases do not fit the normal path, and what completes the work? Write down at least the exceptions the team mentions: a duplicate enquiry, a client who changed terms, missing data, an existing contract, denied access. Separately name the proof of completion: the email was sent and saved, a task has an assignee, the record was updated in the source of truth, a person confirmed the decision.
Gmail illustrates why permission names matter. Google recommends the narrowest scopes the app needs. gmail.readonly permits viewing messages and settings; gmail.compose permits both managing drafts and sending email. A workflow described as “draft only” therefore needs an enforced application boundary if it uses gmail.compose: the model’s instruction to wait for approval is not a restriction on that credential.
You do not need to draw a complex diagram. One sheet or shared document is enough, as long as every step shows the input, source of truth, responsible person, permitted action, exception and proof of completion. Such a map is more valuable than a beautiful demo because it shows what actually needs to be built and what first needs agreement without any development. If two people answer the question “who can change the status?” differently, that is already a result. Do not argue with AI about the accuracy of its answer. First agree on the rule manually, add it to the working description and only then hand it to automation.
3. The first run should be a safe check, not a replacement for the team
A small controlled pilot tests an idea in a narrow area without handing it the whole process. For AI, it is especially useful where the result can be read, checked and reversed. This is not automatically a canary release: Google SRE defines canarying as a partial, time-limited deployment of a service change and its evaluation, comparing the changed portion with a control. A manually reviewed draft pilot can borrow the discipline of limited exposure, evaluation and stopping without being a production canary.
Our recommended first pilot has three characteristics. It takes one type of input, does not change a critical record without a person and leaves a trace that lets you understand why the result appeared. For example, AI could read emails with the subject “enquiry,” extract facts only from the email itself and prepare a response draft together with a link to the source. A manager would approve or reject it manually. This is a hypothetical workflow: not only the AI suggestions matter, but also their link to evidence and human control.
Before such a run, agree on what you will consider an error. Not the abstract “it wrote badly,” but understandable cases: it missed a required detail in the email, mixed up the client, suggested an action without a source or sent something to a queue that does not belong to that queue’s topic. Then the person checking the results does not simply correct everything silently, but returns the process to a specific rule.
A hypothetical pilot, with a decision at the end
Suppose the cleaning company tests 20 enquiries over one working week. These numbers illustrate a plan, not a recommended sample size or a client result. For enquiry E-104, the email requests an address outside the approved service-area list, while the CRM says “ready.” The agreed rule gives the service-area list precedence for coverage. AI prepares a review item containing the enquiry ID, the conflicting records and their links; it leaves the CRM status unchanged. The manager decides whether an exception is possible. Completion means the review item records that decision and its owner, not merely that AI generated text.
Before starting, the team agrees to pause on any wrong-client match or unapproved write, and to record missing facts, unsupported statements and review time for every item. Suppose the pilot ends with 17 drafts needing no factual correction, two missing an address and one matched to the wrong client. The wrong-client match triggers the agreed stop: correct the matching rule and repeat the limited test. The 17 usable drafts do not cancel that error. Compare checking and correction time with the manual task before deciding whether the pilot is useful; this small run cannot establish reliability for every channel or rare exception.
The boundary of a safe launch
A bad first pilot looks different: “let the agent manage all leads on its own.” It contains too many different rules, channels and consequences. When something goes wrong, it becomes harder to distinguish an email-recognition error from an integration failure, a routing rule or a write-permission problem. Such a launch is like repairing the wiring in an entire building with one switch: the lights may come on, but finding the cause afterwards will not be fun.
If you still need an action in the system, choose a reversible one. For example, AI does not change a deal budget but adds a “review needed” tag; it does not send an invoice but creates a draft; it does not close a task but adds it to a confirmation queue. Check whether that tag, draft or queue entry triggers another automation: removing a tag will not necessarily undo what followed it. Before launch, define who reviews that queue and how they reverse a mistaken action. Without this, “automatic” sometimes only means “the error had time to run ahead of you.”
Separating roles matters too. For this first pilot, our recommendation is to keep decisions and accountability with a person, assign preparation to an assistant and allow automation only a limited action under a defined rule. This is an editorial allocation of responsibility, not a technical definition of an agent. In Building effective agents, Anthropic distinguishes workflows following predefined code paths from agents dynamically directing their processes and tool use. Its engineering guidance supports starting simply, checking evidence from the environment, using human checkpoints and defining stopping conditions. A fixed sequence may be sufficient for your handoff; it does not need to become an autonomous agent.
4. What should remain after reviewing the process
The result should not be a report that is easy to put in a folder and forget. It should become a short working description that you and the team can use to make decisions.
Keep a map of one process with seven fields: request source, source of truth, states, owner, AI permissions, exceptions and proof of completion. Add the boundary of the first launch: exactly what AI reads, what it suggests, what it does not change and who checks the result. Mark separately where a human decision is needed and where a pre-agreed rule is sufficient. Record unresolved questions with an owner; do not silently turn an unanswered question into permission to act.
This description is useful not only to a developer. A new manager can use it to understand where to look for the current price. The process owner can see which decisions are hanging between roles. The person responsible for access can give the tool exactly the permissions needed for the first task. And you can compare providers’ proposals not by the promise that “we will integrate everything,” but by whether they see the real boundaries of the process.
After this, you may find that automation should not be built yet. For example, if commercial terms are approved in chat without a single record, and managers use different meanings for the status “ready,” first agree on a manual rule. This is neither defeat nor a rejection of AI. It is a way to remove a broken handoff before adding another service to it. A broken process is inventive enough without AI: it can lose the responsible person even across three tabs.
A controlled workflow can already be proof that the team knows how to build a process with checks. But it does not prove that every next task can safely be handed to an agent. That is why controlled AI systems should be assessed not by how impressive the first screen looks, but by whether it is clear where the agent stops, who sees its action and what confirms the result.
5. Start with one place where the team looks for truth
Do not start with the question “which AI should we buy?” Open one real process where people often compare email, the CRM and a spreadsheet. Take the latest example where someone asked “which status is correct?” and write beside it where there should have been one truth, who could change it and what was missing for completion.
You do not need to decide immediately which system to replace or how to connect everything. Your first task is to name the boundary of the problem in a way the team will recognise. When one case has a source of truth, an owner and proof of completion, you have a candidate for a small launch whose permissions and stop conditions still need testing.
That is the first step. Once it is done, AI has clearer rules for handling three versions of reality and can become a useful assistant in clearly defined work. Its results still need verification.



