This article is part of our Journal archive. Any prior offers reflect its publication date. Read our current services and approach.
We assess paperwork by repetition, inputs, judgment, exceptions, and review effort, then define a small test to decide whether AI, ordinary software, or process cleanup fits.
The best place to start is a recurring task that holds up useful work and has a result someone can check. Copying details from a request into a draft record might qualify. Deciding whether to approve the request is a different job, even when both happen on the same sheet of paper.
We look for that distinction before choosing technology. A narrow change should have a clear business purpose: less repeated entry, a clearer handoff, or fewer incomplete records reaching the next person. Those are goals to test, not outcomes to assume.
As an AI-native strategy and engineering studio, we start with the valuable problem, examine the data and constraints, and define a small test. Short iterations let you inspect the work while the decision is still easy to change. AI belongs in that evaluation when the task calls for it.
Look at repetition and volume together
Choose a specific step, such as preparing a draft intake record from an incoming form. Name its trigger, inputs, output, and next owner. A department or an entire approval process is too broad for a first candidate.
Then look at how often that step happens and how much attention it takes. Use a representative period, including busy and quiet stretches. Count repeated entries, corrections, and handoffs rather than relying only on someone's memory of a difficult afternoon.
High volume alone does not make a task suitable. A frequent task that takes little effort may leave little room for improvement once review and maintenance are included. A less frequent task may deserve attention if it repeatedly blocks important work. We would compare the burden of doing it today with the burden of operating the proposed workflow.
Repetition matters at the level of the action. Documents can look different while requiring the same fields. Identical forms can require very different decisions.
Check whether the inputs are stable enough
Gather approved examples of the actual material. Look for recurring document types, readable text, consistent field meanings, and a way to identify the current version.
Stable inputs do not have to look identical. They do need to contain the information the task requires. A supplier name that changes position on a page presents a different problem from a required reference that never arrives.
We would separate missing information from information that is merely hard to find. A clearer form or a required field may address the first problem. Software might help with the second. Asking a model to guess a missing value creates an answer without establishing that it is true.
Changing requirements also deserve attention. If staff disagree about which fields matter, settle that question before automating their collection. Otherwise each iteration may be testing a different job.
Separate copying from judgment
Write down what a person actually decides at each step.
If the work follows explicit rules, deterministic software may be the better choice. Moving known fields between systems, checking a required reference, or comparing a total with a defined limit does not inherently need AI. Existing software settings or a conventional integration may already support it.
AI may be worth testing when useful information arrives in varied language or document layouts. Its role could be preparing a proposed record with links to the source. That proposal still needs to be evaluated against the original material.
A decision that depends on negotiation, unwritten policy, or context held by an experienced employee is a weaker first candidate. We might narrow the task to gathering evidence for that person. Preparing information and authorizing an action should remain distinct boundaries.
Measure exceptions and the work of review
Sample ordinary cases and difficult ones. Mark which follow the agreed path, which lack information, and which require a different decision. Estimate the exception rate from those records and keep the sample's limits visible.
A high exception rate can mean the scope is too broad. Separating a predictable document type from a mixed queue may produce a better candidate. It can also reveal a process problem that software would merely pass along.
Review needs its own test. Can the reviewer see each proposed value beside the relevant source? Can they correct it, reject it, and identify missing information without reopening a chain of messages? Who receives an exception, and what happens if that person is unavailable?
We would observe the whole task, including corrections and handoff. Quick extraction is not enough if checking the result requires doing the original work again. Consequential actions should remain subject to the review their risk requires, regardless of how certain a generated answer sounds.
Confirm sensitivity, access, and ownership
Before selecting a tool, identify which information the workflow needs and which it can exclude. Sensitive records can change where a test may run, who may inspect it, what a provider may receive, and how long material may be retained. Use approved, minimized samples. Synthetic examples can help test structure, but they cannot establish performance on real variation.
System access can decide feasibility before model choice matters. Confirm whether the source can be read and whether the destination supports the needed fields and an approved connection. Reading an exported file does not prove that a production integration will work. Keep simulated connections explicit.
A first test can often stop at a draft for review. If writing to another system is essential to the question, use an approved test environment and examine duplicate handling, failed transfers, and recovery.
Name a workflow owner who can resolve policy questions and judge results. Also identify who will maintain rules, permissions, and exception routing. A useful demonstration without anyone responsible for its operation is not enough to justify expansion.
Put the first test in writing
Consider a hypothetical task: preparing an internal intake record from a defined type of service request. We would scope the test around that preparation step, leaving acceptance of the request with staff.
A short test brief should record:
- Purpose: the repeated work or blocked handoff we want to examine.
- Boundary: the document type, required fields, permitted sources, and actions outside the test.
- Samples: ordinary, incomplete, conflicting, and unreadable inputs, with some held back for evaluation.
- Output: a draft record, source references, and visible unresolved fields.
- Review: a named reviewer, correction path, and destination for exceptions.
- Decision criteria: acceptable field accuracy, review effort, exception handling, and failures that require stopping.
Set the criteria before seeing results. Compare the proposed workflow with the current process on comparable material. Include operating costs, access constraints, and maintenance effort in the decision, even if the initial experiment cannot establish them fully.
Keep inputs, outputs, corrections, and versions together. Change a specific rule or extraction step, rerun the relevant examples, and inspect the difference. For a model-based step, repeat cases to examine variation. A correction that helps one document should not silently break another.
Choose based on what the test shows
The strongest first candidate combines meaningful repeated work, available inputs, bounded decisions, manageable exceptions, reviewable results, approved access, and clear ownership. A serious gap in any of those can outweigh an attractive document count.
Expand only where the evidence supports it. If a clearer intake form removes the ambiguity, use that. If ordinary software handles the rules, a model may add little. If review remains too demanding, narrow the task or stop.
Before choosing your first workflow, you should be able to name the burden, show representative paperwork, identify the person who can judge the output, and describe the smallest test that could change your decision. That is enough to begin evaluating the right approach without committing the whole process to it.