The credibility problem with the opposite advice
Everyone selling AI has an incentive to find a use for it. The useful version of that expertise is knowing where it loses to something simpler, because a badly-placed model is the fastest way to make a team distrust the whole category — and that distrust then blocks the projects that would have worked.
1. When the rule is actually a rule
If the logic can be written as conditions — thresholds, eligibility, routing by explicit criteria — write it as conditions. A model asked to apply a deterministic rule will apply it correctly almost every time, and almost is a downgrade from a rule that is right always.
The tell is a specification that reads as a decision table. If someone can describe the logic to you in the form “if this and this, then that, unless the other”, you have been handed working code in prose, and converting it to a prompt makes it slower, costlier and occasionally wrong.
The genuine hybrid is worth knowing: rules for the decision, a model for turning messy input into the fields the rules need. That is extraction, which is checkable, feeding logic that is deterministic.
If you can write it as an if-statement, a model is a slower, more expensive if-statement that occasionally disagrees.
2. When you can’t tolerate variance
Anything where the same input must produce the same output — financial calculation, compliance determinations, anything audited — wants deterministic code. Consistency is not something to prompt for; it is something to choose an architecture for.
This matters more than it sounds because the variance is usually small enough to pass testing and large enough to matter in aggregate. Two customers with identical circumstances receiving different answers is a fairness problem before it is a technical one, and it is very hard to explain after the fact.
3. When the volume doesn’t justify the maintenance
An AI feature carries permanent costs: evaluation, prompt upkeep, monitoring, review, and the work of responding when a provider changes something. A task performed twice a month cannot repay that, however well it demos, and the total running cost is mostly staff time.
A rough test before committing: multiply how often the task happens by how long it takes, and compare that against a few hours a month of maintenance forever. Low-frequency, high-judgment tasks fail this comfortably; high-frequency, low-judgment tasks pass it easily. The middle is where honest disagreement lives.
4. When nobody can verify the answer
If no one on the team can tell a correct output from a plausible one, you cannot evaluate it, cannot catch drift, and will not notice when it degrades. That is not a model problem — it is a missing evaluation and the feature should wait.
This is the condition most often waved away, usually with “we’ll spot-check it”. Spot-checking requires the same expertise as doing the task, so if that expertise isn’t available the check is theatre. It is also the specific failure mode behind research output that reads authoritatively and cannot be challenged by its reader.
5. When the real problem is upstream
Teams frequently reach for AI to cope with data that is contradictory, a process that is unclear, or documentation that doesn’t exist. A model does not fix any of those; it inherits them, and contradictions in the corpus produce effectively random answers.
Fix the upstream problem and quite often the AI feature is no longer needed. That is a good outcome, not a failed project — and it is why the knowledge work is worth doing even when the automation is later abandoned.
The cases that look like these and aren’t
Some rejections are wrong. “The output varies” is not disqualifying where the task is genuinely generative — drafting, summarising, suggesting — because there is no single correct answer to be consistent about. “We can’t verify it” is not disqualifying when verification is simply reading the source document the answer cites.
And low volume is not disqualifying where the value per instance is high: a task performed monthly that prevents an expensive mistake can repay maintenance easily. Test against value rather than frequency alone.
How to decline it usefully
Saying no to an AI project is more useful when it comes with the conditions that would change the answer. “Not until the policy documents agree with each other”, or “not at this volume, but revisit if it triples”, converts a rejection into a dated hypothesis rather than a dead end — the same discipline as recording why a feature was switched off.