The first version is not the decision
A capable developer can produce a working AI feature quickly. What follows — evaluation, monitoring, prompt maintenance, model changes, knowledge upkeep — is the actual commitment, and it is what the decision should be made on.
This is why comparing a quote against “we could build that ourselves in a fortnight” usually compares the wrong things. Both statements can be true, and neither addresses who is running the thing in year two.
In-house wins on proximity to the workflow
The hardest part is usually understanding the process being automated and the undocumented rules inside it. People who already know the business start ahead, and they are the ones who can extract the rules nobody wrote down.
It also wins on iteration speed once live. The feedback loop between noticing a bad output and changing something is measured in hours internally and in change requests externally, and that loop is most of what improves an AI feature in its first six months.
The scarce knowledge is how your business decides things, not how to call a model.
Outside help wins on the mistakes already made
Evaluation harnesses, guardrails, retrieval design and permission enforcement are patterns rather than inventions. Paying for the second attempt at those is cheaper than making the first one yourself.
The specific mistakes worth not making: indexing everything and discovering permissions later, shipping without an evaluation and having no way to judge a model upgrade, and building the interface before deciding what happens when the system doesn’t know. Each costs weeks to unwind and is entirely predictable to someone who has done it before.
Where each genuinely fails
In-house fails when the team has no slack. An AI feature added to a full roadmap gets built and then abandoned at exactly the point where it needed attention, and a half-maintained feature is worse than none.
Outside help fails when the supplier never gets access to the tacit knowledge — because the people who hold it are too busy for the interviews — and delivers something technically sound that encodes a process nobody actually follows.
The hybrid is usually right
Bring in help to establish the foundations and ship the first feature, with your team involved throughout, then run it internally. The condition is that handover means what it should — code, credentials, corrections and documentation.
Make the involvement contractual rather than aspirational. Someone from your team present for the workflow interviews, named as the future owner, and doing the review during the first month is the difference between inheriting a system and inheriting a black box.
Ask the maintenance question out loud
Who runs the evaluation, who updates the knowledge, who is called when a provider changes the model. If those answers do not exist, no amount of build quality helps — the ownership question again.
If the honest answer is that nobody internally can own it, that is a real finding and it points at a managed service rather than a build — bought with the questions worth asking any AI vendor, and accepting the trade that you control less.
Whoever builds it, keep the assets
The evaluation set and accumulated corrections are the durable value, not the code. A supplier who keeps those has locked you in far more effectively than any contract clause.
Write it into the engagement at the start: the evaluation set, the corrections data, the prompt history and the documented rules are yours, exportable, at any point. It is an easy clause to agree before work begins and an awkward conversation afterwards.