Scorecard

Build vs Buy Scorecard

Building retrieval is a weekend to a demo and a year to something reliable. This scores the decision on the factors that actually decide it — evaluation, maintenance and who is on call — rather than on infrastructure cost.

Ask about CreobotSee all tools

YOU GIVE ITsix answers about your teamreadcomputereturnWHAT THE SCORE WEIGHSteam capacityevaluationmaintenanceinfrastructure cost is the smallest inputYOU GET BACKa score weighted to maintenance, not buildRuns in your browser. Nothing is sent anywhere.
  • Weighs the real costsEvaluation and on-call
  • Honest either waySometimes build wins
  • Takes two minutesSix questions

Work it out

Your numbers

Result

—out of 100, higher favours building
  • Verdict —
  • Team capacity —
  • Evaluation readiness —
  • Maintenance load —
  • Biggest risk —
  • Do next —

Infrastructure cost is deliberately excluded. It is the smallest and most quoted number in this decision, and it is almost never what makes a self-build fail.

What the numbers mean

Evaluation decides this, not infrastructure

A pipeline you cannot measure will degrade and nobody will notice until customers do. If you have no scored test set, that is the first thing to build regardless of which way you go.

A demo is a weekend, reliability is a year

Retrieval that works on ten documents in a notebook is genuinely easy. Retrieval that stays correct across a changing corpus with real traffic is the actual project.

Content churn is the hidden cost

Every content change means re-indexing, re-checking and re-evaluating. Weekly churn with no dedicated owner is where most self-builds quietly stop being maintained.

Sometimes build genuinely wins

Unusual formats, strict residency, an existing vector store, or retrieval being the product itself. This scorecard says build when those hold, and we would rather it did than pretend otherwise.

Questions

It weights team capacity, evaluation and strategic fit, and it returns build when those are strong. We would rather be trusted on the cases where building is right than win an argument we should lose.

Because it is the smallest number in the decision and the most quoted. Use the RAG Cost Estimator for it separately; it rarely changes the answer.

Then the question is whether you can evaluate it. Run the test question generator against it. If it scores well and someone owns it, keep it.

Want this run against your actual site?

Tell us the page your assistant lives on and the questions you care about. A person runs them and sends back the raw answers.