Without a scorecard, software decisions are made on recency, presenter charisma and the preference of whoever is most senior in the room. A scorecard does not remove judgement; it makes the judgement explicit and reviewable.
Design principles
Weight before you look. Assign weights to criteria before seeing any vendor. Weights set afterwards describe your preference rather than your requirements.
Score evidence, not claims. A score above "partial" requires something you observed.
Score by section. Overall totals hide a vendor that is strong everywhere except the area you cannot compromise on.
Score independently, then discuss. Individual scores first, then reconcile. Group scoring converges on the first opinion voiced.
Record why. A one-line justification per score. Six weeks later, when the decision is questioned, the justifications are what defend it.
A feature promised for the next release is not a feature. Score it as partial with a note, and if it is critical, require a contractual commitment with a date and a remedy.
A workable structure
| Section | Weight | Scored by |
|---|---|---|
| Planning and optimisation | 20% | Planner, operations manager |
| Execution and dispatch | 10% | Dispatcher |
| Driver mobile app | 20% | Drivers, operations |
| Integration and technical | 12% | IT |
| Reporting and analytics | 8% | Operations, finance |
| Administration and security | 5% | IT, compliance |
| Implementation approach and risk | 10% | Project lead |
| Total cost of ownership | 10% | Finance |
| Vendor viability and support | 5% | Project lead |
Adjust to your operation. A route accounting purchase would weight pricing and settlement heavily; a maintenance-focused purchase would weight the workshop module.
Note the driver app weight. In operations with many stops per day, the app is used thousands of times a week; the planning interface is used once. Weight accordingly, and let drivers score that section.
The scoring scale
| Score | Meaning |
|---|---|
| 0 | Absent |
| 1 | Partial — workaround required, or on the roadmap |
| 2 | Adequate — meets the requirement |
| 3 | Strong — meets it well, with evidence |
Four points is enough. Ten-point scales produce false precision and endless argument about whether something is a 6 or a 7.
Mandatory gates
Separate from scoring, define pass/fail gates. A vendor failing any gate is eliminated regardless of total score:
- Offline operation of the driver app for a full shift
- Data export in a documented format at no cost
- Support hours covering your operating hours
- Required integrations demonstrably possible
- Security certification your policy requires
- Financial viability threshold
Gates prevent the outcome where a high-scoring vendor is selected despite failing something you cannot live without.
Running the scoring
- Circulate the scorecard with weights before the first demonstration.
- Each evaluator scores independently within 30 minutes of each session.
- Collect scores before discussion.
- Review divergences — where two evaluators differ by two or more points, discuss. Divergence usually means the requirement was ambiguous or one person saw something the other did not.
- Agree a consensus score with a recorded justification.
- Compute weighted totals by section and overall.
- Review the result. If it contradicts everyone's instinct, examine the weights — sometimes the weights were wrong, and sometimes the instinct was.
Presenting the decision
A one-page summary that survives scrutiny:
- The weighted scores by section, all vendors
- Gate results
- Five-year total cost for each
- The three most significant differentiators, with evidence
- Key risks of the recommended option and how they will be managed
- The recommendation and the reasoning
Attach the detailed scorecard. The summary is for the decision meeting; the detail is for the questions afterwards.
Common questions
How many people should score?
Four to eight, covering the affected functions, including at least one driver and at least one daily user of the planning system. Larger panels dilute accountability; smaller ones miss perspectives.
What if the cheapest vendor scores lowest?
That is the scorecard working. Present total cost of ownership alongside the score so the trade-off is explicit, and quantify what the score difference means operationally — usually in hours, failures or risk rather than in abstract points.
Should vendors see the scorecard?
Share the criteria and weights; withhold the scores. Transparent criteria produce better-targeted proposals and demonstrate a fair process, which matters if a decision is challenged.
How do we handle a tie?
Look at the section scores rather than the total, weight the sections that matter most in your operation, then use implementation risk and vendor viability as tie-breakers. A genuine tie usually means either option would work, in which case commercial terms and cultural fit should decide.
Can we change weights mid-process?
Only with a documented reason applied consistently to all vendors and re-scored. Changing weights after seeing results to produce a preferred outcome destroys the purpose of the exercise and will be visible to anyone who reviews it.