How the ranking works
The scorecard puts 30 products in order, which is worth something only if you can see how the order was produced. This page gives the six criteria, the weights, the scoring bands, the evidence that counted, a worked calculation you can check on a calculator, and a plain account of what this method cannot tell you.
Weights first, scores second
The six weights below were fixed and written down before any product was assessed. This is the only real protection against a rubric being built backwards to justify a winner that was already chosen, and it matters more than the specific numbers. A rubric assembled after scoring is an opinion wearing arithmetic as a costume.
The second protection is publishing every sub-score. A weighted average is easy to manipulate while the inputs stay hidden. All 30 rows in the scorecard print their six sub-scores, so any figure can be recomputed and any disagreement can be aimed at a specific number rather than at the result.
The six weights
These criteria are written for the person who buys, installs and runs address software. That is a different person from the one who integrates an address API, and the rubric reflects it.
| Criterion | Weight |
|---|---|
| Deployment options | 20% |
| Usability without developers | 20% |
| Data quality and certification | 20% |
| Batch and list handling | 15% |
| Pricing transparency | 15% |
| Onboarding and support | 10% |
Why these weights
The three criteria at 20 percent decide whether the software solves the problem at all. Software you cannot deploy the way your organisation works, that your team cannot operate without engineering help, or whose output you cannot trust, has failed regardless of price.
Batch handling and pricing transparency sit at 15 percent because they are real but recoverable. A weak batch tool can be worked around by exporting and re-importing. Opaque pricing can be dragged into the open through procurement, at a cost in time.
Support sits at 10 percent because it varies more by contract tier than by vendor. The same company often delivers attentive support to an enterprise account and a ticket queue to everybody else, which makes it the least stable thing on the list to score.
1. Deployment options. 20 percent
Whether the product runs as a hosted service, as installed software on a desktop, inside your own network, or as a module of a platform you already license. Scored on the range offered rather than on any single mode, because a buyer with a hard constraint has usually lost most of the market before they start. On-premises capability lifts this sub-score substantially, since roughly one buyer in five cannot send addresses to a third party at all.
2. Usability without developers. 20 percent
Whether a marketing or operations person can run the thing alone. Scored on the presence of a real interface, the number of steps between a file and a cleaned file, and whether routine work needs engineering time. A product that needs a developer for every list upload scores poorly here even when the underlying verification is excellent, because that cost recurs forever.
3. Data quality and certification. 20 percent
No head-to-head accuracy test was run, so this criterion scores the strongest available proxies. Held postal certifications carry the most weight because an external body audits them: USPS CASS, Royal Mail PAF and the equivalents elsewhere. After that comes how direct the reference data relationship is, how often that data refreshes, and how specific any published accuracy evidence is. A vendor licensing another vendor's data scores below the vendor that holds it.
4. Batch and list handling. 15 percent
Maximum practical file size, throughput, what comes back beyond a corrected address, and whether deduplication and change of address flags are included or sold separately. Weighted because list cleaning is the job most often brought to software rather than to an API.
5. Pricing transparency. 15 percent
Whether a buyer can work out the cost without a sales call, and whether the published shape survives contact with real volume. Published per-record or per-tier pricing scores highest. A quote-only enterprise product scores low here even when it turns out to be good value, because the buyer cannot know that in advance.
6. Onboarding and support. 10 percent
Time to a working setup, quality of documentation written for non-developers, and what support exists below the enterprise tier. Email-only support for a product sold to operations teams is scored as the limitation it is.
What each sub-score means
Published bands, so that a 7.4 means the same thing in every row. Sub-scores move in steps of 0.2, which is the finest distinction this evidence can honestly support.
| Band | Meaning |
|---|---|
| 9.0 to 10 | Best available. Leads the field on this criterion, with published evidence behind the claim. |
| 8.0 to 8.9 | Strong. Does this well, with a documented gap or two against the leaders. |
| 7.0 to 7.9 | Solid. Meets the need for most buyers without standing out. |
| 6.0 to 6.9 | Adequate. Works, with limitations a buyer would notice inside a month. |
| 4.0 to 5.9 | Weak. Present but thin, or too undocumented to verify. |
| 0 to 3.9 | Absent or close to it. Used where a capability is missing rather than merely limited. |
A worked calculation
Using the top-ranked product, so the most consequential number on the site is the one shown being derived.
| Criterion | Sub-score | Weight | Contribution |
|---|---|---|---|
| Deployment options | 9.2 | 20% | 1.84 |
| Usability without developers | 9.6 | 20% | 1.92 |
| Data quality and certification | 9.4 | 20% | 1.88 |
| Batch and list handling | 9.4 | 15% | 1.41 |
| Pricing transparency | 9.0 | 15% | 1.35 |
| Onboarding and support | 9.2 | 10% | 0.92 |
| Total | 100 percent | 9.32 |
9.32 rounds to the published 9.3. Every other row works the same way, and every input appears in the scorecard.
Reordering the field around your own priorities takes about ten minutes in a spreadsheet. Enter whatever weights reflect what you actually care about, divide each one by their new total so the set adds up to one, then multiply through the sub-scores exactly as printed above. Nothing else needs to change, because the sub-scores are judgements about the products rather than about the weighting.
How the field was narrowed
More than 35 products were catalogued from software directories, vendor sites and the overlapping category listings most buyers start from. Thirty were scored. The rest were excluded for reasons worth stating rather than hiding.
Discontinued or absorbed products. Several named products now redirect to a parent company's platform. Scoring something a buyer cannot purchase wastes their time.
Pure geocoding services. Turning an address into coordinates is a different job from confirming that an address is deliverable. Products that only geocode were left out, while products doing both were scored on the verification half.
Products with no published evidence. Where a vendor publishes no documentation, no pricing shape and no coverage detail, there is nothing to score. Assigning numbers anyway would be invention.
Two awkward fits were kept. A pair of entries near the bottom are not verification software in the strict sense. They are scored because buyers keep shortlisting them, and a comparison that silently omits what people are actually considering helps less than one explaining why those options rank where they do.
What counted, in order
- Externally audited factsPostal certifications and compliance attestations. These can be checked against the certifying body rather than the vendor, which is what puts them at the top of the hierarchy.
- Published vendor documentationTechnical documentation, pricing pages and support policies. Treated as accurate about what the product does and treated sceptically about how well it does it.
- Directory listings and user reviewsUseful for support quality and onboarding friction, which vendors never document honestly. Discounted for accuracy claims, because reviewers rarely measure match rates.
- Vendor case studiesLowest weight, discounted heavily wherever the vendor supplies both the metric and the method used to produce it.
Coverage counts, throughput figures and accuracy percentages anywhere on this site are vendor claims unless attributed otherwise. They are labelled that way rather than presented as findings.
What this method cannot tell you
No accuracy test was run. The data quality criterion scores proxies for accuracy, and a proxy is not a measurement. If accuracy on your data is the deciding factor, test the shortlist on your own worst records. Most vendors offer a free tier sized for exactly that.
Coverage claims are not comparable. A vendor advertising 250 countries might confirm exact delivery points in a few dozen of them and recognise nothing finer than a city name everywhere else. Both situations get described with the same number. Granularity was weighted wherever a vendor published it, and they publish it unevenly.
Support scores are the least stable figures here. Support quality tracks contract tier more closely than it tracks vendor, so a low support sub-score may say little about what an enterprise account receives.
Pricing moves. Every figure reflects what was published in September 2026. Treat published pricing as a starting position rather than a quote.
A weighted average compresses real differences. Two products separated by 0.2 are not meaningfully apart. The scorecard is a way to see the field in order, and the fit table is the better tool once you have a shortlist.
Questions about the method
How is the score calculated?
Each product gets a sub-score from 0 to 10 on the six criteria above. The published score is the weighted arithmetic mean, rounded to one decimal place. Every sub-score is printed in the scorecard, so any row can be recomputed by hand.
Were the weights chosen before or after scoring?
Before, and fixed in writing. Choosing weights afterwards is how a rubric becomes a justification for a result somebody already wanted.
Was an accuracy test run?
No, and the ranking does not claim one. Held certifications, reference data relationships, refresh cadence and published accuracy evidence are scored as proxies. Certifications carry the most weight because an external body audits them.
Why does this differ from developer-focused API comparisons?
The criteria differ, and they should. This rubric rewards deployment breadth, operation by non-technical staff and batch handling. A developer rubric weights SDK quality, latency and documentation. The two produce different orders, most visibly among desktop and on-premises products, which are strong here and weak there.
Can I recompute the table with my own weights?
Yes, which is why the sub-scores are published. Swap in weights that match what you care about, divide each by their new total so the set adds up to one, then multiply through the printed sub-scores. Pushing deployment to 40 percent for a hard on-premises requirement reorders the table substantially.