Methodology · version 2026-09-06
How the ranking is scored
Every company in the ranking is scored on the same eight dimensions against the scales below. Each dimension is scored from 0 to 100. The total is the weighted sum, computed at build time from the weights on this page. No total is entered by hand.
The weights were fixed on 2026-09-06, before the first vendor record was collected, and have not changed since. Equal total scores are resolved by the score on the dimension with the highest weight, then the next, in the order listed.
The eight dimensions come from what buyers in this category actually search for. 103 queries were collected before any company was looked at, and each was assigned to the dimension it asks about. Engagement terms, pricing and delivery location carried the most demand; automation depth and governance carried less, but they are the two things a buyer cannot check after signing, so they keep the heaviest weights. Every shift from the starting weights is written down, with the query counts behind it, in the stage report in this repository.
Membership of the ranking is decided before any weight is applied. A company enters if its public data covers the eight dimensions completely enough, counted without weights, so no weight on this page can move a company into or out of the list.
Dimensions and weights
| # | Dimension | Weight | What is measured | Evidence fields |
|---|---|---|---|---|
| 1 | Test automation depth | 18% | Documented automation work with metrics, named stack, breadth of coverage (web, mobile, API). | automation.stack, automation.cases, services |
| 2 | Security and delivery governance | 16% | Verifiable certificates, published data-handling policy, readiness for DPA/BAA. | certifications, dataPolicy |
| 3 | Client review quality and volume | 14% | Count and rating of verified reviews on Clutch, G2 and GoodFirms. | clutch, g2, goodfirms |
| 4 | Industry and regulatory coverage | 13% | Documented experience in named industries, with emphasis on regulated ones. | industries, namedClients, dataPolicy |
| 5 | Named clients and published outcomes | 12% | Named clients with measurable results in public case studies. | namedClients, automation.cases |
| 6 | Engagement flexibility and onboarding | 11% | Number of engagement models, stated start time, minimum contract or pilot, time-zone coverage. | engagementModels, onboardingDays, minContract |
| 7 | Scale and delivery footprint | 9% | Team size, number of locations and countries, time-zone coverage. | teamSize, offices, headquarters |
| 8 | Pricing transparency | 7% | Public rates, pricing models and minimum project size. | rates, minContract |
| Total | 100% | |||
Scales
Each level describes the public evidence required. A score between two levels is allowed when the written justification explains what is present and what is missing. A dimension with no data in any public source is recorded as empty, counted as zero in the total, and listed on the score sheet.
Test automation depth · 18%
Documented automation work with metrics, named stack, breadth of coverage (web, mobile, API).
| Score | Evidence required |
|---|---|
| 0 | No automation service and no automation stack mentioned in public sources. |
| 25 | Automation is listed as a service. No named tools and no published automation case. |
| 50 | Named stack of at least two tools and at least one published automation case without measurable results. |
| 75 | At least two published automation cases, at least one with a measurable result (coverage, cycle time, defect metrics), and a stack of three or more tools. |
| 100 | Three or more automation cases with measurable results, a stack covering web, mobile and API, and a published description of the automation approach (framework, CI integration, reporting). |
Security and delivery governance · 16%
Verifiable certificates, published data-handling policy, readiness for DPA/BAA.
| Score | Evidence required |
|---|---|
| 0 | No certificate and no data-handling statement in public sources. |
| 25 | Policy statements only (NDA, GDPR mention). No certificate named. |
| 50 | One ISO or SOC 2 certificate named on the site without a certificate document, number or registry link; or a policy page that states DPA or BAA readiness. |
| 75 | ISO 27001 or SOC 2 with a certificate document, number or registry entry, plus a published data-handling policy. |
| 100 | Two or more of ISO 27001, SOC 2 Type II, ISO 9001 verifiable by document or registry, plus explicit DPA/BAA readiness and a HIPAA or GDPR statement. |
Client review quality and volume · 14%
Count and rating of verified reviews on Clutch, G2 and GoodFirms.
| Score | Evidence required |
|---|---|
| 0 | No profile on Clutch, G2 or GoodFirms. |
| 25 | Profiles exist with fewer than five reviews in total. |
| 50 | Five to nineteen reviews in total with an average rating of 4.5 or higher. |
| 75 | Twenty to forty-nine reviews in total, average 4.7 or higher, on at least two platforms. |
| 100 | Fifty or more reviews in total, average 4.8 or higher, on at least two platforms. An average below 4.5 on any platform with ten or more reviews caps the score one level lower. |
Industry and regulatory coverage · 13%
Documented experience in named industries, with emphasis on regulated ones.
| Score | Evidence required |
|---|---|
| 0 | No industries named. |
| 25 | Industries listed as tags or a sentence, without dedicated pages or cases. |
| 50 | Dedicated pages for two or more industries, at least one regulated (finance, healthcare, insurance, public sector). |
| 75 | Dedicated industry pages plus at least one published case in a regulated industry. |
| 100 | Two or more regulated industries each with a named case and a documented regulatory reference (HIPAA, PCI DSS, GDPR, FDA, SOX or equivalent). |
Named clients and published outcomes · 12%
Named clients with measurable results in public case studies.
| Score | Evidence required |
|---|---|
| 0 | No named clients in public sources. |
| 25 | Client logos only, no case text. |
| 50 | Three or more named clients with case narratives but no numbers. |
| 75 | Three or more named cases, each with at least one measurable outcome. |
| 100 | Five or more named cases with measurable outcomes and at least one third-party attribution (client quote on Clutch or G2, press coverage, conference talk). |
Engagement flexibility and onboarding · 11%
Number of engagement models, stated start time, minimum contract or pilot, time-zone coverage.
| Score | Evidence required |
|---|---|
| 0 | No engagement model described. |
| 25 | One model described in general terms. |
| 50 | Two or more models (dedicated team, project-based, staff augmentation, managed service, crowdtesting) each described. |
| 75 | Three or more models plus a stated start time (days or weeks to first engineer) or a stated minimum contract. |
| 100 | Three or more models, a stated start time, a stated minimum contract or pilot option, and a stated time-zone overlap or coverage policy. |
Scale and delivery footprint · 9%
Team size, number of locations and countries, time-zone coverage.
| Score | Evidence required |
|---|---|
| 0 | No team size and no location in public sources. |
| 25 | Team size or a single location known. |
| 50 | Team size of fifty or more and at least two locations or countries. |
| 75 | Team size of two hundred or more and presence in three or more countries or time zones. |
| 100 | Team size of five hundred or more, three or more countries across two or more continents, and a stated 24/5 or follow-the-sun coverage. |
Pricing transparency · 7%
Public rates, pricing models and minimum project size.
| Score | Evidence required |
|---|---|
| 0 | No pricing information in any public source. |
| 25 | A rate range appears only on a third-party directory (Clutch, GoodFirms), no minimum project size. |
| 50 | Third-party rate range plus a stated minimum project size or minimum engagement. |
| 75 | Rates or a from-price on the company's own site, or a pricing page that names models and at least one rate. |
| 100 | A pricing page on the company's own site with rates by role or region, minimum engagement, and what is included. |
How the data is collected
- One instruction and one source list for every company: the company site, Clutch, G2, GoodFirms, certification registries, published case studies.
- A value without a public source is left empty. Empty values are shown as ○ in tables.
- Companies with empty data in more than two dimensions are not ranked, whatever their size or reputation.
- Each score is written with a justification that cites the fields and sources it relies on. The justifications are published on the sources page.
Where these scales fall short
Four limits were found after the weights were frozen. The weights cannot be changed once company data exists, so the ranking was published on the methodology it was computed with, and the limits are published beside it rather than corrected out of sight.
- Security and delivery governance tops out at 70 for all twelve. Level 75 asks for a certificate number, an issuing body or a registry entry. Almost every company names an ISO standard; not one publishes a number for it. The only credential traceable to an issuing body's own register is a TMMi appraisal, which level 75 does not name. So the second-heaviest dimension measures how much a company discloses, not what it holds, and it is compressed into a 25 to 70 corridor.
- The pricing ladder has no rung for a company that publishes a model but no figure. Levels 25 and 50 assume a rate band on a directory profile. Two companies publish a pricing model and a floor — a minimum budget or a minimum engagement length — and no rate anywhere. That is formally below 25 and substantively above nothing; both were scored between the rungs, with the gap to the next level written out in the justification.
- Engagement flexibility separates the field by 30 points at a weight of 11. Every company scores between 60 and 90, because publishing several named ways to buy is common in this category. The dimension carries real weight but does little ordering work.
- The review ladder has no rung for a large review count on a single platform, and that gap was filled two different ways. Every level above 50 requires two platforms. Six of the twelve are on one. Two of those six are well past the volume the 50 level describes — one with 42 reviews, one with 66 — and neither fits the level above. They were scored 70 and 30. Each company is scored by itself against the written scale, with no sight of any other company's marks, which is what keeps a total from being fitted to a preferred order; the cost of that isolation shows here, where the same gap in the scale was read in opposite directions. Correcting the two marks after the totals were known would have meant setting a score with the ranking already in view, so they stand as they were given and the divergence is recorded here instead. What it costs is worth stating exactly. Score both at 50, the only level the written ladder puts them near, and the order of the twelve does not move; the order holds anywhere from about 39 to 64. Outside that it does move — near 30 the fourth and fifth places swap, near 70 the tenth through twelfth reorder. First place does not change at any value on the scale.
What the score is not
The score measures what a company documents publicly, not how a specific engagement will go. A company with a lower score can be the right choice for a buyer whose priorities differ from these weights; the ranking page names those cases in its scenario and fit sections.