
Every provider, Superside included, is scored on the same 100-point rubric, built entirely from public evidence. Four pillars. Yes/no or small-tier checks. No free-form judgment calls. Here's how it works, and why the weightings shift when the topic does.
If you've ever landed on a "Top 10 [X]" article and wondered how the ranking actually got made: Welcome. So did we, which is why we built one.
Our editorial team runs a bunch of comparison lists:
- Best AI design tools
- Best AI design agencies
- Best AI presentation makers
- Best AI ad creative tools
- And more.
Although we also have great editorial pieces with top-of-the-industry leadership insights, we don’t only rely on our authority.
This means that each list we publish gets scrutiny.
From readers, from the companies included, from us.
So we created a rubric that could survive it. And for you to assess it as well.
This article is a look under the hood. Learn about our key pillars, why we chose them and how we adapt them per topic.
The objective: comparable scores from public evidence
The rubric has one job — produce a score that two independent researchers would land on given the same evidence. That means every check is either a yes/no or a small tiered value. No "we felt 7/10 was fair." No hidden vibes.
It also means we only look at what a buyer can look at: G2, Clutch, Trustpilot, the provider's own site, published research, press coverage, pricing pages. If it's behind a paywall or only in a vendor's press release, it doesn't count.
And to keep ourselves honest: Superside is scored on the exact same rubric. No home-team advantage.
The four pillars (and why they exist)
Each provider is graded on 100 points across four pillars. The weights aren't arbitrary: they map to the questions a buyer actually asks when they're picking a creative partner.
1. Client Evidence & Social Proof — 25 pts
The question: Can we trust that real buyers vouch for them?
This is the trust anchor. It's weighted toward third-party review platforms (G2, Clutch, Trustpilot, with Capterra as a fallback). The single biggest check here is the third-party rating itself. It's the one signal a competitor can't manufacture.
Case studies also count, but only when they show quantified outcomes. Logo walls score lower than "we cut cost-per-lead by 42% for a Fortune 500" — because that's how buyers actually evaluate.
2. Market Authority — 20 pts
The question: Do they shape the category or just operate in it?
Three checks, each free and fast to verify: original research or reports, an active content hub and awards or third-party press coverage. Original research earns the most points because it's the hardest to do: it earns citations rather than chasing them.
We deliberately left out things like paywalled domain-rating tools and self-reported analyst mentions. They can't be independently confirmed, and the tools that measure them often disagree with each other.
3. Capability & Fit — 35 pts
The question: Can they do this job at scale?
The heaviest pillar. A comparison list exists to help a buyer choose, so fit-to-buyer outranks any single credibility or pricing signal.
It combines core-service depth (do they actually offer the thing this list is about, with real depth), complementary breadth, AI capability, and five buyer-fit signals: enterprise page, global delivery, a trust/security page, a visible team or talent model, and scalable ongoing capacity.
Together they answer the question a buyer really asks — "can this company actually take on/support my work?"
4. Buyer Accessibility — 20 pts
This pillar measures buyer friction. Pricing disclosure carries the most weight here because published pricing is the clearest self-serve signal. But we don't penalize sales-led models: clear service scope, a stated engagement model (project, retainer, subscription), and a clean starting path each earn their own points.
The question: Can a buyer realistically engage without friction?
Why the pillars behave differently by topic
The rubric is the same across every list. What changes is what "fit" means — and that alone shifts the shape of the results.
Take two examples:
- Best AI design tools — a specialist list. "Core-service fit" means depth in a specific discipline (AI-powered design). A provider that offers 15 loosely related services doesn't get rewarded for breadth here.
- Best AI design agencies — a generalist list. Breadth is the offer. Being an all-around creative partner is the point, so the same checks reward a different shape of company.
The same logic applies to Pillar 1: "case-study depth" on a specialist list means case studies in that discipline, not the company's greatest hits from an unrelated category.
That's why the same provider can rank differently on two of our lists without either result being inconsistent. The pillars are stable, but the lens tightens or widens with the topic.
How the score is reported
We don't publish raw numbers. A total translates into one of five documentation bands:
- Fully documented — 85–100
- Well documented — 70–84
- Partially documented — 55–69
- Limited documentation — 40–54
- Minimal documentation — under 40
Two things worth noting about that language:
- The bands say "documented," never "good" or "bad." A pitch-based agency doing brilliant work can land in a lower band simply because it doesn't publish G2 reviews, case studies or pricing.
- The rubric has a known bias toward transparent, scaled providers. We think that trade-off is worth it — buyers reading a list blog post need signals they can independently verify — but we'd rather be explicit than pretend the bias doesn't exist.
What this methodology doesn't cover
Public-evidence scoring can't tell you what it's like to actually work with a provider. It can tell you who has the receipts to show up in a comparison, who publishes enough to be evaluated fairly and who has structured themselves to be findable by buyers. It can't tell you whose account manager returns emails on a Friday afternoon.
That's why our lists pair the score with editorial commentary. The rubric gets you a ranked shortlist grounded in facts. The commentary tells you which shortlist entry to call.
If you want to see the rubric in action, start with one of our recent roundups — best AI design agencies, best AI design tools, or best AI ad creative generators — and read the tier language in each entry. That's the rubric talking.
FAQs








