/ The Bottom Line Up Front

A model can know your name and still describe you as the wrong kind of company.

Recognition and correct description are two different measurements. In the 1,000-brand study behind this tool, 510 brands were volunteered unprompted and 194 were filed under the wrong category. Zoho is both: named readily, read well, and ranked as CI/CD rather than CRM.

What each probe measures

The audit reads an open model directly and reports what each probe found:

  • Whether the model volunteers your brand unprompted, and which category it files you under
  • How it regards you on twelve buying dimensions, read from activation space rather than generated text
  • How far your brand name surfaces as the signal travels up the model’s layers
Read how the ModelRank brand audit works

How the ModelRank brand audit works

01

Name your domain

We read a handful of your pages to place you in the category taxonomy, and you can name up to seven competitors or let the crawl suggest them.

02

Four probes run live

Recall, category binding, activation-space sentiment and the logit lens. Each one streams back the moment it lands, so the report fills in as it is measured.

03

The channel phase queues

Third-party surfaces, backlinks and own-property facts for you and every named competitor, measured per brand and placed as an ordinal among the peers you named.

04

The audit keeps a link

Every run is stored under its own key, so the finished report can be reopened and shared without running the model again.

What the ModelRank audit looks like

A real audit for Stripe, rendered by the same components your own report uses. Your audit measures your domain live and produces a report of exactly this shape.

Example report for Stripe

ModelRank/payments

79

Stripe

Recitable

deeply imprinted, reproduced on demand

#1 of 6

in payments, the rank that matters

The verdict
Established

The model holds this brand, volunteers it, and files it under the right category.

What to do about this3 actions from this audit, in priority order
  1. 1.

    Name third-party integrations on the features page

    Stripe has no integrations page the crawl could reach.

    Memory depth
  2. 2.

    Publish reference documentation that answers a task end to end

    The model reads Stripe below its named peers on documentation.

    Sentiment
  3. 3.

    Say what the product is in the opening sentence of every page you own

    The brands at the core of payments state their category in their first sentence; Stripe does not.

    Where you sit

Measured on google/gemma-4-12B on one date; every model holds a different imprint of the same brand. Read the method

How much of Stripe the weights hold

M, the first of the four parts of the score, and the one with the most weight on it. Two probes measure it: whether the model volunteers the name at all when asked about the category, and how much of the brand's own published surface is reproducible from the weights.

What the model associates with you

Each node is a named concept axis, fitted from hand-written contrastive prompt pairs before any brand was measured, so the axis means what it says rather than whatever an unsupervised feature happened to latch onto. Node area is how firmly the model holds that axis for Stripe; node state is what to do about it. Their agreement with the payments axis profile is C in the score: 50.0 of 100, read here from the share of Stripe's measured surfaces that also state what it does, which is the live equivalent of the cohort's concept alignment rather than the alignment itself.

Overlay a competitor

The same selection drives the perception radar above and the breakdown table below.

What the model holds, and what to do about it

Every one of the 12 measured axes, grouped by the action this run wrote on it. One state at a time: the tab says how many axes sit behind it.

StripeEase of useEase of use, already held, alignment +1.824ReliabilityReliability, already held, alignment +1.631Pricing valuePricing value, already held, alignment +1.317Trust and securityTrust and security, already held, alignment +0.996InnovationInnovation, target, alignment -0.919TrajectoryTrajectory, already held, alignment +0.793CommunityCommunity, already held, alignment +0.689IntegrationsIntegrations, strengthen, alignment +0.510DocumentationDocumentation, strengthen, alignment -0.299SupportSupport, strengthen, alignment +0.293
  • Target(2)
  • Strengthen(4)
  • Already held(6)
  • How firmly the axis is held

The model's sentiment towards Stripe

Twelve buying dimensions, each read directly from the model's activation space rather than asked in natural language. Green means the model reads Stripe above the cohort on that dimension, red below it. Their mean, scaled to the same clamp the charts here draw against, is S in the score: +0.76.

Compare against a competitor

The same selection drives the concept action network and the breakdown table below.

DocumentationReliabilityIntegrationsProduct depthSupportPricing valueEase of useTrust and security
Stripe, solidAdyen, category leader

How this is measured

Sentiment is read as a direction inside the model's own activation space: contrast pairs of positive and negative words fit a single axis, and the brand's activation is projected onto it, following Tigges et al. (2023). The grey shape overlaid on the radar is Adyen, the highest-scoring brand in payments.

Strongest read

Trajectory: the model reads Stripe above the cohort here, above the payments peer median. This is the claim the model is most prepared to support unprompted.

Weakest read

Ease of use: the model reads Stripe below the cohort here, still above the payments peer median. When an assistant hedges about Stripe, this dimension is the likeliest reason.

What to do about this1 item
  • The model reads Stripe below its named peers on documentation.
    Close this

    Publish reference documentation that answers a task end to end

    Runnable examples, error tables and migration notes on pages that are crawlable without a login, rather than an API reference that assumes the reader already knows the product.

    https://stripe.com/docs

These actions were planned from this audit’s own measurements: the surfaces the crawl found, what your own pages say in their first sentences, and where the twelve dimensions place you against the competitors you named. They are readings of what the model has absorbed, not interventions anybody has run, so none of them promises a movement in a score.

Where you sit, and what separates you

Stripe and its payments peers, placed by a 2D projection of the model's representation space. The map answers two questions: where the model puts Stripe, and which brands it is most likely to be confused with. What separates it from the brands at the core of payments, and what closing that would require, is beside it in units the projection cannot distort.

payments hubStripeStripe 85Stripe, ModelRank 85AdyenAdyen 43Adyen, ModelRank 43AuthorizeAuthorize 43Authorize, ModelRank 43CheckoutCheckout 38Checkout, ModelRank 38BraintreepaymentsBraintreepayments 38Braintreepayments, ModelRank 38SquareupSquareup 38Squareup, ModelRank 38
  • Stripe
  • Category peer
  • Middle of payments
  • ModelRank
  • Sentiment, positive to negative

Scroll or pinch to zoom, drag to pan. A 2D projection, so read it for who sits next to whom, not for how far apart they are.

The gap to the core of payments

The model places Stripe next to Adyen, Braintreepayments, Squareup, and those are the brands it is most likely to be confused with. At the core of payments it holds Adyen, Authorize, Checkout, alongside Stripe: a reference set named rather than averaged into a centroid.

Stripe scores 79, at or above the 62 at the core of payments. There is no gap to close here; the parts below are where the remaining differences are.

Memory, co-occurrence and sentiment, summed+17.2 points
How to read this gapPoints of ModelRank, not projection distance; three components sum, accuracy sits alongside

The headline is a difference in ModelRank score points, because the map beside it is a 2D projection whose distances are not meaningful and cannot honestly be measured off. The reference set is three brands by name rather than an average, so every number here can be checked against those brands. Memory depth, category co-occurrence and sentiment valence sum to the score gap exactly; factual accuracy multiplies the verified score rather than the free one, so it is reported alongside the sum and contributes nothing to the number this card leads with.

These are associations measured across the cohort, not effects anyone has intervened to produce, and none of them promises a score movement. One of them is a warning rather than a lever: social prominence is associated with being named and with nothing about being understood, so volume there can raise recall while leaving the category binding and the read exactly where they are.

Does the model have its facts right

The model correctly binds Stripe to payments, preferring it over four length-matched distractor categories. With a margin of 2.64 over the runner-up, the binding is held firmly.

This is A in the score, read as 1.00 on 0 to 1: zero for a deviated or unaddressed binding, and rising with the margin of a consistent one until it reaches 1.00 at a margin of 0.83.

Expand the exact test we ran

We score five complete sentences of the form “Stripe is a payments platform.” and rank them by mean per-token log probability: the true category against four distractors drawn from the cohort's 333 categories. Distractors are matched to the true label's length, so a short or common category name cannot win on fluency alone. Nothing is generated and nothing is asked in natural language, so the result cannot be shaped by how the model chooses to phrase an answer.

The margin of 2.64 is the gap between the winner and the runner-up. Above 0.83, the median among correct bindings in this cohort, we call a binding firmly held; below it, weakly held.

See how we measure and action this for our clients

Brand perception
Category
HubSpotYOU
Salesforce
PipedrivePipedrive
AttioAttio
Gap
Awareness
81
73
77
68
+4
Competitive
79
64
60
69
+10
Discovery
44
61
30
11
−17
Evaluation
72
66
70
69
+2
Features
68
63
66
61
+2

Competitive breakdown

Stripe against its payments peers: the composite, the assessment the same rule reaches everywhere on this page, whether the model names each brand at all, whether it files each one under the right category, and how it reads them. Same measurements, same model, ordered by ModelRank. Select any competitor's name to compare Stripe against it, which drives the radar, the dimension bars and the concept comparison above.

BrandModelRankAssessmentRecallCategorySentiment
StripeYou85EstablishedThe model names Stripe unprompted when asked what exists in its categoryThe model binds Stripe to payments, which is correct+1.90
43EstablishedThe model names Adyen unprompted when asked what exists in its categoryThe model binds Adyen to payments, which is correct+1.09
38AbsentThe model never named Checkout in any sampleThe model binds Checkout to payments, which is correct+0.45
38AbsentThe model never named Braintreepayments in any sampleThe model binds Braintreepayments to payments, which is correct+1.32
43EstablishedThe model names Authorize unprompted when asked what exists in its categoryThe model binds Authorize to payments, which is correct+2.34
38AbsentThe model never named Squareup in any sampleThe model binds Squareup to payments, which is correct+0.88

Read how the ModelRank brand audit works

Frequently asked questions

What is a ModelRank brand audit?

It is a direct measurement of what an open weights model holds about a brand: whether the brand is volunteered when the model is asked what exists in its category, which category the model ranks highest for it, how the model reads it on twelve buying dimensions, and how far up the layers the brand name surfaces. It reads the model rather than grading a chat answer.

How is this different from asking ChatGPT about my brand?

Asking a chat assistant gives you one sample of a stochastic generation, shaped by the system prompt, the retrieval layer and the conversation so far. The audit measures the weights instead: token probabilities, activations and layer-by-layer resolution. That makes the result reproducible and comparable between runs, which a chat answer is not.

Which model do you measure against?

An open weights model served on our own endpoint, so the measurement can be reproduced and the layers can be read. The audit names the exact model it ran against at the top of every report. Nothing here is a measurement of a closed commercial assistant, and the report never claims to be.

What does it mean if my brand reads as deviated?

It means the model ranks another category above your real one when asked what your brand is. In the 1,000-brand study behind this tool, 194 brands in 1,000 were deviated. Answers about a deviated brand start from the wrong category, so a buyer question gets framed against the wrong set of alternatives before the model gets to your product.

Why do some probes come back unavailable?

Each probe needs a specific capability from the serving endpoint: a generation route for recall, scored continuations for category binding, the stored sentiment direction vectors for the sentiment read, and a logit lens route for the depth read. When one of those is missing the audit reports which probe could not run and why, and draws nothing in its place. It never fills the gap with an estimate.

What is the channel phase?

Phase two of the audit. It fans out across you and every competitor you name, measuring third-party surfaces, the backlink footprint and the facts a crawl can establish about your own pages. Each metric is reported as an ordinal placing among the brands you named, because at four to eight peers a percentile would be a false precision.

Can I compare myself against named competitors?

Yes, in the channel phase. You can name up to seven competitor domains, and the audit measures each one on the same channels you are measured on. The model phase is measured for your brand only, since a model measurement of a peer would need its own consent and its own run.

Is the report shareable?

Every audit is stored under its own key and has a permanent link. Reopening the link renders the stored audit, including whichever part of the channel phase has finished, without running the model again.

How long does the audit take?

The model phase usually finishes in one to three minutes, longer when the serving endpoint has scaled to zero and has to start cold. The channel phase is queued behind it and keeps running after you close the tab, which is what the stored link is for.

Is the tool free?

Yes. Enter your domain and your work email and the audit runs. The email is what turns an anonymous request into a report we can send you a link to, and it is the same gate the other free tools on this site use.