How a ShipScore is computed

Every verdict on this site comes out of the same pipeline, on the same harsh scale, with the evidence printed underneath. This page is the full provenance — models, data, calibration, and where we've been wrong.

Part I — The pipeline
§1 — Parse (Haiku)

Your idea or URL is parsed into a product description, a category, and up to four competitor-search keywords. Cheap, mechanical, fast.

§2 — Evidence (the corpus)

The keywords are matched against the App Radar corpus — a living database our collectors refresh daily: 1,400+ tracked products with verified MRR time series, AppSumo trajectories, funding events, marketplace listings (what similar businesses actually sell for), and Reddit pain threads. Word-level matching with a ≥2-hit gate so "AI tracker" doesn't match every tracker.

§3 — Verdict (Opus)

The evidence pack is scored on MJ DeMarco's CENTS framework (Control, Entry, Need, Time, Scale) by Claude Opus 5 — the same referee for every idea, including the ones we generate ourselves. The rubric is harsh by design: no market evidence caps Need at 5; a well-loved incumbent caps Entry low; the weakest commandment is always named. Inspired by MJ DeMarco's CENTS framework.

§4 — Stamp & re-spin (Sonnet)

The score becomes a stamp — FULL AHEAD, DEAD SLOW, or STOP — with kill switches, the pace test, and the cheapest next step. Re-spins (business-model variations engineered to beat the original) are generated by Sonnet, chosen in calibration tests specifically because it doesn't inflate its own scores.

Part II — The two instruments
§1 — CENTS: the score

Every idea is scored on five commandments, 0–10 each, by the same referee with the same evidence standard. The composite is the plain mean — no weighting games — and the weakest commandment is always named in the verdict, because a business fails at its weakest link, not its average.

C — ControlOwnership of the customer relationship. Gatekeepers (app stores, search algorithms, one marketplace, one ad platform) that can tax or kill you score this down. Direct relationships score it up.
E — EntryBarriers to imitation. High is good once you're in. A well-loved incumbent (4.5★+ with loyal users) is a barrier working against you, and caps this low — ratings alone can only soften that cap when the incumbent's revenue is provably collapsing (see §2).
N — NeedDemand proven with money: verified rival revenue, funding flowing into the space, exit multiples, complaint volume. Ideas with no market evidence cap at 5 — "people should want this" is not demand.
T — TimeDetachment of income from hours. Software that sells while you sleep scores high; services, marketplaces needing constant operation, and content treadmills score low.
S — ScaleReach without linear cost. Marginal cost per customer near zero, a reachable market big enough to matter, no geographic or regulatory ceiling that stops growth at a wall.

Inspired by MJ DeMarco's CENTS framework (attribution only — the rubric, caps and wording here are ours).

§2 — Moat trend: the trajectory

A rating tells you how loved a rival is. It doesn't tell you whether that love is durable. For every product in the corpus with a verified revenue series, we compute a moat trend — the direction of its economic moat, from its own numbers:

↑ wideningRevenue compounding ≥ +10%/mo (monthlyized). The incumbent is getting stronger — attacking it gets harder every month you wait.
→ stableHolding between −5% and +10%/mo. A fortress, but a static one.
↓ narrowingDeclining −5% to −10%/mo. Users still love it; the money is quietly leaving. Possible opening.
↓↓ erodingCollapsing faster than −10%/mo. A well-loved product whose economics are dying — the classic disruption seam.

The rules that keep us honest:

Where you'll see it: trend badges on rival evidence in every report. Where it works behind the glass: our daily-idea engine closes lanes led by widening incumbents and discounts lanes whose leader is eroding.

Part III — What the evidence is (and isn't)
Provenance grades
Verified MRRIndependently tracked revenue (TrustMRR-class sources) — the strongest signal we have. Not self-reported.
Ratings / upvotes / reviewsPlatform-reported (App Store, Product Hunt, AppSumo). Useful, gameable — weighted accordingly.
Funding eventsPublic announcements cross-checked across feeds; company-level dedup.
Marketplace multiplesAsking prices and sold outcomes from Acquire.com — what buyers actually pay for profit.
Reddit painComplaint threads as demand signal, not proof. Volume over sentiment.
Calibration

Scores anchor to a calibration ladder of real products (the "6.6 attractor" = real revenue, crowded market). New models and prompt changes run a blind-recheck protocol — flag at Δ ≤ −1.5 — before they touch production. A 7-model public bake-off (2026-08-01) decided the current architecture; the results are published in the repo.

Part IV — Where we've been wrong
§1 — The chimera apps

A collector once mis-paired sponsor names to favicons, manufacturing products that didn't exist with phantom revenue. Caught when a name didn't match its logo. Fix: sponsor filter + duplicate-favicon guard + an audit comparing raw source data to stored names. The wrong rows were deleted publicly.

§2 — The flaky oracle

Apple's review feed returns empty results at random. An early version cached the empty result and concluded "no complaints exist" — a demand-killing false negative. Fix: empty is treated as an error, never an answer; dry apps retry on the next cycle.

§3 — The generous generator

Models asked to invent opportunities score them ~1 point higher than adversarial verification allows. Every generated candidate — including our own daily ideas — now goes through the unchanged harsh referee. When nothing clears 7.0, we publish "closest, and here's what kills it." The misses stay in the archive.

Part V — The pledges
Privacy
Money-back
Honesty