Benchmark

Orty, measured.

How Orty stacks up — feature by feature, and how its engine is measured internally. The rubric is a Client Opportunity Report: commercial signals, need, fit, a confidence level, and the evidence behind every conclusion. No marketing fog: every method is published, and the harness ships with the app.

Orty vs the alternatives

Seven capability dimensions, five tools, one 0–10 scale. Every bar is labeled.

OrtyChatGPT (free)Apollo.ioInstantly.aiSpreadsheet + Gmail

Client qualification rubric

Orty9/10
ChatGPT (free)5/10
Apollo.io6/10
Instantly.ai4/10
Spreadsheet + Gmail2/10

Email rules enforcement

Orty9/10
ChatGPT (free)4/10
Apollo.io5/10
Instantly.ai6/10
Spreadsheet + Gmail1/10

Built-in CRM pipeline

Orty8/10
ChatGPT (free)1/10
Apollo.io9/10
Instantly.ai8/10
Spreadsheet + Gmail3/10

Setup time for a freelance editor

Orty9/10
ChatGPT (free)7/10
Apollo.io3/10
Instantly.ai3/10
Spreadsheet + Gmail6/10

Commercial / opportunity signals

Orty9/10
ChatGPT (free)4/10
Apollo.io5/10
Instantly.ai5/10
Spreadsheet + Gmail2/10

Discovery workflow (no scraping)

Orty8/10
ChatGPT (free)6/10
Apollo.io7/10
Instantly.ai3/10
Spreadsheet + Gmail4/10

Price fit for solo freelancers

Orty8/10
ChatGPT (free)9/10
Apollo.io4/10
Instantly.ai4/10
Spreadsheet + Gmail9/10

Feature-based scoring on a 0–10 scale, compiled from public pricing and feature docs as of Sep 2026, and updated for the v2.1.3 rubric (opportunity scoring, discover-without-scraping workflow, and freelancer types beyond video editors). "Price fit" favors tools sized for solo freelancers. ChatGPT wins general flexibility — it's a chat tool, not a pipeline.

How the AI engine is scored

Orty's rating engine runs against an internal harness: 10 curated mock leads — strong fits, maybe-fits, and obvious skips — each rated live by the engine. Every output is then checked against 9 pass/fail dimensions:

Opportunity score valid — an overall 0–100 integer, so the run fails outright if the engine returns nothing usable.
Client Opportunity Report present — an overall 0–100 integer, a confidence level, and Commercial-signal / Need / Fit ratings (each 1–5), so a prospect is never reduced to a single unexplained number.
Evidence and why-not present — the report must list the evidence behind its conclusions and an explicit why-not; unsupported claims must read “Unknown” / “Not enough evidence”. No guessed dollar budgets — invented figures fail.
Signals explain the score — the output must list the positive and negative signals it actually observed, so every score has a why.
Band behavior — 70–100 must pitch, 40–69 must pitch but flag hesitation, 1–39 must skip. No straddling.
Subject line present — every drafted email ships with a subject line, or the dimension fails.
Email length ≤ 150 words — long-winded pitches fail. The cap is enforced on every output.
Specificity + no generic openers — the email must reference the actual lead and may not use banned filler like “I love your content.”
CTA present — every email ends with one clear, low-friction call to action.

The overall score is the average across dimensions — a failed dimension drags the run down, so there is no cherry-picking. The harness ships at tests/benchmark_rating.py in the app repo, and the results JSON is published after each run — anyone with an API key can reproduce the numbers.

Latest run

Run scheduled

Numbers publish here right after the next engine run (every dimension is pass/fail, so results are reproducible by anyone with an API key).

Convinced by the numbers?

Free tier keeps 3 ratings/day + 3 Co-Pilot messages/day. Pro is $19.99/mo via Whop.