Orty, measured.
How Orty stacks up — feature by feature, and how its engine is measured internally. The rubric is a Client Opportunity Report: commercial signals, need, fit, a confidence level, and the evidence behind every conclusion. No marketing fog: every method is published, and the harness ships with the app.
Orty vs the alternatives
Seven capability dimensions, five tools, one 0–10 scale. Every bar is labeled.
Client qualification rubric
Email rules enforcement
Built-in CRM pipeline
Setup time for a freelance editor
Commercial / opportunity signals
Discovery workflow (no scraping)
Price fit for solo freelancers
Feature-based scoring on a 0–10 scale, compiled from public pricing and feature docs as of Sep 2026, and updated for the v2.1.3 rubric (opportunity scoring, discover-without-scraping workflow, and freelancer types beyond video editors). "Price fit" favors tools sized for solo freelancers. ChatGPT wins general flexibility — it's a chat tool, not a pipeline.
How the AI engine is scored
Orty's rating engine runs against an internal harness: 10 curated mock leads — strong fits, maybe-fits, and obvious skips — each rated live by the engine. Every output is then checked against 9 pass/fail dimensions:
The overall score is the average across dimensions — a failed dimension drags the run down, so there is no cherry-picking. The harness ships at tests/benchmark_rating.py in the app repo, and the results JSON is published after each run — anyone with an API key can reproduce the numbers.
Latest run
Run scheduled
Numbers publish here right after the next engine run (every dimension is pass/fail, so results are reproducible by anyone with an API key).
Convinced by the numbers?
Free tier keeps 3 ratings/day + 3 Co-Pilot messages/day. Pro is $19.99/mo via Whop.