BaselineUXToby Biddle

Baseline UX · Founded by Toby Biddle, founder of Loop11 · Melbourne

Stop arguing about the experience. Measure it.

Baseline UX runs benchmark studies that tell you exactly how your digital experience performs — against your competitors, and against itself last quarter. Real users, real tasks, hard numbers your executive team can act on.

The problem

Everyone has an opinion about the experience. Almost nobody has a number.

Research teams got cut. AI made it trivial to produce something that sounds like insight. So the one thing that has become genuinely scarce — evidence from real users doing real tasks — is also the one thing that still moves a roadmap or a board.

Benchmarking closes that gap. It gives you a defensible score for your experience, the same score for every competitor you care about, and the same score again next quarter so you can prove whether the work you funded actually did anything.

I have been building this measurement machinery for twenty-five years — first as co-founder of one of Australia's largest UX research consultancies, then as the founder of the platform thousands of teams use to run these studies themselves.

01You can't prioritise what you can't rank.Six competitors, one league table, and suddenly the roadmap argument is a five-minute conversation.
02Nobody funds a redesign twice without proof.A re-run benchmark shows the delta. That is what protects next year's budget.
03Executives don't read insight decks.They read numbers with a trendline. Give them one they can track.
04Qual tells you why. Only quant tells you how much.Every benchmark carries both — the score, and the verbatim that explains it.

The deliverable

This is what lands on your desk.

Every study produces a league table: your brand and your competitors, ranked, on the same tasks, with the same measures. No adjectives.

Illustrative example · brands anonymised · not client data

Example brand league table from a competitive UX benchmark study
Rank Brand SUS Grade NPS Avg task success Key observation
1Brand A 84Good +31 91% Only brand to keep success above 85% on every task
2Brand B 78OK +14 83% Strong discovery, loses users at account creation
3Brand C 73OK +2 79% High success, low confidence — users finish unsure
4Brand D 66Poor −9 71% Pricing task drives the entire deficit
5Brand E 61Poor −17 64% Severe disorientation on comparison (lostness 0.81)
6Brand F 48Fail −34 52% Half of participants abandoned before task two
90+ Excellent 80–89 Good 70–79 OK 50–69 Poor <50 Fail

Behind the table sits the detail: task-by-task breakdowns, the funnel stage where each brand loses people, coded verbatims mapped to the moment they occurred, and five to seven prioritised recommendations ranked by failure frequency and severity.

Engagements

Four ways to work together.

Most clients start by measuring themselves, then add competitors, then keep it running. Pricing in AUD, excluding GST and participant incentives.

A benchmark run once tells you where you stand. Run every quarter, it tells you whether anything you did about it worked — which is the only version that changes decisions.

Starting point

Deep Dive

Your brand, from a published benchmark · 1 week

If your brand appears in one of the industry benchmarks I publish, this is everything the public report doesn't say about your result: every task, every participant verbatim, and the exact point in the funnel where you lose people relative to the category.

  • Task-level detail
  • All verbatims
  • Competitor gap analysis
  • Priority list
From $4,500Only for brands in a published study

Start here

Baseline

Your product on its own · 3 weeks

The five tasks that actually matter, put in front of thirty to forty representative participants. You get a scored result, a ranked list of what breaks — how often, and what it costs you — and a one-page executive readout: one number, three decisions.

  • Five tasks
  • 30–40 participants
  • Unmoderated, recorded
  • Ranked failures
  • Executive readout
$9,000Fixed scope, fixed fee

Core offer

Benchmark

Competitive benchmark study · 5–6 weeks

Your product plus three to five competitors, put in front of real users on four to six realistic tasks. You get the league table, the full metric set, the funnel analysis, coded qualitative themes, and a prioritised list of what to fix first.

  • Task success
  • SUS
  • NPS
  • Lostness
  • Ease & confidence
  • Time on task
  • Coded verbatims
From $18,000Fixed scope, fixed fee

Keep it current

Tracker

Re-run every quarter · ongoing

The same study, same tasks, same measures, every quarter. The instrument already exists, so each round costs less than the first. UX performance becomes a metric with a trendline rather than a project with an end date — and most clients re-run after a significant release, which is how you find out whether the release worked.

  • Quarterly waves
  • Trend reporting
  • Executive one-pager
  • Competitor watch
From $6,000 a wave$6,000 to re-run a Baseline, $12,000 for a Benchmark. Billed per wave or monthly.

Leadership

Fractional Research Leadership

One to two days a week

For teams that need the function built, not just a study delivered. I set the research strategy, stand up the operations, run the measurement program, manage vendors and panels, and get your product managers and designers doing credible work themselves.

  • Research strategy
  • ResearchOps
  • Panel & vendor management
  • Team coaching
  • Hiring support
From $8,000/moMinimum three months

Entry point

Advisory

By the day

An expert review of an experience before you commit to testing it, a second opinion on a research plan, a critique of a vendor's proposal, or a working session to get a stalled program moving again.

  • Expert review
  • Research plan critique
  • Vendor selection
  • Workshops
$3,500/dayRemote or on-site in Melbourne

Method

How a benchmark runs.

The analysis rules are written down and agreed before any data is collected. That is what stops a benchmark becoming an exercise in explaining away whatever the numbers turned out to be.

Below is a competitive Benchmark across five or six brands. A Baseline on a single product runs the same way in three weeks.

STEP 01

Scope

We agree the competitor set, the tasks that represent your real customer journey, and the measures that matter for your category.

Week 1

STEP 02

Lock the rules

A written output mapping fixes the thresholds, the ranking logic and the claims the report is not permitted to make — before anyone sees a result.

Week 1–2

STEP 03

Field

Unmoderated, task-based studies with recruited participants, run on Loop11 — the platform I founded and still run — so scale costs almost nothing.

Week 2–4

STEP 04

Report

League table, task-level analysis, coded qualitative themes, and prioritised recommendations. Presented to your team and to your executives.

Week 4–6

Published research

Industry benchmark reports

Each quarter I benchmark an entire category — five or six of the best-known brands in it, tested the same way on the same tasks — and publish the industry findings openly.

If your brand appears in one, you can commission the deep-dive on your own result: every task, every verbatim, every place a competitor beat you and by how much.

Schedule

SuperannuationIn field
Health insuranceQ4
Retail bankingQ1
UniversitiesQ2
Your category?Open

About

Baseline UX is Toby Biddle.

You are not buying a team of strangers. You get me, and twenty-five years of doing this on both sides of the problem.

In 2001 I co-founded U1 Group and spent sixteen years growing it into one of Australia's largest UX and design research consultancies, running a team of around thirty and delivering research programs for enterprise and government clients.

In 2009 I founded Loop11 because the designers we worked with kept asking for a way to run their own studies. It became one of the first online usability testing platforms, and it is now used by teams in over 100 countries. I still run it.

That combination is unusual and it is the reason this offer exists. I have run the research function at scale, so I know what an executive team will actually act on. And I own the measurement infrastructure, so studies that would be expensive for an agency to field are routine for me.

I work with product, design and digital leaders who are past the point of wanting more opinions — in Australia and New Zealand primarily, and globally where the timezone allows.

2009 — PRESENT

Founder & CEO, Loop11

Remote usability testing platform used by teams in over 100 countries.

2001 — 2017

Co-founder, U1 Group

Grew to ~30 people; one of Australia's largest UX & CX research consultancies.

ONGOING

Writing & speaking

Published in UX Magazine, UXmatters and Boxes & Arrows.

BASED IN

Melbourne, Australia

AEST. Working across AU/NZ, Asia-Pacific, Europe and North America.

What's your number?

If you can't answer that about your own experience, that is the conversation. Thirty minutes, no deck — tell me what decision you're stuck on and I'll tell you whether a benchmark would settle it.