patrickz.aiLet’s talk ↗

Start with AI / Guide

Compare AI tools on your own tasks

Build a small comparison sheet based on work you can judge.

About 35 minutesBeginnerRead free · No sign-up

Before you start

Access to one or two AI tools and three non-sensitive tasks with known answers. No new subscription is necessary.

Why this lesson exists

Patrick’s model-bench posts use roles and evidence instead of a permanent “best model” ranking. The same comparison can be repeated whenever tools change.

Do the exercise

  1. Choose representative tasks

    Use one factual summary, one creative edit and one structured extraction. Write expected facts and formatting rules before testing.

  2. Keep the input fair

    Give each tool the same source, instructions and constraints. Record the date and displayed model or mode when available. Do not assume all accounts expose the same features.

  3. Score blind where possible

    Hide tool names and score correctness, usefulness and editing effort. Use a known answer or trusted source for factual tasks. Record waiting time and any visible usage cost.

  4. Assign roles

    Choose a tool for each task type, not one winner for everything. Keep a difficult example as a future regression test and repeat the comparison after meaningful changes.

A prompt to adapt

Replace the bracketed parts with your own practice details.

Create a scoring rubric for these three tasks: [tasks]. Score factual accuracy, instruction following and human editing effort from 0 to 2 with concrete definitions. Keep speed and cost separate. Give a blank results table and a rule for choosing a tool by task.

What this looks like

A summary that is fast but changes the deadline should lose to a slower accurate answer. A beautiful paragraph does not compensate for incorrect extracted data.

Check your result

Use evidence from your output. A confident explanation from the AI is not enough.

  • Inputs and constraints are identical.
  • The factual task has an independent answer key.
  • Your choice names a task and a tradeoff.

If it isn’t working

If outputs are too similar, use a realistic edge case. If one tool had web access and the other did not, record that difference instead of treating it as a model-only comparison.

Where this came from

Adapted from the archived LinkedIn theme “Stop Asking Which AI Is Best. Build a Bench..” See the post coverage and editorial method.

Prepared September 2026. Tools and interfaces change; use current official setup instructions. Session lengths are estimates.

KEEP GOING

Your next useful step

Browse all 59 guides ↗