LezzworkRead the latest
Method

How we test

A measured test at Lezzwork has four parts, and all four are public.

Last reviewed 8 Sept 2026

  1. 01

    Fixed inputs

    Every test runs on a dataset we publish before the test: synthetic emails, invoices, meeting notes or records built to look like a small business's real data without containing anyone's real data. You can download it from the article and run the same test yourself.

  2. 02

    A rubric written first

    We decide what "correct" means before running anything: fields to extract, categories to assign, actions to produce. The rubric is in the article. We do not change it after seeing results.

  3. 03

    Scored outputs

    Each tool's output is scored against the rubric. We report the score, the failures, the time it took and what it cost. Where a tool has a free tier or a trial, we say which one we used.

  4. 04

    Limitations left in

    A test on twenty emails is a test on twenty emails. We say what the test cannot tell you, and we date every check because tools change.

Datasets

Every dataset is synthetic: no real people, companies or invoices. Files are small and plain (TXT, CSV, JSON, Markdown) so they open anywhere. Download, run the same test, compare.

Invoice extraction v1
Ten synthetic invoices as plain text, bundled as JSON, with an answer key and the rubric.
Measured test scheduled for the week of 15 September 2026.
Email triage v1
Twenty synthetic emails with the triage rubric.
Measured test scheduled for the week of 15 September 2026.

What we do not do

  • Call something "best" without a comparison set and a rubric.
  • Quote prices or limits we did not check near publication.
  • Present desk research as a test. Articles say "tested" only when we ran it.
  • Give medical, legal, tax or investment advice.