Skip to content
Reference Plate

Index of methodology

Methods

The standing procedures. Every run cites the clause it ran under; clause ids are stable and never reused.

A method clause is a procedure, not a result. A clause marked drafted means the rig exists and has been checked but no run using it has published — we would rather show the empty slot than imply a measurement we have not made.

Clauses M-01, M-02 and M-03 are written out at length, with a worked example and the arithmetic, in how we weigh a reference plate.

Clause index
ClauseProcedureStatusOwner
M-01Weighed-plate reference — food energyIn use. Rig RP-P1 built and calibrated 2026-05-30.Teodora Vrabel
M-02Logging protocol — one plate, every app, one photographIn use.Teodora Vrabel
M-03What we report, and what counts as agreementIn use.Nandini Rege
M-04Sleep staging and duration — against polysomnographyDrafted. Rig built and checked 2026-04-21; no run published yet.Callum Reith
M-05Step counting — against a hand tallyIn use. Rig RP-S2 built and calibrated 2026-06-11; first results published 2026-08-10 as RP-RUN-2026-02, read out in RP-REP-2026-08.Callum Reith
M-06Body-mass scales — against certified massesDrafted. No run published yet.Callum Reith
M-07Heart rate — against a 3-lead ECG harnessDrafted. Harness checked 2026-04-21; no run published yet.Callum Reith
M-08Price checkingIn use. First sweep 2026-08-10.Teodora Vrabel
M-09Numbers we did not measure ourselvesIn use.Nandini Rege

M-01 · owner Teodora Vrabel

Weighed-plate reference — food energy

In use. Rig RP-P1 built and calibrated 2026-05-30.

Every component of a test meal is weighed separately, raw and again as served, on a balance reading to 0.01 g. Energy is computed from the weighed masses against published composition data for the specific item, not against a generic category entry. Cooking losses are measured, not assumed: the pan is weighed before and after.

The balance is checked against a certified 200 g mass at the start and end of every session. If the two checks disagree by more than 0.05 g the session is discarded, not corrected.

Restaurant and assembled plates are handled differently and worse, and we say so: a plate somebody else built can only be weighed as served and disassembled by hand, which introduces an error we cannot fully bound. Those plates are counted and reported separately from plates we cooked.

M-02 · owner Teodora Vrabel

Logging protocol — one plate, every app, one photograph

In use.

One photograph per plate, taken once, from a fixed height and angle, under the same light. That single photograph is what every app under test receives. Re-shooting until an app agrees is not testing; it is coaching.

Apps are logged in a rotating order so no product consistently gets the first or last pass. Where an app offers a correction step after its estimate, we record both the first estimate and the corrected one, and we report which is which. Manual and barcode entry are tested as separate paths from photo estimation and never blended into one figure.

Accounts are ordinary paid accounts opened at retail. No app is told it is being tested.

M-03 · owner Nandini Rege

What we report, and what counts as agreement

In use.

For each product we publish mean absolute percentage error against the reference, the signed bias, the sample size, and the interval around the estimate. Absolute error without bias hides a product that is consistently high; bias without absolute error hides a product that is wildly wrong in both directions and averages out. Both, or neither.

Two products are called different only when their intervals do not overlap. Anything closer than that is reported as a tie, in the table and in the prose, no matter how convenient the alternative would be.

Sample size is decided before the run starts and is not extended because a result is close. Every figure is published with its n attached, and a figure that appears anywhere on this site without one is an error.

M-04 · owner Callum Reith

Sleep staging and duration — against polysomnography

Drafted. Rig built and checked 2026-04-21; no run published yet.

Consumer sleep trackers are compared against a Type II ambulatory polysomnograph worn on the same night by the same person. Epoch-by-epoch agreement is scored against the scored PSG record for total sleep time, sleep onset, wake after sleep onset, and stage classification.

Nights are scored by a person, blind to which device produced which trace. Devices are rotated between wrists across nights, because wrist dominance moves the numbers and a single-wrist protocol quietly bakes that in.

We do not report stage-level accuracy for any device from fewer than the pre-declared number of nights, and we will publish the count before the result.

M-05 · owner Callum Reith

Step counting — against a hand tally

In use. Rig RP-S2 built and calibrated 2026-06-11; first results published 2026-08-10 as RP-RUN-2026-02, read out in RP-REP-2026-08.

A fixed treadmill course at declared speeds, plus a fixed outdoor course, counted by hand with a mechanical tally by an observer who is not wearing any of the devices. The tally is the reference. Two observers count the first and last session of every run and their counts must agree within 0.5% or the session is repeated.

Devices are worn simultaneously, positions rotated between sessions. We separately report the low-cadence case — walking slowly while carrying something — because that is where step counters diverge most and where a single brisk-walk figure flatters everyone.

M-06 · owner Callum Reith

Body-mass scales — against certified masses

Drafted. No run published yet.

Mass accuracy and linearity are checked against certified calibration masses across the working range, plus repeatability from ten consecutive placements at a fixed load.

Body-composition estimates from bioimpedance are reported as repeatability only — how much the same body on the same scale varies within a session and across a day. We do not own a reference method for body composition, we are not going to pretend a second consumer device counts as one, and so we do not publish body-fat accuracy figures at all.

M-07 · owner Callum Reith

Heart rate — against a 3-lead ECG harness

Drafted. Harness checked 2026-04-21; no run published yet.

Optical heart-rate sensors are compared beat-by-beat against a 3-lead ECG reference recorded simultaneously. Results are reported separately for rest, steady-state work and interval work, because the interesting failure is at the transitions and an all-day mean absolute error hides it completely.

Skin tone, tattoo coverage and wrist position all affect optical sensors. Where our participant sample cannot cover that range, we say so rather than generalising from it.

M-08 · owner Teodora Vrabel

Price checking

In use. First sweep 2026-08-10.

Prices are read at the point of purchase — the checkout or in-app purchase sheet the buyer actually sees — and not from a marketing page. Where the two disagree, the price charged is the price we publish, and the discrepancy gets an entry in the corrections log so a reader who checks the vendor's own site is not left thinking we made it up.

Every price on this site carries the date it was checked. A price without a date is an error. We publish the storefront and the tier alongside the figure, because a single number with neither is unciteable.

We publish introductory and renewal pricing only where the seller states it. Where a seller does not state what a price renews at, we say nothing about renewal. Inferring a renewal figure from a discount is the kind of invention that is impossible to walk back, and on a page about prices it would be the worst error available to us.

M-09 · owner Nandini Rege

Numbers we did not measure ourselves

In use.

Any figure this desk did not produce is labelled with whoever produced it, in the same row as the number. There is a required source field on every published figure for exactly this reason: a number with no source cannot be entered into the system at all.

A manufacturer's own accuracy figure is a vendor claim. It may well be correct, and we will report it as what it is — a claim by the party with an interest in it — never as independent validation, and never averaged together with an independent figure into a single number that is neither.

A figure gains real weight when a second, unaffiliated lab reproduces it on its own samples. Cross-lab replication is the standard we apply to other people's results, and by our own log entry RP-COR-2026-01 it is a standard this desk does not yet meet for its own.

Page last reviewed . Standing pages are re-read whenever a method changes or an entry is added to the corrections log, and the date above moves in the same commit.