Index of methodology
Methods
The standing procedures. Every run cites the clause it ran under; clause ids are stable and never reused.
A method clause is a procedure, not a result. A clause marked drafted means the rig exists and has been checked but no run using it has published — we would rather show the empty slot than imply a measurement we have not made.
Clauses M-01, M-02 and M-03 are written out at length, with a worked example and the arithmetic, in how we weigh a reference plate.
| Clause | Procedure | Status | Owner |
|---|---|---|---|
| M-01 | Weighed-plate reference — food energy | In use. Rig RP-P1 built and calibrated 2026-05-30. | Teodora Vrabel |
| M-02 | Logging protocol — one plate, every app, one photograph | In use. | Teodora Vrabel |
| M-03 | What we report, and what counts as agreement | In use. | Nandini Rege |
| M-04 | Sleep staging and duration — against polysomnography | Drafted. Rig built and checked 2026-04-21; no run published yet. | Callum Reith |
| M-05 | Step counting — against a hand tally | In use. Rig RP-S2 built and calibrated 2026-06-11; first results published 2026-08-10 as RP-RUN-2026-02, read out in RP-REP-2026-08. | Callum Reith |
| M-06 | Body-mass scales — against certified masses | Drafted. No run published yet. | Callum Reith |
| M-07 | Heart rate — against a 3-lead ECG harness | Drafted. Harness checked 2026-04-21; no run published yet. | Callum Reith |
| M-08 | Price checking | In use. First sweep 2026-08-10. | Teodora Vrabel |
| M-09 | Numbers we did not measure ourselves | In use. | Nandini Rege |
M-01 · owner Teodora Vrabel
Weighed-plate reference — food energy
Every component of a test meal is weighed separately, raw and again as served, on a balance reading to 0.01 g. Energy is computed from the weighed masses against published composition data for the specific item, not against a generic category entry. Cooking losses are measured, not assumed: the pan is weighed before and after.
The balance is checked against a certified 200 g mass at the start and end of every session. If the two checks disagree by more than 0.05 g the session is discarded, not corrected.
Restaurant and assembled plates are handled differently and worse, and we say so: a plate somebody else built can only be weighed as served and disassembled by hand, which introduces an error we cannot fully bound. Those plates are counted and reported separately from plates we cooked.
M-02 · owner Teodora Vrabel
Logging protocol — one plate, every app, one photograph
One photograph per plate, taken once, from a fixed height and angle, under the same light. That single photograph is what every app under test receives. Re-shooting until an app agrees is not testing; it is coaching.
Apps are logged in a rotating order so no product consistently gets the first or last pass. Where an app offers a correction step after its estimate, we record both the first estimate and the corrected one, and we report which is which. Manual and barcode entry are tested as separate paths from photo estimation and never blended into one figure.
Accounts are ordinary paid accounts opened at retail. No app is told it is being tested.
M-03 · owner Nandini Rege
What we report, and what counts as agreement
For each product we publish mean absolute percentage error against the reference, the signed bias, the sample size, and the interval around the estimate. Absolute error without bias hides a product that is consistently high; bias without absolute error hides a product that is wildly wrong in both directions and averages out. Both, or neither.
Two products are called different only when their intervals do not overlap. Anything closer than that is reported as a tie, in the table and in the prose, no matter how convenient the alternative would be.
Sample size is decided before the run starts and is not extended because a result is close. Every figure is published with its n attached, and a figure that appears anywhere on this site without one is an error.
M-04 · owner Callum Reith
Sleep staging and duration — against polysomnography
Consumer sleep trackers are compared against a Type II ambulatory polysomnograph worn on the same night by the same person. Epoch-by-epoch agreement is scored against the scored PSG record for total sleep time, sleep onset, wake after sleep onset, and stage classification.
Nights are scored by a person, blind to which device produced which trace. Devices are rotated between wrists across nights, because wrist dominance moves the numbers and a single-wrist protocol quietly bakes that in.
We do not report stage-level accuracy for any device from fewer than the pre-declared number of nights, and we will publish the count before the result.
M-05 · owner Callum Reith
Step counting — against a hand tally
A fixed treadmill course at declared speeds, plus a fixed outdoor course, counted by hand with a mechanical tally by an observer who is not wearing any of the devices. The tally is the reference. Two observers count the first and last session of every run and their counts must agree within 0.5% or the session is repeated.
Devices are worn simultaneously, positions rotated between sessions. We separately report the low-cadence case — walking slowly while carrying something — because that is where step counters diverge most and where a single brisk-walk figure flatters everyone.
M-06 · owner Callum Reith
Body-mass scales — against certified masses
Mass accuracy and linearity are checked against certified calibration masses across the working range, plus repeatability from ten consecutive placements at a fixed load.
Body-composition estimates from bioimpedance are reported as repeatability only — how much the same body on the same scale varies within a session and across a day. We do not own a reference method for body composition, we are not going to pretend a second consumer device counts as one, and so we do not publish body-fat accuracy figures at all.
M-07 · owner Callum Reith
Heart rate — against a 3-lead ECG harness
Optical heart-rate sensors are compared beat-by-beat against a 3-lead ECG reference recorded simultaneously. Results are reported separately for rest, steady-state work and interval work, because the interesting failure is at the transitions and an all-day mean absolute error hides it completely.
Skin tone, tattoo coverage and wrist position all affect optical sensors. Where our participant sample cannot cover that range, we say so rather than generalising from it.
M-08 · owner Teodora Vrabel
Price checking
Prices are read at the point of purchase — the checkout or in-app purchase sheet the buyer actually sees — and not from a marketing page. Where the two disagree, the price charged is the price we publish, and the discrepancy gets an entry in the corrections log so a reader who checks the vendor's own site is not left thinking we made it up.
Every price on this site carries the date it was checked. A price without a date is an error. We publish the storefront and the tier alongside the figure, because a single number with neither is unciteable.
We publish introductory and renewal pricing only where the seller states it. Where a seller does not state what a price renews at, we say nothing about renewal. Inferring a renewal figure from a discount is the kind of invention that is impossible to walk back, and on a page about prices it would be the worst error available to us.
M-09 · owner Nandini Rege
Numbers we did not measure ourselves
Any figure this desk did not produce is labelled with whoever produced it, in the same
row as the number. There is a required source field on every published figure
for exactly this reason: a number with no source cannot be entered into the system at all.
A manufacturer's own accuracy figure is a vendor claim. It may well be correct, and we will report it as what it is — a claim by the party with an interest in it — never as independent validation, and never averaged together with an independent figure into a single number that is neither.
A figure gains real weight when a second, unaffiliated lab reproduces it on its own samples. Cross-lab replication is the standard we apply to other people's results, and by our own log entry RP-COR-2026-01 it is a standard this desk does not yet meet for its own.