Open the app

Testing it on yourself: telling an effect apart from noise

Trying something on yourself is the only available way to find out whether it works for you specifically. The catch is that people are poor instruments for measuring their own changes: they start observing at the worst moment, change several things at once, and remember confirmations better than misses.

Why "I tried it and it helped" means little

Not because anyone is lying. Because several effects converge in self-observation, each capable of producing an improvement with no intervention at all.

This applies well beyond supplements. It is exactly how people become convinced about eating patterns, techniques and devices.

Noise inside numbers that look precise

MeasureIts own variation
Morning weight±1–2 kg within a day from water, glycogen and gut contents
Eyeballed food intakeSystematic underestimate, often by a quarter or more
Sleep stages from a wristbandInferred from pulse and movement; total sleep time is more reliable than structure
Calories burned per a trackerDisagreement with laboratory measurement reaches tens of per cent
Self-rated wellbeingDepends on time of day and on what the rater expects to find

If the expected effect is smaller than a measure's own variation, self-observation cannot see it. That is a matter of resolution, not of effort. More on instrument error in fitness trackers and smart scales.

How a decent single-person experiment is built

The format exists in medicine, where it is called an n-of-1 trial. A methodological review of randomised trials of this type sets out what separates them from plain observation: alternating intervention and control periods, an outcome defined in advance and, where feasible, blinding.

None of that transfers home completely, but three parts of it do.

  1. One variable. Exactly one thing changes. If sleep and diet both changed, the experiment ended before it started.
  2. The outcome named first. Not "I felt better" but a specific number: average weight over a week, sets completed at a given load, average sleep rating. An outcome chosen afterwards is always found.
  3. Alternating blocks. Two weeks on, two weeks off, then repeat. A single switch-on is indistinguishable from regression to the mean.

How long to run it

For body weight, no less than three or four weeks, comparing weekly averages rather than individual mornings. Weekly averaging removes most of the water-driven variation; comparing single days lets noise swamp any realistic effect.

For subjective outcomes — sleep, alertness, appetite — the same duration applies, with the caveat that you cannot blind yourself, so the result stays open to question. That is a limit of the method, not a flaw in the execution.

For anything that changes slowly, a home experiment is simply the wrong tool: a month will show nothing.

Where almost every home experiment breaks

Not at the device and not at discipline, but at food accounting. Most of the protocols people test affect appetite one way or another, while the person continues to believe they are eating "about the same". Underestimating intake is the most stubborn error in the whole field, and on its own it can explain both the success and the failure of any experiment.

So the first step of an honest experiment is not choosing the intervention but establishing the baseline: how much is eaten when nothing has changed yet. The method of recording matters less than using the same method before and after — otherwise the comparison is between counting methods rather than between periods. What goes wrong most often is covered in counting calories in home cooking.

Knowing when to stop

A pre-set duration also gives the experiment an ending. Without one it turns into an untested habit: someone keeps doing something for months on the strength of how the first week felt.

A reasonable statement at the start sounds roughly like this: "if after four weeks the weekly average weight has not moved by more than this much, I stop." It is uncomfortable precisely because it makes failure visible. That is what it is for.

Count it from a photo

Frequently asked questions

Why can I not trust the feeling that something helped?
Because self-observation combines regression to the mean, expectation, several simultaneous changes and selective memory. Each of those can produce an apparent improvement with no intervention at all.
How much does body weight fluctuate on its own?
By one to two kilograms within a day, from water, glycogen and gut contents. That is why weekly averages are the sensible comparison rather than individual mornings.
How long should a self-experiment run?
For body weight, at least three to four weeks, with alternating periods on and off the intervention. A single switch-on is indistinguishable from a natural return to the mean.
Can several things be changed at once?
They can, but then the result cannot be attributed to any of them. If the goal is to learn what works, exactly one variable changes at a time.
What ruins these experiments most often?
Food accounting. Most tested protocols affect appetite, and eyeballed intake is systematically underestimated, so the change in calories goes unnoticed and gets credited to the protocol.

Read next

This article is for general information. It is not medical advice, a diagnosis, or a prescription for treatment or a diet, and it does not replace a consultation with your doctor. If you have a health condition, are pregnant, take medication, or follow a diet prescribed to you, decisions about food belong with your doctor.

Figures from regulations, guidelines and studies are given as they stood when this article was prepared and may since have changed; check them against the primary sources. This article is not advertising, an offer, or individual advice, and neither the author nor the site owner is responsible for decisions taken on the basis of it.