Open the app

AI nutritionist and nutrition chatbots: what can you trust

A phone in a hand over a kitchen table with a plate of salad, the phone screen is dark, a glass of water is nearby

Asking a chatbot about your diet has become easier than making an appointment with a specialist, and the answer comes detailed, polite, and structured. Tests of such answers exist, and the picture in them is twofold: for general questions, the quality is quite good, and the collected errors are grouped into recognizable topics where making a mistake is most costly.

What exactly was tested

Studies of this kind are structured similarly: a set of questions, model responses, and expert evaluations across several axes—scientific accuracy, completeness, clarity, quality of reasoning, and sometimes reproducibility upon repeated queries.

Compared to humans, chatbots performed comparably on nutrition questions: for some questions, the scores for scientific accuracy and clarity were higher than those of specialists.

In more specialized areas, the picture is worse. A review of sports nutrition questions evaluated accuracy, completeness, clarity, quality of reasoning, and reproducibility and described the accuracy as moderate and inconsistent between models.

A separate study on vegetarian and vegan diets found that all tested models provided responses capable of causing harm, and—more importantly—grouped them by topic.

Where errors repeat

A list of recurring topics is more useful than a general accuracy score because it indicates exactly where to double-check.

TopicNature of error
Vitamin B12 in plant-based dietsUnreliable sources cited as sufficient
Iodine and seleniumAdvice given without considering the upper intake limit
Pregnancy, children, elderlyGeneral advice provided to a vulnerable group without caveats
Chronic diseasesRecommendation given without considering condition and medications
SymptomsDietary advice provided instead of a referral to a doctor
ReasoningConfident tone with an untraceable source

What the first five rows have in common is that they are situations where an upper limit or contraindication is more important than the recommendation itself. What the last one has in common is that, based on the tone of the response, it is impossible to distinguish a substantiated claim from a plausible one.

Why confidence is not linked to correctness

A model generates the most probable continuation, not a verified fact, and there is no "I am not sure" form in this scheme. A generally accepted recommendation and a claim that no one can verify sound equally smooth. This mechanism and ways to bypass it are analyzed in detail in the article on AI errors.

Added to this is reproducibility: the same question asked twice can receive different answers. This is useful to use as a check—a discrepancy between two answers indicates that there is no solid foundation for the question.

Questions that a chatbot answers well

Questions where the answer must be double-checked

How to ask a question so the answer is more useful

  1. Provide context. Age, weight, activity level, restrictions. A general question gets a general answer.
  2. Ask for the mechanism, not the verdict. "Why is it so" is verifiable; "what should I eat" is not.
  3. Ask about boundaries. "How much is too much" and "who should avoid this" are the very places where answers are most often incomplete.
  4. Ask the question twice. A discrepancy in answers is a signal to double-check.
  5. Separate calculation from advice. You can accept the arithmetic, but you should verify health recommendations.

And the boundary that does not move: a chatbot does not diagnose or prescribe treatments and diets. For illnesses, pregnancy, medication use, and any conditions requiring dietary control, decisions should be made with a doctor.

Count it from a photo

Frequently asked questions

Can you trust nutrition advice from a chatbot?
For general questions, the quality of responses in tests proved comparable to those of specialists, and in some areas, they were even more accurate and clear. However, in specialized fields, accuracy is considered moderate and inconsistent, and all tested models produced responses that could potentially be harmful.
In which topics do errors occur most frequently?
Sources of vitamin B12 in plant-based diets, iodine and selenium dosages without considering upper limits, advice for vulnerable groups without caveats, recommendations for chronic illnesses, and providing nutritional advice instead of referring to a doctor when symptoms are present.
Why does the answer sound confident even when it is incorrect?
The model generates the most probable continuation rather than a verified fact, and it has no way to express uncertainty. It is impossible to distinguish a generally accepted recommendation from a plausible-sounding statement based on tone alone.
How can I verify an answer if I am not a specialist?
Ask the same question a second time and compare: discrepancies indicate a lack of a solid foundation. Additionally, verify links by clicking on them—the model may provide a plausible-looking link to a non-existent study.
What is a chatbot best suited for?
For explaining mechanisms, decoding food labels, calculating ingredient quantities, and finding recipe substitutions. It is also useful for formulating questions to ask your doctor.

Read next

This article is for general information. It is not medical advice, a diagnosis, or a prescription for treatment or a diet, and it does not replace a consultation with your doctor. If you have a health condition, are pregnant, take medication, or follow a diet prescribed to you, decisions about food belong with your doctor.

Figures from regulations, guidelines and studies are given as they stood when this article was prepared and may since have changed; check them against the primary sources. This article is not advertising, an offer, or individual advice, and neither the author nor the site owner is responsible for decisions taken on the basis of it.