How accurate is calorie counting from a photo?

6 min readAuthor: Doğan Varkan

Estimating calories from a photo has become genuinely useful in the last few years, but using it without knowing what it measures is misleading. What the model sees is not the food — it is an image of the food.

What a photo actually carries

An image tells you two things: what is on the plate and roughly how much space it takes up. Recognising the food — rice or bulgur, chicken breast or thigh — is comparatively easy work for today’s vision models. The hard part is the second step: turning volume into mass, and mass into calories.

The gap between recognising and measuring sits exactly there. On academic datasets of around a hundred food classes, recognition accuracy has been above 90% for years; portion estimation has no comparable number, because one photograph does not carry the scale of the scene. A plate twice the size, shot from twice the distance, fills the frame identically. The camera’s geometry cannot separate those two cases.

Two further unknowns follow. Density is the first: leafy salad comes in around 0.1–0.2 grams per cubic centimetre, cooked rice approaches 0.8, oil sits at 0.92. The second is what you cannot see — the oil that went into the pan, the sugar in the sauce, the butter stirred through the rice. Fat carries 9 kcal per gram; a tablespoon of olive oil adds roughly 120 kcal while leaving no trace in the photo.

The spread in energy density is why this matters: lettuce is 15 kcal per 100 grams, olive oil 884. Sixty times apart, on the same plate, in one frame.

What the model is forced to guess

When you ask a vision model how many calories are on a plate, this chain runs underneath: recognise the food → estimate the portion → recall typical values per 100 grams → multiply. The weakest link is the middle step, and the error leaves it magnified: 30% over on the portion is 30% over on the calories.

Errors here multiply rather than add. Ten per cent off in recognition, 25% off in portion and 10% off in the reference table do not cancel; on a bad day they all point the same way.

That is why two photos of the same meal produce two different numbers. The model is not being indecisive; the second photo genuinely holds less information — the angle changed, the rim of the plate fell outside the frame, the light swallowed the shadows.

The reference table brings uncertainty of its own. “100 grams of rice” does not say cooked or raw, and the difference is not small: rice absorbs water and roughly triples in mass, pasta gets close to two and a half times, grilled meat sheds 20–25% of its weight as water. One name, three different numbers.

Cooking method is usually unreadable too. A baked potato and a fried one share the same shade of gold, yet the second is heavier by whatever oil it absorbed — 100 to 150 kcal per 100 grams. Breaded chicken behaves the same way: the coating is a thin layer, the calorie difference is not. The image will not tell you.

The benchmark is a scale, not a person

When photo estimation looks bad, one question tends to get skipped: how good is the alternative? Trained dietitians estimating portions from images land in the 20–25% error band. Ordinary people keeping written logs have been shown, in doubly labelled water studies, to under-report their intake by 20–30%.

Photos have one advantage typing does not: they cut omissions. What gets forgotten in a written log — the bread at the edge of the plate, the sugar in the tea, the cooking oil — is still sitting in the frame. An item weighed badly beats an item never recorded.

Four things that shrink the error

Put a scale in the frame. A photo with a fork, a spoon or a standard glass beside the plate makes volume estimation markedly easier. Without a reference object the model has to assume the plate’s size, and dinner plates run from 22 to 30 centimetres — that range alone is close to a two-fold volume difference between two plates that look alike.

Shoot at an angle, not straight down. A top-down photo shows surface area but not height. Thirty to forty-five degrees carries both in one frame.

Break a mixed plate apart. For a stew or a composite salad whose contents are not visible, naming the components separately beats asking for one number. Giving the model what it cannot see moves the estimate towards a measurement.

Call out the oil and the sauce. Most invisible calories come from there, and this single correction usually matters more than the other three combined.

Drinks are their own problem. How full a glass is, and how thick its wall is, do not read reliably from a photo, and sweetened drinks carry a lot of calories per millilitre. Typing those in by hand is faster than fighting the angle.

When you actually need a scale

Photo estimation is useful in daily tracking because it is roughly right and can be done every day. A scale is accurate but not realistic at every meal. One practical way to combine them: weigh the 10–15 meals you eat most often, once each, learn those weights, then carry on with photos. A scale is a calibration tool, not a habit.

Some cases genuinely need one: oil, nuts, cheese, tahini — anything with high calories in small volume; and goals where grams decide the outcome, such as contest prep or a prescribed diet. A handful of almonds can be 25 grams or 45, and that gap is 120 kcal.

The scale has a second job: training your eye. Someone who weighs for a few weeks and checks each reading against their own guess estimates far better from a photo afterwards, and the scale can then live in a drawer — the thing that learned was you.

Consistency beats accuracy

The point of tracking is not absolute accuracy but seeing the direction. Even a method that systematically undercounts every meal by 10% is useful read alongside the weight curve: if after two weeks the scale is not moving as expected you shift the target, and the constant error is absorbed by the correction.

What breaks it is inconsistency. Someone who weighs everything to the gram one day and leaves two meals unrecorded the next cannot compare those days at all, and the weekly average stops representing either intake or change — the log ends up describing the method rather than the thing it measured. Applying one method every day beats a more precise method applied irregularly. That is the difference.

How this works in FitDex

The photo analysis in the app separates the items on the plate, estimates each one and leaves the result editable — the person who knows the grams best is the one holding the plate.

Which vision model does that work is not picked at random either. The criterion is not accuracy on a single photo but the same photo returning the same number again and again; a model that says 620 one day and 890 the next is useless for tracking even if it is right on average.

The photo calorie tool here does not upload photos; it answers a different question. If your estimates leave a systematic margin, it works out what that margin comes to in calories and kilos over weeks — and whether the four corrections above are worth applying.

aicaloriesmeasurement

Run the numbers from this guide yourself

Related guides

This guide is for information only; it is not a substitute for medical diagnosis or treatment. Talk to a health professional before making a lasting change to your diet.