Tracking & tools

Comparing food trackers: the error nobody puts in the table

Dennie Ordynskyi · August 2026 · 8 min read

If you are comparing food trackers side by side, the short answer is this. Broadly there are two archetypes: the depth tracker, built on a curated database with micronutrient coverage and a preference for verified entries, and the speed tracker, built on an enormous crowd-contributed database with barcode scanning and fast repeat entry. Pick depth if you have a specific micronutrient question you are trying to answer. Pick speed otherwise. That choice is real, and it takes about thirty seconds to make.

The reason comparison tables run to three thousand words anyway is that they are measuring the smaller error. Database quality is a genuine difference between tools and it is almost never the thing that makes someone's log wrong. Two other sources of error dominate it, and neither appears in any feature grid I have seen: the meals you never enter, and the portions you guess.

Missing-meal bias: why your log is wrong in a specific direction

Here is the pattern I would most like other people building in this space to name and design against. I will call it missing-meal bias: in any self-logged food record, the meals that go unlogged are not a random sample. They cluster around unstructured, late, social, and high-energy eating.

Think about the mechanics of your own logging. The 7am porridge gets logged, because you eat it alone, at a table, from a packet with a barcode, at a moment when your day has not started going wrong yet. The 11pm handful of whatever, the second helping, the restaurant meal with no listed ingredients, the four drinks: these are the entries that require the most effort and arrive at the moment you have the least appetite for effort. So they are the ones that quietly drop out.

This matters more than it sounds, because the missingness is not noise. Noise averages out. Bias does not. A log with eighty per cent coverage that is missing a random twenty per cent of meals gives you a usable estimate scaled down. A log missing exactly the twenty per cent of meals with the highest energy density gives you a confident, precise, systematically low number. And it does it while showing you green rings.

The uncomfortable consequence: a tool with a slightly worse database and materially lower friction per entry will usually produce a truer picture of your week than a rigorous one you use on four days out of seven. Precision applied to a biased sample is just a well-dressed wrong answer.

So when you are comparing trackers, the question is not "which database is better". It is "which of these will still have my Friday night in it".

Portion estimation is the bigger error term, and I will make that falsifiable

Second claim, stated plainly so you can disagree with it: for a mixed, home-cooked meal, portion estimation contributes more error than the choice of food database does.

The test is cheap and you can run it tonight. Serve yourself a dinner you would normally eyeball. Before eating, weigh each component. Then compare your honest pre-weighing guess against the scale. For rice, pasta, oils, nut butters and cheese, most people are not out by three per cent, which is roughly the sort of discrepancy a database argument is about. They are out by a third. Oil in a pan is the worst offender, because the quantity is invisible by the time the food reaches the plate and it carries more energy per gram than anything else on it.

Now put those two errors in order. You can pick the tracker with the most rigorously verified entry for cooked white rice, and it will not help you if the entry says 100g and the bowl held 165g. The database argument is an argument about the second decimal place of a number whose first digit is uncertain.

This is the honest case for photo-based estimation, and also its honest limit. A photo does not know what is under the surface and it cannot see the oil. What it does do is remove the failure mode that produces missing-meal bias, because the cost of recording a meal drops to the cost of taking a picture. In building Numi the reasoning was explicit: I would rather have an approximate record of every meal than an accurate record of the meals somebody had the patience to search for. Coverage beats precision when what you are looking for is a pattern rather than a single number.

That trade only works if the estimate is presented as what it is. A vision model will commit to an answer whether or not it is right, and if you show that answer as a verdict, every near-miss chips away at trust in a way the user cannot articulate. Presented as a proposal you can correct in one tap, being wrong becomes an ordinary part of the interaction rather than an error state. Whichever tool you pick, notice how it behaves when it is wrong. That tells you more than the marketing page.

Check the target, not just the log

Every calorie tracker comparison examines one side of the arithmetic. There are two.

Your target is typically derived from the Mifflin-St Jeor equation, which is unremarkable and fine, multiplied by an activity factor. The factor is where tools quietly diverge. The published range runs from about 1.2 for genuinely sedentary through to roughly 1.9 for people doing hard physical work or two training sessions a day. Plenty of onboarding flows compress that into three buttons: sedentary, normal, active.

Do the arithmetic on that compression. For a person with a basal rate around 1,600 kcal, the difference between a 1.375 and a 1.55 multiplier is roughly 280 kcal a day. That error is not a one-off. It repeats every single day, and it is larger than anything you will win or lose by preferring one food database to another. If you sit at either end of the range, and lots of people do, a three-option activity picker gives you a target that is wrong before you log a single thing.

So during a trial, open the settings and look at how many activity levels there are and whether you can see the multiplier. If the tool is willing to show you how the number was produced, you can sanity check it. If it is not, you are trusting a black box for the most consequential number in the app.

Three questions no comparison table asks

If I were choosing between two trackers with a fortnight of free trial in each, these are the three things I would check, in this order.

1. Does the target move when the scale moves?

Daily weight is mostly water, food volume in transit, glycogen and timing. If your calorie target changes because Tuesday was up 0.8kg, the app is reacting to noise and presenting it as feedback. That feels responsive and it teaches you, over a few weeks, to distrust the number entirely, because you can see it lurching for reasons that have nothing to do with your behaviour. What you want instead is a seven-day trend, with the stored target changing only on a persistent shift. Weigh yourself daily if you like; just make sure the app reads the line, not the point.

2. When progress stalls, does it tell you or does it silently adjust?

Detection and action are different jobs, and collapsing them is one of the most common design mistakes in this category. An app that quietly lowers your target after two flat weeks has made a decision on your behalf that you cannot see, did not authorise, and cannot argue with. It should say: your weight trend has been flat for sixteen days at a stated intake, here are the options. Then you choose. The distinction matters generally: goal intent belongs to the user, and calories, macros and hydration are derived from it. A tool that blurs the two will eventually overwrite something you decided on purpose.

3. Can you interrogate the score?

Most trackers now surface some composite daily number. Ask what goes into it. If the answer is a model, you cannot check it, cannot learn from it, and cannot tell whether today's 72 differs from yesterday's 78 for a reason you care about. A deterministic score with stated inputs is often less sophisticated and far more useful, because you can look at it and know which lever moved. An explainable number beats an accurate black box, and in practice the black box is not more accurate anyway.

A fourth, related check: are the insights about you? "Great choice!" is worse than silence, because it is visibly not about your day and it tells the reader that nothing in the app is really watching. An insight has to reference something specific and real, what you have already eaten, what is left, what you have done four days running, or it should not appear at all.

A ten-minute way to decide

Rather than reading another feature grid, run this:

  1. Log yesterday, from memory, in both. Time it. Not the barcode-scanned yoghurt, the awkward meal: the curry someone else cooked, the thing you ate standing up.
  2. Count the abandonments. Note every point where you nearly gave up or entered a rough substitute. Those points are where your future missing meals will come from.
  3. Open the target settings. How many activity levels? Can you see how the number was calculated? Can you override it and have the override stick?
  4. Enter three days of plausible weights, one up, one down, one flat. Watch whether the target twitches.
  5. Read the insight it gives you. Ask whether it could have been written about anybody.

Whichever tool survives that is the right one for you, and the answer will differ by person, which is why a universal verdict was never available. If you already own a food scale and enjoy the ritual of precise entry, the depth tracker will reward you and micronutrient coverage is a real advantage. If you have started and abandoned manual logging more than once, the constraint is friction, not data quality, and no amount of database rigour will fix it.

One last thing worth saying, because the comparison framing hides it. The purpose of logging is not the log. It is the pattern you can only see across weeks: that your protein collapses on the days you train, that Thursday is structurally different from Wednesday, that the deficit is real on five days and cancelled on two. A tool that gets you an imperfect but complete record and then reads it back to you honestly is doing the job. A tool that gets you a beautiful partial record is producing a very tidy artefact about a person who does not exist.

← All writing