Compared to Whom? The Bias Debate Forgot to Ask "Compared to When."

Elodie Aishwarya Remoissenet
Elodie Aishwarya Remoissenet
October 6, 2026·11 min read
Compared to Whom? The Bias Debate Forgot to Ask "Compared to When."

Three fields are arguing about the wrong axis. A dataset of seven thousand dogs shows why.

Brain size predicts how well a dog breed scores on memory and self-control. It predicts almost nothing else.

Across ten cognitive tasks and more than seven thousand dogs, Horschler and colleagues found a breed-level signal for executive function (short-term memory, impulse control) and nothing for six of the ten tasks outright. The associations that did show up for following a human's pointing gesture didn't survive the controls: once you account for training history, or estimate brain size from skull shape instead of body weight, they fall away. What's left is narrow and honest. Breed tells you a little about one slice of cognition. About the rest of the animal in front of you, it is close to silent.

That finding comes from one of the largest dog-cognition datasets ever assembled, and a separate line of evidence points the same way. In 2022, a genomic study of more than 2,000 dogs, paired with surveys from over 18,000 owners, found that breed explains only about 9 percent of the variation in behaviour between individual dogs. There is too much variation inside a breed to see much of a breed signal at all.

Hold that thought. It is not about dogs.

Every system that judges an individual by a population label is making the same bet: that the label carries the signal. A credit model bets the ZIP code carries it. A hiring filter bets the school carries it. A clinical algorithm bets the diagnostic category carries it. A behavior model bets the breed carries it. Sometimes the bet pays. Mostly it pays far less than the confidence attached to it.

We have spent a decade arguing about that confidence. We called the argument bias. We have at least three different fields each holding a different thing by that name, and not one of them is naming the actual error.

Three "biases," one blind spot

In machine-learning engineering, bias is one half of the bias-variance tradeoff. It is the error you get from assumptions that are too simple: a model that underfits, that flattens the world into a line it was never shaped like. The standard cure is to add complexity until the line bends to the data, then stop before it starts memorizing noise. Useful. But notice what this version of "bias" never questions: the reference distribution itself. It assumes the population you fit to is the population you'll be judged against. A population reference is a simplifying assumption about the individual. The bias-variance decomposition has no term for getting that assumption wrong.

In machine-learning ethics, bias means unfairness across groups. Here the field has split into two camps. Group fairness asks for parity between protected groups: similar acceptance rates across race, sex, and so on. Individual fairness, in Dwork's original formulation, asks that similar individuals receive similar predictions. The two camps spend most of their energy on the tension between them; a 2026 review of the trade-offs finds the field still at an early stage of maturity, and a separate line of work argues the two can be formally incompatible.

That incompatibility is not hypothetical. It is the COMPAS story: the recidivism score ProPublica dissected in 2016 was calibrated across racial groups yet produced unequal false-positive rates, and the impossibility theorems that followed proved no score could satisfy both criteria at once when base rates differ. A decade of argument, all of it about which sideways comparison to privilege.

But look at what they share. Group fairness compares you to a group. Individual fairness compares you to other people who resemble you. Both are cross-sectional. Both answer the question compared to whom. Both are looking sideways, across a population, at a single instant, for the reference that tells them whether you are normal.

And in dog cognition, bias shows up as the breed stereotype: the assumption that the label on the animal predicts the animal. Horschler's data is the cleanest natural experiment we have against it, precisely because it is not an ethics argument. No one is arguing about fairness to Border Collies. It is just measurement, at scale, and the measurement says the within-breed spread swamps the between-breed signal for most of what cognition is.

Same blind spot, three times. Every one of these debates asks compared to whom. None of them asks compared to when.

The axis nobody is on

Here is the move the bias debate keeps skipping.

If you want to know whether an individual is changing (declining, recovering, drifting, breaking) there is exactly one reference that is not a guess about which crowd they belong to: their own past.

A calm dog and an anxious dog have different baselines. Comparing both to the breed average tells you which one is "more typical." It does not tell you that the calm one has quietly slid two standard deviations off its own normal over the last month, because that slide can stay comfortably inside the population's wide band the whole way down. The population reference doesn't see the change. It was never built to. It is the wrong axis.

This is not a refinement of individual fairness. Dwork's similar-individuals-treated-similarly is still a sideways comparison: it references you against a cross-section of people who look like you right now. The longitudinal frame references you against yourself across time. One asks are we treating like cases alike. The other asks is this case still itself. They are different questions, and the second one is the one that detection actually requires.

You can debias the dataset and still miss the decline

The strongest counter is the obvious one: so fix the reference. Balance it. Debias it.

The state of the art is genuinely good at this. A 2024 MIT method called D3M finds the specific training examples that drive a model's failures on underrepresented subgroups, removes just those, and retrains, recovering worst-group accuracy without gutting overall performance, and without needing the subgroups to be labeled in advance. It is the most accessible debiasing tool I know of, and it works.

It also operates entirely at the subgroup level. It makes the population reference fairer. It does not make it yours. You can rebalance the dataset, equalize the subgroups, even match yourself to the most similar individuals in the cohort, and still miss the decline, because the decline is defined relative to a baseline that no cross-sectional fix contains. The error was never in the balance of the data. It was in the choice of reference frame. Debiasing is the correct answer to a different question.

The same lesson, found three more times

I am not the only one arriving here, and that matters more than if I were.

In maternal health, a 2025 analysis put it in its title: aggregated trends can be misleading. The authors show that population-level pregnancy monitoring overlooks individual variability, and argue for N-of-1 wearable analysis as the missing standard. In psychiatry, a 2025 Scientific Reports study trains a separate model per patient to predict psychotic relapse, framing the individual-level approach explicitly as the way to handle population heterogeneity. In respiratory medicine, current COPD work establishes each patient's personalized normal range from their first days of data and watches for change-point deviations against that, not against a cohort.

Different organs, different teams, no contact between them, the same correction. Each one independently discovered that the population's "normal" is the wrong thing to be normal against. That is what a real mechanism looks like: it keeps getting rediscovered by people who weren't looking for it.

It is the same mechanism I wrote about in the healthcare algorithm Obermeyer's team dissected: the model that used cost as a proxy for need, read equivalent illness in Black patients as lesser need because the system had historically spent less on them, and was corrected only when the target was reframed, lifting the share of Black patients flagged for extra care from 17.7% to 46.5%. That was a reference-class failure too. The proxy was standing in for the individual, and the proxy was systematically off.

The mechanism, made legible

You can watch the trap operate on a bench.

I built a head-to-head on fully synthetic data (no real animals, no sensors, the decline injected and known) to test one thing: does an individual baseline actually beat a population reference at catching a genuine change? Both arms use the identical detector, the identical robust statistic, the identical calibration window. The population arm even gets the pooled-cohort advantage. The only difference is the frame of reference.

As between-individual spread grows, the population's "normal" band widens until a real decline can hide inside it. At a matched false-alarm rate, the individual reference holds full detection across every heterogeneity level tested; the population reference falls toward roughly half. By area under the ROC curve, individual lands at 0.988 versus 0.935 for population. Same detector, same data, opposite outcome, decided entirely by what you compared the animal to.

I want to be exact about what this is and isn't, because overclaiming is the disease under discussion. This is a proof of concept, not a validation on real dogs. The advantage is conditional, not universal: when the cohort is nearly homogeneous, the population reference does about as well, and is marginally better by AUC. The premises are supported by existing science: individuals vary, population labels are weak proxies for them. The mechanism is reproducible under controlled conditions. The real-world comparison remains the open step. Three claims, three different strengths. Keeping them apart is the whole point.

Compared to when

The bias debate has been a fight about whose average to measure you against. Which group, which cohort, which set of people who look like you. It was a real fight and it produced real tools.

But it was the wrong axis. For anything that moves (a patient, a pregnancy, a recovering mind, an aging dog) the question was never compared to whom. It was compared to when. The most honest reference for an individual is not a fairer crowd. It is the individual, yesterday.

Breed is a weak proxy. So is every population average held up against a single life. The fix is not a better stereotype. It is to stop using one as the baseline.


Notes

Individual baseline AUC 0.988 vs population reference 0.935, at matched false-alarm rates. Full methodology and code: github.com/labs-barkley/barkley-reference-architecture (head-to-head results in results/HEAD_TO_HEAD_RESULTS.md).

This is the second piece in the reference-class trap series. The first: "Your Model Doesn't Have a Bias Problem. It Has a Reference-Class Problem."

Sources

Horschler, D. J., Hare, B., Call, J., Kaminski, J., Miklósi, Á., & MacLean, E. L. (2019). Absolute brain size predicts dog breed differences in executive function. Animal Cognition, 22(2), 187–198. https://doi.org/10.1007/s10071-018-01234-1

Morrill, K., et al. (2022). Ancestry-inclusive dog genomics challenges popular breed stereotypes. Science, 376(6592), eabk0639. https://doi.org/10.1126/science.abk0639

Geman, S., Bienenstock, E., & Doursat, R. (1992). Neural networks and the bias/variance dilemma. Neural Computation, 4(1), 1–58. https://doi.org/10.1162/neco.1992.4.1.1

Dwork, C., Hardt, M., Pitassi, T., Reingold, O., & Zemel, R. (2012). Fairness through awareness. Proceedings of ITCS 2012, 214–226. https://doi.org/10.1145/2090236.2090255

Trade-offs between individual and group fairness in machine learning: a comprehensive review (2026). arXiv:2602.00094. https://arxiv.org/abs/2602.00094

Friedler, S. A., Scheidegger, C., & Venkatasubramanian, S. (2016). On the (im)possibility of fairness. arXiv:1609.07236. https://arxiv.org/abs/1609.07236

Angwin, J., Larson, J., Mattu, S., & Kirchner, L. (2016). Machine bias. ProPublica. https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing

Kleinberg, J., Mullainathan, S., & Raghavan, M. (2017). Inherent trade-offs in the fair determination of risk scores. Proceedings of ITCS 2017. https://arxiv.org/abs/1609.05807

Chouldechova, A. (2017). Fair prediction with disparate impact: a study of bias in recidivism prediction instruments. Big Data, 5(2), 153–163. https://doi.org/10.1089/big.2016.0047

Jain, S., Hamidieh, K., Georgiev, K., Ilyas, A., Ghassemi, M., & Mądry, A. (2024). Improving subgroup robustness via data selection. NeurIPS 2024. https://arxiv.org/abs/2406.16846

Goodday, S. H., et al. (2025). Maternal health aggregated trends can be misleading: the power of N-of-1 level wearable data analysis for personalized pregnancy monitoring. medRxiv. https://doi.org/10.1101/2025.10.20.25338363

Relapse prediction using wearable data through convolutional autoencoders and clustering for patients with psychotic disorders. Scientific Reports, 2025, article 18806. https://pmc.ncbi.nlm.nih.gov/articles/PMC12122716/

Differentiating the start of an exacerbation from day-to-day variation in people with COPD: a systematic review. European Respiratory Review, 35(180), 250212, 2026. https://publications.ersnet.org/content/errev/35/180/250212

Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342

Remoissenet, E. A. P. (2026). Barkley Reference Architecture (software). Zenodo. https://doi.org/10.5281/zenodo.20369863

Disclosure

The research question, experimental design, analysis, and conclusions are mine. AI was used for drafting assistance and language editing. I disclose that use openly. I'm proudly AI-assisted.

Elodie Aishwarya Remoissenet
Elodie Aishwarya RemoissenetArtificial Intelligence, Data Science & Analytics, Machine Learning, Technology, Entrepreneurship

Founder and independent researcher. I run Barkley Labs: behavioral intelligence for dogs (Barkley AI), an open evaluation protocol for hiring (IREP), and a cryptographic provenance protocol for music (ACTA MUSIC). I write about decision intelligence and the reference-class trap: systems that judge each individual against a population average instead of against their own trajectory, in health, credit, hiring, music, and the companion-animal work where I test the mechanism. Individual over population, always. Proudly AI-assisted.