Every serious evaluation of a new preclinical model reaches the same question, usually about twenty minutes in. Someone stops asking whether the biology looks more human, because they can already see that it does, and asks whether the data will hold. Will the model give the same answer next month, in a different technician’s hands, from a different batch of cells? That is the right question, and in vitro model reproducibility is the hinge most adoption decisions actually turn on. A model can be the most human system in the building and still lose credibility if a scientist cannot trust it to repeat.
So let me take the question head on, because I think the usual framing gets it backward.
The reproducibility problem is real, and it is not a primary cell problem
The numbers behind the reproducibility crisis are sobering. A 2015 analysis in PLOS Biology put the annual cost of irreproducible preclinical research in the United States at roughly 28 billion dollars, and estimated that about half of preclinical work cannot be reproduced as reported. When Nature surveyed more than 1,500 researchers in 2016, over 70 percent said they had failed to reproduce another scientist’s experiment, and more than half had failed to reproduce their own.
Here is the part that matters for anyone evaluating a gut model. The PLOS Biology team traced the single largest contributor to irreproducibility not to statistics or study design, but to biological reagents and reference materials: the cells, the antibodies, the materials themselves. The biology you start with is where most of the trouble begins.
That should reframe the whole conversation. Reproducibility is not a weakness peculiar to primary human cells. It is the central problem of the entire field, and the incumbent models are not exempt.
Caco-2 is not the reproducibility benchmark people assume
The reflex is to treat Caco-2, the immortalized line that has anchored permeability screening for decades, as the stable reference and primary cells as the risky newcomer. The data does not support that reflex. When laboratories run the same compounds through Caco-2, apparent permeability values disagree by about three-fold on average, and by as much as 24 to 44-fold at the extremes. A well-known 2008 study that sent five compounds to ten laboratories found interlaboratory ranges up to 44-fold, and even within a single lab the values swung up to 16-fold.
The model most people reach for by default, precisely because it feels safe, can put the same drug in two different labs and disagree by more than an order of magnitude. Passage number drifts, monolayers mature differently, protocols diverge, and the phenotype moves with them.

So the honest framing is not primary cells versus cell lines. Every in vitro system carries variability. The useful question is whether that variability is understood, separated into its parts, and controlled. That is a question you can actually answer, and it is the one worth asking of any vendor, including us.
Signal versus noise: where the variability actually comes from
The mistake is to treat all variability as one blob of uncertainty. It is not. In a primary cell gut model, the scatter you see comes from two very different places, and they call for opposite responses.
The first is biological signal. Cells from different human donors behave differently because people are different. Intestinal drug-metabolizing enzyme expression, to take one example, varies roughly 10 to 30-fold across individuals. When two donors give you different numbers, that is often not error. It is the human population showing up in your assay, which is the entire reason you left a single-genotype cancer line behind.
The second is process noise. Seeding density, confluence, coating, replicate design, handling, and lot release all introduce scatter that has nothing to do with biology and everything to do with control. This is the variability you want to drive out, and it is the variability a well-built model and a disciplined protocol are designed to contain.
Confusing the two is how good models get rejected and bad conclusions get published. Read biological signal as noise and you pool away the human relevance you paid for. Read process noise as signal and you chase ghosts. Getting reproducibility right starts with telling them apart.
Here is how the pieces break down, and where each one gets addressed.
| Source of variability | Signal or noise | How it gets controlled |
|---|---|---|
| Donor-to-donor differences | Biological signal | Donor qualification, and pooling or single-donor design chosen on purpose |
| Lot-to-lot differences | Mostly noise | Release testing against defined criteria before a lot ships |
| Well-to-well and plate-to-plate | Noise | Seeding, confluence, and barrier-integrity gates like TEER |
| Assay transfer between labs | Noise | Standardized kit format and protocol, not tribal knowledge |
| Replicate design | Neither, it is inference | Biological replicates, not repeated wells from one run |
Each of those deserves its own treatment, so this cluster breaks them out. I have written separately on how to read donor variability and when to pool, what lot release testing actually verifies, getting well-to-well consistency from seeding through confluence, why an assay drifts when it moves between labs, and why technical triplicates are not biological replicates.
How a human model earns trust
Reproducibility is not a property a model has. It is a discipline you build around it. With RepliGut®, that discipline runs on a few principles that map directly onto the sources above.
Primary cells are donor-qualified before they enter the system, so the biological starting point is defined rather than assumed. Lots are released against set criteria, so a batch that will not perform does not reach a customer’s bench. Barrier integrity is checked before a compound is ever dosed, so a monolayer that has not formed properly is caught early rather than misread as a drug effect. And the format is standardized so that the protocol, not undocumented habit, carries the result from one lab to the next.
None of that removes biological variability, and it should not. The goal is to preserve the human signal while stripping out the process noise, so the range you see reflects real people rather than inconsistent handling. That is what makes a human-relevant model usable rather than merely interesting.
This same discipline underwrites the RepliGut platform across applications, from permeability and intestinal metabolism to inflammatory and barrier biology. The biology changes with the question. The reproducibility standard does not.
Why this matters now
The regulatory direction makes the reproducibility question sharper, not softer. The FDA’s April 2025 Roadmap to Reducing Animal Testing in Preclinical Safety Studies leans on human-relevant new approach methodologies precisely because animal models translate so poorly, noting that more than 90 percent of drugs that look safe in animals still fail in humans. As human in vitro systems move closer to decisions that used to rest on animal data, the bar for trusting them rises. A model that cannot demonstrate control over its own variability will not carry that weight.
Reproducibility is not the enemy of human relevance. It is the discipline that makes human relevance usable. The models worth adopting are the ones that can show you both the human signal and the control around it, and tell you which is which.
If your team is weighing a human intestinal model and reproducibility is the question on the table, that is the right instinct. Talk to the Altis team about how the data holds up, and ask the hard version of the question. A model built to be trusted should welcome it.


