Essay
The Scientific 99.99%
29 July 2026
All science is either physics or stamp collecting -attr. Ernest Rutherford
I certainly don’t agree with the quote, but I understand the sentiment. I used to agree, generally, but then I started learning about the science of science - ‘scientometrics’, and realised how many of works of science are the direct result of this scientific gruntwork.
- RNA vaccines were first postulated in 1989 before breaking out in 2020
- CRISPR-Cas9 first appeared in 2002. Ten years before Doudna’s Nobel-prize winning work
- Large Language Models can be found throughout the 1990s before ChatGPT launched in 2022
Modern science is reported as the story of leaps. 45K papers have received more than 1,000 citations, forming a backbone of discoveries that have propelled science. This is compared against 123M papers on OpenAlex, which form the incremental steps that have lead to these discoveries reaching technology maturation and a strengthened community of researchers? Modern science is about a community identifying the gaps in knowledge and filling them, and that’s what the scientific method’s all about. Filling the gaps, knowing the new the gaps and deciding where to go next.
The problem here, is that the corpus in fragmented by discipline and the 45K papers that make the glam-mags overshadow a vast majority. But what is this 99.99% of papers doing?
Most reading this will understand for early career scientists, the projects you take are mostly based on skills which take time to hone, rather than the domain knowledge which can be easily mopped up. It might have been a weird choice for a 20-year-old Welsh boy with crippling accent anxiety to then do a masters in Welsh sheep genomics, but what I was doing was much more important than looking at Welsh sheep breeds.
I was discovering valuable mutations that have been selected for. Usually breeds of animal in commercial senses are bred for productivity, and with that comes susceptibility to disease and stress when compared to their “wilder cousins”. Think of the Texel, which is now crossed with other breeds to increase hardiness. So what was important was adaptation to local environment, but also how sheep play their own role in the local environment: genes in mitochondrial biogenesis, oxidative stress pathways. Due to 2020 happening, I was able to verify none of this in the lab to prove if these were activating or suppressing. Arguably these are the missing pieces of the puzzle, but in that we build up databases of potentially useful mutations with seemingly beneficial effects.
The reason we do this (or at least what is put on the grant applications) is to understand these mutations, transplant them to other breeds with the view of making them much hardier to the oncoming environmental stress iun a warming world. The way it’s done is by researching the literature around the mutations your bioinformatics pipeline suggest are under selection and then identify why it might be important, and see what fits the story. It’s not the most rigorous approach, but people are doing this all around the world: Russian sheep breeds, Chinese sheep and Brazilian sheep to name some examples that have cited my work.
So looping around on the point of this article, I move onto the idea of AI scientists. I’m hearing a lot about autonomous labs, AI models that will be able to observe phenomena, test hypotheses and change accordingly. The capabilities of these AI scientists is largely unknown, to which, my question is: will they be doing the science or the stamp-collecting? We can hope that they could do the science. The US is spending $5b to find out: That’s 10% of an annual NIH budget, and about 4x the annual UKRI budget.
The gruntwork mentioned throughout this thought piece is what the models get trained on. Repositories like what I mention with my sheep example make things like protein language models possible. Which was the basis of my PhD. It’s the reason why big pharma companies are raiding stockpiles of lab books for proprietary data to train their own models. I think this could all be federated to achieve more generalisable models quicker, without surrendering training data, but companies be companies.
What I do think is possible and more achieveable, is more akin to the Modern Synthesis: introducing the mechanics of Gregor Mendel’s inheritance laws with Darwinian evolution. Now the scientific corpus is more searchable, autonomous systems can be made to scrape the corpus, evaluate the gaps and draw our those precious non-obvious inferences that are cross-disciplinary or stem from work that was largely overlooked. Could discoveries in cancer therapies help develop more resiliant crops? This would firstly, keep our human scientists in a job, but also reposition them with the best information, and the most prioritised work to tackle a world that is becoming more uncertain and more fragmented by the problems that face it.
Could AI scientists be the tools to unpack what has been left in 123M papers? This is the scientific 99.99%.