Great science, real science, is all about suspicion and gossip, deception, and occasionally divination to spot lemons, all so you don't get hosed in your quest for glory. At least, you might be left with that impression after your first few lab meetings.
One day you bring an exciting paper to discuss with the team, having spent hours working out what all these words mean, and you proudly exclaim:
"This is super cool, I think we can build on these findings and — "
"No, wait, we've been burned by this lab before —"
"Yeah. Every time we try to replicate their findings we can't find the same effect sizes."
Everyone nods along. Maybe it was fraud, or sloppiness, or maybe the paper just left out the one detail that makes it work. You stand there, somewhat dazed, wondering how you might learn to read the tea leaves and suss this out next time.
Many labs contain a memory like this: hard-won, battle-tested, accumulated over hours of chasing dead ends that other teams published. For competitive labs at the frontier, it's an edge — you can out-compete other labs that get stuck burning months on the wrong leads. Of course, some of this intel can be shared with vetted collaborators, but it's not for everyone.
Evidently, in the 21st century, science runs on gossip, a whisper network of sorts built on informal reputation systems, orally transmitted.
It's tempting to interpret this as a story about fraud, and sadly there's a bit of that, but there are plenty of other things that get in the way of trusting a paper, including intentionally vague methods sections to slow down competitors so the original lab can beat them to market on follow-on ideas.
I previously argued that the journal as a coordination technology bundles a few jobs — disseminating your work, vouching that it's good, and signaling its importance. Indeed, I believe that any strong system of science enables the efficient dissemination of validated knowledge to whoever needs it in a way that they actionably understand.
What constitutes validation, and our ability to establish trust? How have we tackled that problem before, and how will that evolve as we develop new capabilities that change how we think about executing science?
As machine generation becomes abundant, we must learn to navigate a new landscape of validated generation, where trust at a distance is reimagined as distributed public infrastructure. It's Boyle's virtual witnessing again, and, once again as co-scientists join the network in years to come.
The Honorable Robert Boyle's Literary Technology
In 1659, Robert Boyle unveiled an elaborate scientific instrument: a glass globe the size of a beachball, mounted on a brass cylinder with a hand-cranked piston that could pump the air out of it. When the air was gone, he observed that candles died, and animals would struggle and suffocate. Boyle had conjured a pocket of near vacuum on his desk and used it to show that air is necessary for fire, and for life1.
At the time, the very idea of a vacuum was disputed. Many serious thinkers held that "nature abhors a vacuum", that truly empty space couldn't exist. His findings were controversial, and worse, were produced by a machine that others did not own, one that was fiendishly hard to operate.
The physical reality of the pump was a mess. The glass globes cracked under pressure, so Boyle would patch them with a paste of water, quicklime, and cheese scrapings; the piston was sealed with leather soaked in salad oil, the joints with homemade cement. Because of these imperfections, Boyle could never create a perfect vacuum, only a partial one.
In one experiment, he hypothesized that two smooth marble discs held together by the pressure of air would fall apart once the air was pumped out. They didn't. No matter, residual air must have crept in through the weak seals. When his experiments yielded messy or incomplete data, he blamed the leaks. Outcomes that matched the theory proved the pump had worked; outcomes that didn't proved the pump had leaked. His fiercest critic, Thomas Hobbes, saw the deeper problem: Boyle had no independent way to determine whether the machine had worked.
How do you know the pump ever held a real vacuum? Only because it produced the results Boyle's theory predicted. And what was the evidence for this new theory? Well, the results of the good runs. It was a closed loop with no exit by logic alone. Today, philosophers of science call this the experimenter's regress, a common occurrence in frontier research2. Hobbes' underlying critique was that he who calibrates the instrument controls the definition of the fact.
Boyle escaped this not with logic, but with two structural inventions to enable social consensus. First, a literary technology: he wrote his experiments in obsessive, tedious detail and documented each leak, failed run, and even published a detailed engraving of the actual machine. This dense reporting allowed a reader from far away to picture the scene vividly enough to become, in Shapin and Schaffer's phrase, a virtual witness3.
The second invention was social: he ran the experiments in front of a room of gentlemen and wrote down their names. Their word was binding — after all, these were the good old days when an Oxford professor was a more reliable witness than an Oxfordshire peasant. Dense reporting and collective, high-status assent turned a leaky machine's performance into an accepted "matter of fact".
What Changed Since Boyle: Scale and Settlement
I'd like to think we've come a long way since the 1600s, but the mechanism we use to establish trust hasn't changed significantly. Boyle's twin pillars of prose and reputation cannot survive the sheer math of modern scientific expansion.
His time was a time of little science, with gentleman scholars, few enough that an informal reputation system could sustain the network. And because so little had been established before them, they were almost by definition working at the frontier, where consensus and discourse are epistemically essential.
After the Second World War, we massively expanded the number of scientists, and the number of papers produced exploded along with them. Derek J. de Solla Price tried to describe this through what he called the coefficient of immediacy: the ratio of scientists alive now to all who have ever lived. His estimate was about seven in eight — roughly 87% of everyone who ever did science was his contemporary4, and it's only more lopsided today. Statistically, this is the greatest moment there has ever been to be a scientist. Buried in an avalanche of information, it rarely feels that way.
Scientific participation became global, aided by the internet, translation software, and now AI. The open science movement reshaped how results get shared. And most recently, we've got emerging capabilities like co-scientists that run science autonomously — from difficult math problems to designing and executing experiments — and cloud labs edging toward full automation5.
The character of the work changed too. Not every discovery today is operating at the frontier; many are iterative advancements within established frameworks, in fields at different stages of maturity, where things like instruments and methodology might be more settled.
If you measure these initiatives across two axes, how well they validate a result and how well they disseminate it, you start to see a picture like this:

The internet and open science pushed us on the dissemination front but didn't necessarily scale validation. Cloud labs and co-scientists can give us real new power to check findings, not just run new experiments, but they sit in the hands of private companies.
So what sits in the empty corner, with strong validation and strong dissemination? I think we get there by properly networking capabilities we have: first with public infrastructure that kills the need for divination, and then by evolving new primitives for discovery, selective disclosure, and reputation.
The Shape of Trust
For Boyle and other gentlemen of the time, trust came from literary technologies and honor. Today, we face new problems driven in part from scale, and from a shifting social order.
Even with pockets of little research tucked inside the big, there is simply too much research for anyone to vouch for. Worse, literary technology is weaker today than in Boyle's time because publish-or-perish incentives combined with AI tools that generate plausible prose overwhelm our nascent peer review system, creating an asymmetric verification bottleneck for publishers and societies, weakening our ability to read a paper and see it and believe it6.
And of course, the human-to-human nature of science is changing; soon we'll have to reckon with:
- A person trusting another person's research
- A person trusting a co-scientist's research
- A co-scientist trusting a person's research
- A co-scientist trusting another co-scientist's research
In a time of abundant ideas, novelty comes from being able to properly recombine what already exists7 — and you can't recombine what you can't trust.
In thinking about this, I've found it helpful to separate objective and subjective notions of trust.
Objective trust looks like provenance and attestation:
- Is this paper actually the original paper (version of record)?
- Did you actually do what you said you did?
- Do the cited papers exist, and did you cite them correctly8?
- Do the numbers add up, and does the code produce this table9?
- Can we reproduce these findings? Not "are these claims true", but the narrower, answerable "under this exact setup, do we see a similar change"10
These checks typically have a right answer, and to the extent possible, their implementation should be automated away in the background, like an invisible layer that gives you confidence, the way HTTPS quietly does on the web. Failures here are unforced errors, and given how measurable they are, that number should be measured and aggressively pushed down to zero.
No, this shouldn't demand extra work on the side of scientists; the tooling is already here to make sure we capture this natively, as exhaust during work rather than a form to be filled out after. The closest analogy is a software bill of materials (SBOM) for a research paper11.
Subjective notions of trust are trickier because they're social, and often left to interpretation and debate. Track record matters, not as a static mark that follows you but a trend measuring shifts in reliability. In frontier science, when issues like experimenter's regress crop up, scientists need to triangulate, debate theory, evaluate experimental design, and interpret claims properly based on the available evidence.
This subjective layer sits separately from the objective core. A manifest can attest that some code produced this table, not whether the instrument measures what you think it measures. That question is irreducibly discursive, and the point of the objective layer is not to resolve it but to ensure the debate starts from an attested ground truth. For example, as a reader, I only want to worry about whether your design and interpretations support your hypothesis, not whether you actually did what you said you did.
There is a varying emphasis on objectivity and subjectivity depending on how settled a field is. Frontier research based on new theory with new instrumentation, or new models, or paradigm shifts, typically has more burden on subjective notions of trust. Iterative research that's advancing ideas in relatively more established fields still needs this, but the objective aspects start to have a greater impact on overall rate of progress.
Because interpretation and collaboration are inherently social, trust requires boundaries. Your institutional container, your lab, company, university, even country all come into picture, so you'll likely give a competitor, collaborator, and stranger different views of your work.
The record in this world would have to allow for selective disclosure. Access can be managed in concentric rings using selective decryption keys: grant full visibility to your immediate lab, partial parameters to a vetted collaborator, an encrypted attestation to a reviewer, and nothing at all to a direct competitor until the paper is published. We already do a crude version of this today, ferrying PowerPoint slides and spreadsheets through the clandestine secrecy of email threads and machine-hating file formats.
In practice, you might start with a cryptographic commitment to timestamp your priority claim on an open registry, while executing the underlying code or protocol inside a trusted execution environment (TEE). The system can issue a remote attestation proving your result ran unmodified and produced a particular output12. More tactical details coming in a separate essay.
Literary technology breaks when generation is abundant; the replacement isn't better prose, but three things working together:
- Deterministic capture of what actually ran (manifest with signed attestation)
- Automated re-execution (co-scientist or cloud lab that re-runs the method instead of reading, imagining, and believing)
- Synthesis at the point of reading
Where Boyle's literary technology was meant to capture scenes vividly so you could reconstruct it, these new primitives capture enough detail, with minimal extra work to the scientist, that anyone can reconstruct exactly what happened.
A co-scientist witnessing another co-scientist doesn't read the methods section, but pulls a signed trace, re-runs it, and checks the attested output against the claim. A human virtually witnessing a cloud lab's run doesn't need to take the vendor's word for it; they open an interactive view which renders the run's provenance, like an audit trail.
The ultimate value of automated re-execution is in its speed. Traditional replication, if it happens, can take months to years, and only notifies the broader field through citations or editorial notices like corrections and errata long after a shaky house has been built on top. In the future, we'll be able to witness the claim as it lands, regardless of where we are.
The paper shifts from carrying a burden of contextualization-on-write, remembering only what the authors chose to say at the time, to more like a runtime that enables someone to understand what happened on-read: has it been reproduced? which attestations still hold? Witnessing shifts from an act of persuasion to a continuous act of synthesis by the reader.
The witness in the future is whoever, or whatever, is building on existing work, and what they'll see is no longer just a narrative, but a re-executable record.
Fix the Floor, Then Raise It
The US alone spends $28 billion a year on preclinical research that can't be reproduced; if you're brave enough to count the downstream work built on shaky findings, the house of cards estimate runs anywhere from $13 billion to $270 billion a year13. This is our tax for failing to maintain a public record of what holds up.
There's clearly an imperative to focus on the objective baseline, and I think we can accomplish it in 3 phases:
- Capture a research "bill of materials" on submission — a manifest that attests to the objective aspects of trust.
- Promote the development and maintenance of a federated reproducibility commons — a public, shared record of which papers and methods are and aren't reproducible, starting with software-based fields and reaching wet-lab fields as the cloud-lab ecosystem matures. Note, not everything will fit the automation paradigm (e.g. clinical trials, or field studies).
- Develop primitives that enable selective disclosure across networked contexts, preserving priority claims and the natural collaboration scientists are accustomed to (note: I personally view wanting to be first to a discovery as a relatively fixed behavioral quality of scientists for at least the next few generations).
A bill of materials for research papers is a pressing and tractable need — it would alleviate pressure off submission and peer review, and having it in place starts to rebuild trust in the published record.
Concretely, for a computer science or finance paper that's mostly computational, a manifest might capture things a machine can emit on its own — the exact container image, dependency locks, dataset versions, their DOIs, code at a pinned commit, bundled into an execution trace11. For a wet-lab paper, that might be a machine-readable protocol itself. In both cases, the manifest is a byproduct of doing work with modern tools, captured as exhaust, not a new form to fill out afterwards.
We can run a similar class of tooling retrospectively over existing papers too, assembling a shared map of statistical inconsistencies, or, from a forensic angle, signs of image manipulation of the kind that eventually surfaced in a landmark Alzheimer's paper14.
The real unlock is using co-scientists and cloud labs to execute reported methodology and publish a shared, public view of whether findings reproduce as a way to build awareness of what holds up and what doesn't.
Who pays for all this? Cost varies across disciplines. In software fields, it's cheaper — re-running a container is renting compute, and the community already treats it like a volunteer sport (the ML reproducibility challenge pays about $500 a paper and now runs as an official NeurIPS track15). There, the commons can afford to check most things, and the sane default is to make it a submission gate: funder or venue requires a manifest, a CI-style service re-runs it, and a graded badge attaches to the result.
Wet-lab economics are more expensive and resource-bound today, even with cloud labs16. In some cases, instrument availability could be a blocker. Rather than reproduce everything, we can triage and spend a re-execution budget based on the stakes of the project being funded and what they build on.
As frameworks from The Institute for Progress suggest, targeting replications at highly influential, early-stage research maximizes the return on investment, showing that a well-designed program can pay for itself many times over by pruning dead ends before they compound17. Demand pulls the funding forward rather than forcing a public commons to bankroll everything.
Once reproduction is reported openly, teams will have more reason to report their methods more accurately, lifting the quality of methods sections across the board. There will be bumps on the road here. Cloud lab capacity is still thin and expensive16, and today's best systems reproduce an easy paper often and a hard one about a fifth of the time18. But it's only a matter of time as these systems get better. Career implications have to be factored in, specifically around how that information is (and is not used) in influencing career decisions.
From a policy perspective, we'd be wise to ensure this information exists publicly. Our status quo behavior of private knowledge accrued by incumbent groups is an information rent — valuable because it is hoarded — and it sustains a market for lemons where newcomers can't easily tell the good from the bad, forcing everyone to discount everything19.
We competed away that rent once before: prior to 1933, information about whether a stock was sound was inside knowledge, until the SEC forced companies onto a standard public ledger and let outsiders trust a market they couldn't personally inspect20.
Verification is science's version, and it belongs in the public. Leave it to the market, and it goes the path of large commercial publishers exacting a private toll on trust itself, capturing the verification stack as an aspect of their prestige moat21.
To be clear, public does not mean a monolithic agency checking each paper or hosting a centralized repository — that would be too easy for political or corporate interests to capture, and scales poorly across highly diverse fields.
What I imagine is closer to a federated network, possibly run by field-specific non-profits, academic societies, and accredited university nodes. We see this emerging organically across disciplines already. The state's role is not to operate the machinery but to fund the open standards and provide baseline infrastructure grants, similar to how NIST maintains absolute measurement standards while leaving local deployment to thousands of independent laboratories22. Execution has to be open and decentralized.
Ultimately, this points toward a new institutional architecture. Just as the 17th century birthed The Royal Society, and the 20th century scaled the research university to absorb industrial funding, the era of abundant machine generation requires individuals and institutions organized around a shared, public protocol.
Do Androids Dream of Peer Review?
This doesn't mean the end of the paper. Narratives can and should still exist to aid in contextualization and understanding, but they should be a dynamic viewport, rather than a static view upon which we place the full weight of our trust.
When Boyle invited gentlemen to observe his air pump, witnessing was a sensory, localized event. You were either in the room, or you read prose so detailed you could imagine being there. How does a machine "virtually witness" a result? Ideally by pulling a signed execution trace, spinning up a runtime environment, and verifying its execution.
Moving forward, our collective rate of progress will be determined by how well our infrastructure helps us verify, filter, continuously synthesize, and even find surprise from the firehose. Scaling verification and trust through new primitives for virtual witnessing is necessary in a world of abundant generation.
I don't think existing discourse has painted an optimistic enough picture of what that could be. Most imagine a scenario where co-scientist capability and capacity from the future will enter the system state today, which naturally creates anxiety for many overworked and underpaid scientists.
In reality, cheaper verification means more trustworthy surface to build on, which means more buildable ideas per scientist. If we get the foundation and capabilities right as the system matures, the knock-on effects will induce more demand (Jevons Paradox), and we'll be able to grow with it. I'll unpack this more in a future essay.
The web only became the web once trust moved from "do you know this webmaster" to a cryptographic handshake, a certificate authority vouching for the site, and public transparency logs keeping those authorities honest23; science needs that equivalent. Trust quietly shifts into the infrastructure, effectively reducing the number of people you have to vouch for.
When done right, all of this should hum in the background. An investigator, whether they're a human scientist at the frontier or a co-scientist executing experiments in more settled fields, shouldn't feel anything except confidence in the work they build on, backed by the strength of a verified public record.
Maybe, just maybe, we can lift the malaise that subdues modern research, allowing scientists to experience the rare, unburdened joy of building on what is known.
More soon,
Ashish
P.S. I said earlier that "any strong system of science enables the efficient dissemination of validated knowledge to whoever needs it in a way that they actionably understand". You can apply that to most information systems, only the stakes are different!
P.P.S. You can subscribe below to be notified of new essays in this series.
P.P.P.S. If you read all this, tell me your favorite sentence. If it's the same as mine, drinks on me when we meet up!
-
Collins, H. (1985/1992), Changing Order: Replication and Induction in Scientific Practice — the "experimenter's regress." source ↩
-
Shapin, S. & Schaffer, S. (1985), Leviathan and the Air-Pump: Hobbes, Boyle, and the Experimental Life — virtual-witnessing definition from Shapin, "Pump and Circumstance" (1984). source ↩
-
Price, D. J. de Solla (1963), Little Science, Big Science — the "coefficient of immediacy," the ratio of living scientists to all who ever lived (≈ 7:8, ~87.5%); science doubling every ~10–15 years. source ↩
-
Autonomous / self-driving labs and AI "co-scientists": "The rise of self-driving labs in chemical and materials sciences," Nature Synthesis (2022); Ginkgo Bioworks' automated biology foundry; Google "AI co-scientist" (Feb 2025, company claims). source, source, source ↩
-
"AI in scientific publishing: Slower, worse, and more expensive," Science. source ↩
-
Weitzman, M. (1998), "Recombinant Growth," Quarterly Journal of Economics — new ideas are recombinations of existing ones. source ↩
-
"Fabricated citations: an audit across 2.5 million biomedical papers," The Lancet (2026). source ↩
-
Nuijten et al. (statcheck) — a statistical-consistency error in roughly half of psychology papers using null-hypothesis tests; and the AEA Office of the Data Editor (since 2018), which re-runs the code behind every empirical paper it accepts, raw data to tables, before publication. source, source ↩
-
Reproducibility baseline: Open Science Collaboration (2015), "Estimating the reproducibility of psychological science," Science; Baker, Nature (2016) — >70% of researchers have failed to reproduce another's work. source ↩
-
Provenance and attestation standards — SLSA supply-chain framework; Sigstore. source, source ↩↩
-
Freedman, Cockburn & Simcoe (2015), "The Economics of Reproducibility in Preclinical Research," PLoS Biology — ~$28.2B/yr (50% of $58.4B) US preclinical research that is not reproducible. source ↩
-
Charles Piller, "Blots on a field?," Science (2022) — image manipulation found across influential Alzheimer's amyloid papers (Lesné et al., Nature 2006, later retracted in 2024). source ↩
-
ML Reproducibility Challenge — community re-execution of published ML papers, now run as an official NeurIPS track. source, source ↩
-
The cloud-lab reality check — only a few real commercial cloud labs; Emerald Cloud Lab entry reportedly ~$250–300k/yr; much benchwork resists automation. source ↩↩
-
"How Much Should We Spend on Scientific Replication?" The Institute for Progress. source ↩
-
Starace et al. (2025), "PaperBench: Evaluating AI's Ability to Replicate AI Research," ICML — the best agent reproduces frontier ML papers ~21% of the time vs. ~41% for human PhDs; far better on well-specified tasks. source ↩
-
Akerlof, G. (1970), "The Market for 'Lemons': Quality Uncertainty and the Market Mechanism," QJE — asymmetric information collapses trust; certification as cure. source ↩
-
Securities Act of 1933 / Securities Exchange Act of 1934 — mandatory public disclosure to reduce information asymmetry. source, source ↩
-
Larivière, Haustein & Mongeon (2015), "The Oligopoly of Academic Publishers in the Digital Era," PLOS ONE — top-5 publishers held >50% of papers by 2013. source ↩
-
NIST metrology — Standard Reference Materials (SRMs) and metrological traceability. source, source ↩
-
HTTPS / TLS, certificate authorities, and Certificate Transparency. source, source ↩