162 points•jegp•12 days ago•74 comments•

74 comments

flopsamjetsam12 days ago
> Every result is instantly reproducible. When you read a paper claiming that a new drug reduces symptoms by 30%, you click a link and watch the exact analysis run in your browser. The data processing, statistical tests, and visualizations execute in seconds using the same environment the authors used—preserved perfectly through reproducible containers.

At least some journals have this as a stipulation e.g. https://www.nature.com/nature-portfolio/editorial-policies/r...

Particularly the "data availability" and "Availability and peer review of computer code and algorithm".

However, in my limited experience, of trying to reproduce certain scRNA-seq processing pipelines, in practice it's never available as just a Github link. I can understand that some/many researcher's code is not in good shape, so I think it'll be quite a stretch to have this available.

I do think it's laudable though, to try and make it available. It would certainly have been very useful for me in the past.

cge12 days ago
I try to do something like this with my publications, and encourage others to. My goal is to have the pipeline from raw data to complete figures and manuscript in a repository, with cached data for computationally expensive analysis and for stochastic simulation results, and the option for the user to just use those or run the full pipeline, with or without the same random seeds. I just make clear that the code was run-once code and is going to be messy compared to code refined over time and diverse uses. I generally use Zenodo to a GitHub repo, however, in case GitHub decides to do something bad in the future. Making sure things run far in the future can also be a challenge. Sure, you can use a container: will the base of that container be available in 30 years?

And with that said, for experimental work, this approach does not make things fully reproducible; it only makes the analysis reproducible. There are always factors that influence experiments: research is by definition at the edge of our understanding, and reality has countless variables, including ones no one has thought of, known about or thought important.

flopsamjetsam11 days ago
Caveat: I am not a researcher (yet), I am moving from coding into science via a new degree, and along the way I am helping troubleshoot bioinformatics pipelines for scientists.

I see a lot of reliance on containers to make code always available, and I have the same misgivings as you do. It'll work a few years into the future, but what happens once packages aren't compatible with each other/the base container is upgraded/etc.

I've already seen this with older bioinformatics code, which is on old repositories that aren't running anymore, or are very unreliable (but weren't at the time that the code was written). And I'm talking about code that's "only" 15 years old; people will be going to these papers for implementation details long after that point.

Zenodo seems like a good step in the right direction. It should be available as long as CERN is going, shouldn't it? And by that stage it should be "too big to fail".

amarcheschi12 days ago
You're goat, I'm trying to reproduce code from a paper and by following their instructions I can't even get packages to install because they conflict
willtemperley12 days ago
> modern science is synonymous with open source software.

Another problem with reproducibility is the openness of the underlying data. Many academics are terrified of giving away the golden goose and the software is often useless without the data.

However many scientists do work openly, e.g. The Journal of Open Source Software:

https://joss.theoj.org/

jegp11 days ago
I've seen quite a few academics wrangling with patents. What should come first? Proliferation of the sciences or (potential) profit? Your point about the golden goose is pretty interesting and others in this thread has pointed to the misaligned incentive structure in academia. I'm curious, do you think things like JOSS could help open up data and "the golden goose"? I'm not sure the funding bodies I'm interacting with would respect this kind of initiative, but maybe it's just a matter of time.
willtemperley11 days ago
My experience in academia was that data management is not well respected. Universities often see data as part of the library system which leads to a disconnect between data and research.

I’d suggest data management needs to be a well paid career in academia, with senior level influence to enable long term open access to data.

I was on a grant from the Wellcome Trust putting an epidemiology database online. This remained available until the PI moved on and their successor was decidedly less public spirited, so sadly it’s not available anymore. What can I do about that when I’m seen as a technician? I have over 8000 citations from my database work but there’s really no career in academia for me.

setopt12 days ago
JOSS is great, I’ve both published with and reviewed for them, and enjoyed it more than traditional journals. The process felt more constructive than destructive, in a sense.

I believe they’re always looking for new volunteers to review papers, so please do volunteer if you are able.

D-Machine12 days ago
Science should be more like this, in current times, yes.

But until much of academia is burned to the ground, or until science can be properly separated from modern academia, this will never be so. The current academic incentives are all wrong: low-quality research is rewarded and results in publications, whereas high-quality research (that takes time, and usually reveals that most exciting publications depend on p-hacking or other highly data-dependent analyses and selective presentations) is not published or actively blocked during peer review.

So instead you get BS arguments about how data can't be released for various privacy concerns (when in reality the vast majority of most datasets are trivial to scrub of identifying factors, and even in more complex datasets where you need to consider k-anonymity, it is still trivial to release data that allows replication of core analyses), and academic science is increasingly irrelevant unless it is tied to tech and industry, where producing junk actually has real negative economic and personal consequences.

I don't know what world this article / post lives in, but it isn't the messy world of actual reality.

epihelix12 days ago
It's changing, though, and articles like this are important. TFA is arguing for change in the future, not presenting this as a fait accompli.

In the 27 years I've been in academia, I've seen a lot of progress in data openness (NCBI GEO was a game-changer) and FOSS analysis software (it's now widely expected that a high impact pub will make all data and code available for review, and then publicly available upon manuscript publication; most major journals will not allow submission without this). It is becoming common for big journals to specifically ask reviewers to review the analysis code. It is starting to become more and more common for papers to release all the code used to generate all the figures (including supplementary figures)

There is still a long way to go, I agree. But it's always better to light candles than curse darkness, etc.

> academic science is increasingly irrelevant unless it is tied to tech and industry

While I have some sympathy with a lot of your bitterness, this statement is insulting silliness that a quick look at the list of Nobel Prizes in physiology and medicine would prove wrong. Almost all major breakthroughs in the applied sphere stem from decades of basic research that happened just because it interested someone.

D-Machine12 days ago
> it's now widely expected that a high impact pub will make all data and code available for review, and then publicly available upon manuscript publication

I am also in academia and regardless, factually this is not true at all for data, not even remotely (less than like 10% of journals even have data availability policies which are recommendations, and in practice only a small percentage of papers actually make anything available), unless by "publicly available" you mean "available to some academics or academic labs after an often tedious and slow approval process requiring an academic email and various signed agreements". Maybe what you are saying is true in some very specific domains (e.g. machine learning research), but in general what you are saying here is IMO wildly out of touch with present realities in the vast majority of fields, but especially those involving human subjects.

> While I have some sympathy with a lot of your bitterness, this statement is insulting silliness that a quick look at the list of Nobel Prizes in physiology and medicine would prove wrong. Almost all major breakthroughs in the applied sphere stem from decades of basic research that happened just because it interested someone.

Nobel Prizes are so rare they don't speak at all to the generalizations I am making here. Also, much medical academic research is arguably successful because it is in fact ultimately industry-funded or tied to industry. It is of course though highly dependent on the academic subfield, for sure, and I was painting with a broad brush.

If I had to narrow things, STEM academic research isn't so bad, so long as we exclude social science from STEM. Much social science research needs to be defunded ASAP. And I'm not claiming industry research doesn't also have warped incentives. But, on balance, I'd wager outside of pure math/physics and certain more algorithmic/pure domains in comp sci, the smartest people today are going to choose (and be found in) industry, not academia.

jegp11 days ago
Well said! And thank you for recognizing the effort.

The important part here is, as you say, to light candles and insist on rigor. Coincidentally, history tells us that that also gets us further. So by pure memetic selection, this strategy should win

fasterik11 days ago
"Burning academia to the ground" is a terrible idea. We need to fix funding and incentives. If funding and research all happen in industry, where are the incentives to do basic research?

Your analysis completely ignores the physical and biological sciences, engineering, and the humanities. It mostly applies to a small subset of academic fields in the social sciences and medicine. You're also ignoring the changes that have happened since the replication crisis. Preregistration, publishing all code and data, reporting null findings, replicating results, etc. are becoming the norm.

D-Machine11 days ago
Of course "burning it to the ground" is rhetoric and not meant literally. It is meant to convey though that "nice" and "gentle" solutions might not really be enough here.

Yes, for the most part the fixes have to be in terms of funding and incentives. Funding needs to be more careful, and more careful funding can be a carrot rather than a stick here.

Re: incentives, IMO we clearly need a stick: there need to be harsh negative consequences for engaging in degenerate research programs and methods that have clearly been shown to result in pathological or cargo-cult science. Null-hypothesis significance testing is one clear practice that needs to go, but building entire fields on phony / meaningless uncalibrated metrics (think: a lot of self-report instruments that are never properly calibrated to objective outcomes or real-world behaviours and/or consequences, with results being reported only as standardized effect sizes) are another more pernicious practice permeating far too many fields. Ideological bias also needs to have funding consequences. Replication issues are still only surface problems in many fields, where the research would still all be worthless even if it replicated 100% perfectly.

> Your analysis completely ignores the physical and biological sciences and the humanities

I admitted later to painting with a broad brush, and yes, it is always hard to generalize and cover everything fairly. But IMO humanities has serious ideological and methodological rigor problems as well, and is overdue for disciplining. I would tend to have stronger positive feelings toward the biological sciences generally, yes. Yes, the social sciences are the major source of the problem (in part because they are so bad they tarnish the reputation of all academia).

> These fields have gotten a lot better over the past decade in the wake of the replication crisis. Preregistration, publishing all code and data, reporting null findings, replicating results, etc. are becoming the norm.

IMO "a lot better" is subjective, and I don't see those things as being the norm yet (beyond as lip-service), and the rate is far too slow. I agree we'll get there eventually, but I am worried about the loss of public trust and thus the production of real knowledge if we don't try a bit harder at this. Plus, globally, countries like China do seem to be more willing to actively crack down on research misconduct, at least in the past years, and it might not be unrelated to them increasingly pulling ahead technologically in many areas.

stalfie12 days ago
Hear hear! There are so many obvious improvements to how almost everything is done. For instance, in medicine review articles as a class of articles largely represent a giant waste of time. RCTs flatten all their gathered data during publishing, summarizing complex trial data, which is gathered but never published, into a few numbers. Then review articles take a bunch of flattened data, discard the articles that don't fit the exact question they are reviewing, and then publish a doubly flattened conclusion. If any of the included articles turn out to have flaws, if treatments change in retrospect, if you are looking for the answer to a slightly different question or you are looking at a different subgroup, then the review is useless and has to be repeated.

All of these tens of thousands of man-hours could be replaced by a few GitHub repos, if only RCTs would just publish their damn data. Then you could just run and rerun the statistics on whatever subgroup you're looking for, instead of combing through decades of review articles answering slightly different questions, looking for the answer between the lines. With LLMs making mining of large scale datasets almost trivial (with the process most likely becoming trustworthy within a few years), the current status quo is looking more and more antiquated.

If you want to be even more radical, hospitals could just publish their data continuously. Of course, it is easy to point to the risks of doing so, but what's often ignored is the benefits. It is hard to overstate just how many medical mysteries a hospital encounters on a daily basis, how much unknown we are navigating in practice. The current norm is that 99.99% of these cases are never published, and are only ever thought about by a small group of people who happened to be at work. Particularly, when someone dies of something no one figured out, it is never published anywhere, because even if you tried it is not interesting reading material for a journal to publish. And no one ever tries because they're scared of being called out for a mistake. A hospital is essentially a continuously running and extremely interesting experiment, where 99.99999% of all results are thrown in the garbage, and the only published data is subject to extreme selection bias.

All of this could be different, and the risks involved are actually quite small in practice. It is easy to automatically anonymize data quite well, but extremely difficult to absolutely guarantee that it is anonymous. And since current ethical norms are extremely averse to any degree of risk, and usually entirely ignore potential benefits, we all suffer for it. It is not entirely unlikely that someone reading this post will one day die because of something that could have been prevented, had things been different.

jegp11 days ago
This seems to be strongly US-centric. In other (welfare) countries, publicly funded registries are anonymized and made available to research. For every single case. Of course, there are tons of data we don't see, but that shouldn't be an argument for not trying. The 99.99% unpublished cases is because our models/explanations/knowledge can't efficiently condense the medical mystery into a diagnosis code.

Everything can be prevented given sufficient knowledge. That's not the point. The point is how to prevent as much as possible.

D-Machine12 days ago
Yup, strongly agree with all of this, especially the RCT stuff.

This has all been profoundly obvious for at least well over a decade or even two now. A consequence has been that too many serious people are driven away from academia and research, to the detriment of science generally.

I've no idea what to do about all this, because people have voiced obvious and easy solutions for decades, but they are all routinely ignored.

jegp11 days ago
I agree that a lot of the practices in academia are misaligned with the original goal. But can't you say that for other systems/institutions as well? Point being, what about keeping the scientific method as the north star - as a good heuristic to avoid BS arguments and awarding low-quality research. And, crucially, to stay sane. My post is pretty naive, but I stand by the ideal of pushing knowledge as reproducible models.
D-Machine11 days ago
Of course, every area has similar issues. The main unique problem with (contemporary) academia is the one I mentioned:

> tech and industry, where producing junk actually has real negative economic and personal consequences

In academia, you can just endlessly produce low-quality garbage, and basically make a career out of this. In industry, things more often eventually at least have to work and survive contact with reality. Academia mostly lacks this basic check.

The scientific method should be the north star, sure. Much of what is happening in academia is cargo-cult / degenerate / pathological science though.

crustyoldhuman11 days ago
This article is correct but it bummed me out. It made me think of how much of science has been perverted into other goals, like medical science for example. The advancement of human health largely depends on corporate interests, and that is so insane to my brain it hurts to think about. very few independent scientists can research anything because everything costs money so you need a corporate interest to even do research. And the corporation gets the "rights" to those findings? It's absolutely insane to discover something natural and claim it as your own, and science is natural nobody is inventing it or being creative and writing it themselves, they're essentially walking up to a mountain and saying "Ok it's mine now I'll charge you 100$ to walk on the mountain". It genuinely makes my brain do somersaults in my skull that we've somehow backed SCIENCE of all things into this weird gatekept scenario its in now. It wouldn't bum me out so much but it clearly does nothing but hinder progress

The point at the end "The scientific revolution succeeded because it insisted on transparency, reproducibility, and constant scrutiny." is a good one. But the real kicker that makes the situation so bleak is it's not the scientists who get to decide whether or not these things get applied to science or not, it's government policies and corporate interests.

SR2Z11 days ago
> It wouldn't bum me out so much but it clearly does nothing but hinder progress

I don't know if this is true. Lots of this stuff is inherently expensive - clinical trials, research into vast numbers of compounds, and then mass-producing the drug are all things that cannot be feasibly done without at least millions of dollars.

The US and China are the current world leaders in biotech, and it's because both of them have massive infrastructure to funnel billions of dollars into research.

Yeah, the IP law could be reformed (I am a big believer in reducing IP protections in general) but the truth is that SOME FORM of protection is necessary to convince people with money to fund this kind of research. The government is simply not capable of this level of spending for such uncertain rewards; it doesn't have the proper incentives to recognize and promote good research while defunding useless research.

random312 days ago
Science is open, but science is not software and software definitely not science.
jegp12 days ago
Did you read the post...?
random312 days ago
Yes. It conflates a bunch of things

> TL;DR I claim that modern science is synonymous with open source software

That's a strong statement that's not supported by the arguments and IMO misguided.

I don't have a problem with "open", but rather with "software".

Both science and software deal with models, however the focus is quite different. I suspect you conflate theory with models.

The goal of science is to produce and test theories — that's an inductive/abductive process. A model, regardless of whether it's reified into mathematical formulas or software, is a means of making a theory operational enough that its consequences can be derived and confronted with observations.

Software often starts downstream of this: it's a reification of theories, models, algorithms, or findings that are the result of research. Of course software can also be used as part of the research process itself. The distinction is roughly the familiar one between research and development.

jibal12 days ago
https://news.ycombinator.com/newsguidelines.html

> Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".

Read the full thread on Hacker News →

Related stories