The Blandification of Science was a Mistake
This is an essay version of some of the ideas from my 2026 Scipy Talk: 'FAIRer Data: The case for data advertising in the era of Agentic AI'.
At some point during my PhD, I came across this article: Research Papers Used to Have Style. What Happened?. Judging by the date it was published, I think it was around the time I was struggling to get my first paper through peer review, and I was both pretty fustrated and disappointed with the process.
I remember being pretty irritated that the copy writer had proposed a bunch of edits to my writing, which I recall either a. reverting changes reviewers had insisted on, or b. changing the language of the paper in a way I didn't find particularly appealing. I wasn't confident enough to push back after the ordeal of peer review - so I just accepted all the edits and got on with my life. Apparently though, I never really forgot about it, because here I am, four years later, still grumbling about the experience.
I've never really been one for scientific prose. I generally tend to write things for myself up in loose, freewheeling, stream of thought prose. Personally, this tends to be a lot more readable for me. The linked article gave me something to latch onto: 'What if the problem wasn't me? What if it was everybody else!?!'.
Always one to be humble, here I am asserting that I am right, and everybody else is wrong...
N.B: I'm extremely fond of listening to podcasts involving Rory Sutherland and Stephen Wolfram. Lots of the ideas presented here are theirs - I've just latched onto them to justify my grumpiness with the world not immediately recognising that I'm obviously the greatest genius to ever live (sarcasm hopefully apparent)
A complete history of the entire scientific publication process (abridged)
Lifted verbatim from the linked article, because I can't state it any better:
It’s not hard to find scientists who think proper scientific writing should be devoid of style (I was a teacher for 6 years, and IMRAD remains the de facto style taught in schools). But there has been an undercurrent of dissenting voices dating back to the very origins of scientific publications. In 1661, four years before the first issue of the Philosophical Transactions of the Royal Society, Robert Boyle wrote:
“And yet I approve not that dull and insipid way of writing, which is practiced by many...for though a philosopher need not be solicitous that his style should delight his reader with his floridness, yet I think he may very well be allowed to take a care that it disgust not his reader by its flatness…”
... (Reordered chronologically)
In 1667, Thomas Sprat urged members of the Royal Society to “reject all the amplifications, digressions, and swellings of style; to return back to the primitive purity, and shortness, when men delivered so many things, almost in an equal number of words.” Some 200 years later, Charles Darwin said much the same: “I think too much pains cannot be taken in making the style transparently clear and throwing eloquence to the dogs” (Aaronson, 1977).
The advocates of the new science in the seventeenth century so reacted against the excesses of stylistic artistry that a reluctance to use any artistry at all seems to have prevailed ever since.
— Whitburn et al., 1978
// Aside 1
- At the risk of being needlessy divisive, I think this argument has become even more pronounced since the rise of LLM's (November 2022 onwards). Go onto any reddit thread, facebook post, etc., and you'll find (often founded) accusations of posts being bland, content devoid, insipid AI slop. This is totally fair - LLM's have ground up the total corpus of human text, and when prompted to write, produce something akin to a statistical mean of that. It's unsurprising that their writing is somewhat uninspired.
// Aside 2.
- The arguments I'm making about writing and/or content delivery here are not true for code. Bland, AI assisted code is a good thing, in my mind. Anything that makes code more immediately interprable is net positive, to my mind. You'll frequently see people bemoaning the youtube algorithm for driving extremism/division - the jogging to running to marathons to David Goggins pipeline. This kind of algorithmic force makes peoples' interests easier to predict. In contrast, writing clear, well documented code makes LLM assistance more effective. Thus, AI autocomplete trains you to write easier to understand code.
Anyway, getting back to the point at hand, the problem is that for some reason, scientists decided that making their ideas as unpalatable as possible was the maximally honest way to behave. That is:
- Good ideas should stand on their own two feet
- Advertising your idea and/or making it palatable gives legs to an idea that shuold never have had them.
- Therefore, science which was written to be palatable or enjoyable to read should be suspect, because the only reason for which the author might have to do so is in order to hoodwink unsuspecting marks into believing their rubbish ideas.
It therefore became a mark of authenticity to make your scientific writing as bland and unappealing as possible.
It is really important to note that this culture did not emerge in a vacuum. From my extremely cursory research, it seems Sprat was writing in response to perceptions that the rhetoric bandied about during the English civil war had stoked divisions, and ruined the country for what seemed like all of living memory. Sprat's argument invoked a return to standards of objectivity and 'properness', yet was in of itself in some sense marketing. Not without irony, it seems.
I also want to note the 'Englishness' of this argument. England has a strange culture, where the nouvelle riche (New Rich) are looked down at for their propensity for buying nice things. It's considered gauche to have earned your own money and therefore ostentatious to display wealth. Somewhat weirdly, in England, the working and lower-middle classes aspire to own a nice car as a status symbol. Amongst the upper-middle and upper classes, owning a shit car is the status symbol.
Also, nouvelle riche and gauche? Peppering your language with random words inherited from French is the 'real' marker of sophistication there, presumably due to the legacy of the Norman conquest. See also beef/cow, pork/pig, poulty/chicken. These are all the legacy of the norman/saxon divide. There's an excellent Tom Scott video on this.
// Aside 3.
- When I went to sea at the start of 2020, Brian King taught me a lot about the history of oceanography. The global oceans thermohaline circulation was first discovered empirically in I think 1751 by a chap called Captain Henry Ellis, who lowered a modified bucket down more than about 1600 metres in the middle of the Subtropical North Atlantic. He discovered the water was surprisingly cool, and remarked how wonderful it was that they had a mechanism for cooling their wine and baths. What a lovely (and memorable) passage!
I discovered, by a small thermometer of Fahrenheit’s, made by Mr. Bird, which went down in it, that the cold increased regularly, in proportion to the depths, till it descended to 3900 feet: from whence the mercury in the thermometer came up at 53 degrees; and tho' I afterwards sunk it to the depth of 534 feet, that is a mile and 66 feet, it came up no lower. The warmth of the water upon the surface, and that of the air, was at that time by the thermometer 84 degrees. I doubt not but that the water was a degree or two colder, when it enter'd the bucket, at the greatest depth, but in coming up had acquired some warmth for I found, that the water, which came up in the bucket, having stood 43 minutes in the air (the time of winding it up) the mercury rose above 5 degrees. When the air had. render’d it equally warm with the water on the surface, I tried their weight, by weighing equal quantities very exactly, as also by the hydrometer, and found from great depths the heaviest, and consequently the saltest water.
This experiment, which seem'd at first but mere food for curiosity, became in the interim very useful to us. By its means we supplied our cold bath, and cooled our wines or water at pleasure; which is vastly agreeable to us in this burning climate.
Anyway, the key point here that I want to make, is that the peculiar quirk of English culture which drives a suspicion of showmanship has become deeply embedded throughout science, in a way that is deeply pernicious. I don't think most people even see it. But science has a deep aversion, almost an allergy, to promotional content.
You might occassionally hear the aphorism 'Never tell a story without making a point, and never make a point without telling a story'. Sprats philosophy somehow missed all but the first four words - and science has suffered ever since. And somehow, it's so all encompassing, we struggle to even notice it, wondering why we seem unable to win trivial debates with climate change 'skeptics', young earth creationists, and flat eathers.
There are two young fish swimming along, and they happen to meet an older fish swimming the other way, who nods at them and says, "Morning, boys. How's the water?" And the two young fish swim on for a bit, and then eventually one of them looks over at the other and goes, "What the hell is water? - David Foster Wallace
David Ogilvy
In the interests of brevity, I'm going to avoid giving such a flowery description of David Ogilvy. Suffice to say, however, he is often referred to as 'The Father of Advertising'. Ogilvy took a somewhat different view to Sprat - perhaps unsurprising.
In 1957, Ogilvy launced this ad campaign, for Rolls-Royce, in America:
"At 60 miles an hour the loudest noise in this new Rolls-Royce comes from the electric clock".
This headline was really remarkable and groundbreaking - not for what it said, but what it implied. Elegance. Sophistication. Refinement.
Even more remarkably, it was totally true. Ogivly didn't even write it himself - he apparently spent three weeks reading anything he could get his hands on about the car, and then lifted the headline - a quote from a technical report. What Ogilvy noticed wasn't even, at it's core, at odds with what Sprat advocated for. It was that you could present the facts - even for something as boring as the acoustic performance of a car - in way that moves people emotionally.
I once heard someone (I think Sam Harris) talking about AI and whether artificial consciousness would ever be possible - what is it that makes us conscious. He pointed out there was nothing magical about the fact that we're all made of meat, so why couldn't we build a conscious artificial brain? To my mind, the mistake of post-enlightenment Britain - and this feeds all the way through to perfectly logical, but totally unworkable ideas such as communism - was to try to pretend that we aren't made of meat.
We aren't disembodied intellects, roaming the cosmos in search of truth and reason. We have insticts, desires, emotions. Some things make us angry, others happy, others bored. Sprat completely missed this point, but Ogivly clocked it.
// (Gratuitious aside) So did the Star Trek writers:
- 'Sell the sizzle, not the steak' - Ferenghi Rule of Acquisition #153
- 'You can't free a fish from water' - Ferenghi Rule of Acquisition #217
Squaring this all up
To be perfectly honest, most of what I know about David Ogilvy, I've learnt from listening to Rory Sutherland. So I'm going to switch tacks a little bit, and take it from that angle, in the hopes it will read better - it's certainly easier for me to write.
If you spend as much time as I do listening to him, you'll be familiar with this quote:
"Nothing kills a bad product faster than good advertising"
Remember the Rabbit r1? In early 2024, it was one of the most hyped AI products. It sold itself as being the badge from Star Trek - touch it, ask the AI for something, get what you want. They raised $20 million in funding. The thing was a total flop - just google it.
What's the point here? It's that if you have a good product, and you advertise it, then lots of people will use it. The market will reward you for making something useful. If you have a bad product, and you advertise it, lots of people will buy it, be upset that they've wasted their money, and the market will punish you. If you make a product that is genuinely useful, but has lots of rough edges, there's a good chance you'll get plenty of useful user feedback. Hopefully, you'll be able to iterate, and refine the product - until you reach success.
Crucially, if hardly anyone uses the product, it will remain unpolished forever.
What this really boils down to is that consumers aren't stupid - but you can't just build it and expect them to come. Unless you're really famous, if you build something useful, you have an obligation to promote it if you want anyone to use it. How else will anyone find out if it's any good otherwise?
Ogilvy put it way better than I could:
"The consumer isn't a moron, she's your wife"
This is all well and good, but what the hell does it have to do with FAIR data?
FAIR data: Findable, Accessible, Interoperable, Reusable.
I've been working at ACCESS-NRI for roughly the past two years, in the MED (Model Evaluation & Diagnostics) Team. Predominately, I work on our data cataloguing utilities - tools that make it easier for scientists to discover and use data. When I joined, our flagship product in this space (really, it still is!) was the ACCESS-NRI Intake Catalog.
This tool is an intake plugin that makes it really easy to find data - provided you meet a couple of preconditions, which I'll get back to. When I was a PhD student, the process for getting your hands on model output data looked something like:
- Find out by chance that someone has a model output that might be useful to you.
- Find the right person in the building.
- Ask them where the files you need are on the HPC system.
- Do a lot of a faffing around & writing scripts/globs etc. to access them
- Do your science.
Steps 1-4 have very little scientific value, but could take weeks. 1 alone could be a total work killer, stopping ideas from ever getting off the ground. The ACCESS-NRI Intake Catalog lets you do this in seconds:
import intake
cat = intake.cat.access_nri
experiment_name = cat.search(frequency='1mon', realm='ocean',variable='tos') # For example
esm_datastore = cat[experiment_name]
xr_ds = esm_datastore.search(variable=['tos','sos'], frequency='1mon').to_dask()
# Analysis time
When I first started working on something, I thought it was magic. It solved all the problems I wish could have been solved for me during my PhD. In a lot of ways, I still do - none of the criticisms that I'm about to level at it contradict or mitigate how much of a step forward from what we had previously it was and is. Shout out to Dougie Squire for building such an impressive tool! But there were a few issues and limitations that needed addressing:
- Scientists had to learn how to interact with a new software package. This is a barrier to adoption. I call this the 'Hire Car Problem'.
- It's walled into an HPC system.
- HPC groups can make the whole thing behave a bit weirdly.
Of all of these, 3 is the easiest to solve. You can infer group issues from the filesystem. 1 & 2 are a bit tougher
The Hire Car Problem
I coined this term after describing a variant to a friend over some beers one friday evening. He gave me a much more elegant version that occured to him.
As Darcy described it to me, he had flown to Melbourne for the weekend, and hired a car. As he drove out of the airport car park, he touched the brakes, and was thrown forwards, almost out of his seat. Good brakes! So what did he do? Parked the car back up, and went back to the car hire desk. "You've got to give me a car with worse brakes! I'm only here for the weekend, I don't have time to get used to this!"
Fundamentally, the Hire Car Problem is that better does not necessarily feel better - if it forces you to update your intutition. I gradually came to realise that despite me been so keen on what was very clearly a better solution to me, lots of scientists preferred to do things the way they were used to. We needed a spoonful of sugar to make the medicine go down.
Being walled into an HPC System
HPC is the embodiment of Sprat's ethic, calcified into a multi million dollar infrastructure project. To use our intake catalog, users would first need to:
- Acquire an NCI login. Easily done for an Australian Academic or Government worker - provided they know that it's easy, and they know there is something they want on there.
- Join the correct projects on Gadi. To manage resources, HPC systems will put resources into 'projects' - basically, a linux group with storage/compute resources.
- Open the intake catalog, and search for the data they need. Only data from the projects they've joined will show up.
Do you see the fundamental issue here? Our intake catalog is in some sense a mechanism for discovering data - abstracting away details like 'experiment X is in project Y under path /g/data/Y/X/...'. Yet, because of the way HPC file access is structured, if they don't know it's under Y, they won't be able to find it via intake. So to discover data, they already need to know (approximately) where it is.
Issue 1 is even more of a barrier. What if you don't know that you're entitled to an NCI login via your work (I've seen this). What if you don't even know that NCI/Gadi exists? There is signposting, but it's all on the NCI documentation, which is very ... functional. The breadcrumbing here is non-existent.
Going back to the FAIR acroynm.
I'm going to assert without any evidence that our intake catalog satifies the AIR part. But what about Findable? Notionally, sure - it's open source, you can get a Gadi login and join the correct projects, and then search it pretty trivially (perhaps with a little help from the documentation).
In practice? Not so much. I've met heaps of people who were surprised this existed at all, and when I showed them, the reactions were generally 'oh my god, this is so useful, I wish I'd known about it sooner'. To my mind, this mean we're failing the findable test.
I want to be super clear here - I'm not taking aim at our intake catalog because there's anything specifically bad about it. This pattern is really commonplace throughout academia. I'm just using it as an example, because it's something I'm super familiar with.
Rory Sutherland (can you tell I like him?) has an even better way of putting this: he asserts that technical people (I think in 'Alchemy' he takes aim at either economists or the Finance Department at this point) often think:
'Sure it works in practice - but does it work in theory'
and worse, don't bother to ask whether it works in practice if it works in theory. Of course, in theory, there's no difference between theory and practice, but in practice, there is.
Our intake catalog was falling victim to this - findable in theory, not in practice.
When we think a bit more about these sorts of 'findable in theory' datasets, I want to also include another point: the competition for attention is getting fiercer and fiercer, and attention spans concomitantly shorter. We live in a world that has tiktok, where social media companies copy mechanisms from slot machines to capture our attention. In science, we see these practices and instinctively turn away from them, reasoning that these are bad faith actions and so we should move in the opposite direction, ever further towards Sprat's ideal.
Sure, this might be ideologically pure and philosophically reasonable, but it makes us even less competitive for attention. Should we be putting slot machine dynamics into our data tooling? Probably not - although the variable reward schedule of agentic coding does give us something of a compelling case for it - but we should at least recognise that we now live in an attention landscape that looks a lot like the modern food landscape of McDonalds, Pizza and Coca Cola.
We need to make our tools more palatable, not less.
Data Advertising
Hopefully, the point I'm trying to make here is now starting to emerge - that if we want to build tools that people actually use, we need to make them truly findable (and accessible, which follows).
Findability does not mean the principle of chucking your data onto an FTP server and hoping that someone who needs it eventually figures out how to download it. The real world equivalent would be turning your living room into a supermarket, leaving your front door unlocked, and hoping somebody wanders in one day to buy something.
If you want something to be findable, you need to do three things:
- Build a great product, which people will actually want to use.
- Make it openly accessible.
- Advertise it.
For our intake catalog, this looked like the following:
- Build the ACCESS-NRI Intake Catalog
- Mirror the relevant parts to openly accessible object storage, accessible over the web, not just via a credentialed HPC login.
- Create an interactive web interaction. Breadcrumb users - give them popup dialogues, hover info, code generation, interactive filtering, fuzzy searching, quick links to relevant documentation - and overall, an intuitive, easy to use experience.
Note how 'thin' steps two and three are - they're not about lying to the user, misleading them about capabilities, or trying to sell them some junk they don't need. They're about creating a funnel from discovery to use, and not putting both on the wrong side of a login wall.
Data advertising is about putting in the work to make your data findable - so that it's quality can speak for itself, and actually be heard.
I'm going to leave the rest of the talk for a separate post, because that's about practical examples - really, this is a story with two parts.