Lecture Eight. The Criticisms That Landed. This is the eighth of ten talks about ontologies. For seven lectures I've been explaining how this technology works and how to build with it, and last time I left you with a governance body that happens to produce a file. Today I'm going to make the case against the whole enterprise, as strongly as I can, using three critics who were right about different things. I've put this lecture here deliberately. If you hear the criticism before you understand the machinery, it sounds like grumbling. Now that you know what a reasoner does and what a governance body costs, the criticism has something concrete to bite. Three critics, three layers. Doctorow attacks the incentives. Shirky attacks the range of cases where the approach applies. Bowker and Star attack the consequences. In my judgement the first two have been substantially answered and the third has not. Start with Cory Doctorow. His essay is called "Metacrap," subtitled "putting the torch to seven straw-men of the meta-utopia," first drafted in May two thousand and one and circulated in August — months after the Scientific American article. That essay is routinely filed as the reply to that article. It isn't. Doctorow never mentions the Scientific American piece, never mentions the Semantic Web, and never mentions Berners-Lee. His stated target is broader, and he says so in his opening: explicit, human-generated metadata, he writes, has enjoyed recent trendiness, especially in the world of X-M-L, the Extensible Markup Language. Read it as an attack on hand-declared metadata generally and it lands hard. His thesis in one sentence: a world of exhaustive, reliable metadata would be a utopia, and it is also, and I quote, "a pipe-dream, founded on self-delusion, nerd hubris and hysterically inflated market opportunities." He gives what he calls at least seven insurmountable obstacles, and I'll give you all seven, because the list is better than its reputation. One. People lie. Metadata exists in a competitive world where suppliers compete for wallets and attention, and when poisoning the well benefits the poisoner, the water gets toxic fast. His examples are of their time — a search for a common term turning up a porn link in the first ten results, spam with subject lines reading "Re: the information you requested," press releases with gargantuan lists of empty buzzwords — but the mechanism is permanent. Anyone who has looked at search engine optimisation has watched this happen in real time. Two. People are lazy. Your clueless aunt sends email with no subject line. Half the pages on Geocities are called "please title this page." Your boss stores files on his desktop called UNTITLED dot DOC. His test case is downloading ten random music files and finding that at least one has no title, artist, or track information — despite the ripping software having a button that fetches all of it in one click. This laziness, he says, is bottomless, and no amount of ease-of-use will end it. Three. People are stupid, by which he means careless, and his example is the best in the essay. On eBay every seller has a strong financial incentive to spell a listing correctly, because a misspelled listing never appears in a correctly-spelled search and therefore attracts fewer bids and a lower price. Search eBay for "plam," p-l-a-m, and you would find nine typoed listings for Plam Pilots. You can almost always get a bargain on a Plam Pilot. Money on the table, and people still get it wrong. Four, and this is his own title: Mission Impossible — know thyself. In the meta-utopia everyone weighs their own stuff and accurately reports its properties. When Nielsen used log-books to gather the viewing habits of sample families, the results skewed heavily towards Masterpiece Theater and Sesame Street. Replacing the log-books with set-top boxes that reported what the set was actually tuned to showed what the average American family was really watching. People, in his phrase, are lousy observers of their own behaviours, and asking them to describe those behaviours in a form is asking for fiction. Five. Schemas aren't neutral. This is where the essay turns from mockery to argument, and it's the part that has aged best. A manufacturer of small, environmentally conscious washing machines draws a hierarchy that leads with energy consumption, water consumption, and size. A manufacturer of glitzy, feature-laden machines leads with colour, size, and programmability. Both are honest. Neither is neutral. Any hierarchy of ideas asserts that some axes matter more than others, and the belief that competing interests will come to easy accord on a common vocabulary ignores the power of organising principles in a marketplace. Six. Metrics influence results. A common yardstick privileges whatever scores high on it, regardless of overall suitability. Intelligence tests privilege people who are good at intelligence tests. His sharpest example: Nielsen couldn't generate ratings for three-minute mini-programmes, so MTV couldn't demonstrate the value of advertising on its network, so MTV stopped showing videos. A measurement system reached back and changed the thing it was measuring. Seven. There's more than one way to describe something. "No, I'm not watching cartoons, it's cultural anthropology." "This isn't smut, it's art." Reasonable people can disagree forever about how to describe the same object, and requiring everyone to use one vocabulary, in his phrase, denudes the cognitive landscape. Now, what does Doctorow actually conclude? Not that metadata is worthless. He says so explicitly. He endorses implicit metadata, and his example is Google, which at the time derived the reputation of a page from the number and quality of links pointing at it — data produced as a by-product of what people did, not as a declaration of what they claimed. That sort of observational metadata, he writes, is far more reliable than the stuff human beings create for the purpose of having their documents found. Observed signal against declared signal is the durable core of the essay. My own gloss on why: the first four obstacles are all arguments against asking a person to declare something, and none of them has much purchase on a signal you simply record. So use obstacles one to four as an incentive audit, two questions long. Who enters this data? And what do they gain from entering it well, compared with entering it badly? If the answer to the second is "nothing" or "less," your ontology will fill with garbage no matter how good the model is. Use obstacles five to seven as a design review: name the axes your model privileges, and say out loud whose interests they serve. That split has become more useful since he wrote it, not less, because language models cut the list in half. They change the economics of obstacles two, three, and four by extracting structure nobody wanted to type — laziness, carelessness, and being a poor witness to your own behaviour are all costs of making a person fill in a field. They do nothing at all about one, five, six, and seven, because those are about incentives and power rather than effort. So the half of the list that no interface improvement could fix is the same half no extraction can fix. That last conclusion is my inference from the shape of the list, not his claim; he was writing in two thousand and one. Now Clay Shirky, and the essay "Ontology Is Overrated," a heavily edited concatenation of two talks he gave in two thousand and five — one at the O'Reilly ETech conference in March, one at IMCExpo in April. Shirky is more often cited than read, and what he actually says is narrower and better than the version that circulates. His own wording is that the strategy is both widely used and badly overrated in terms of its value in the digital world — and "in the digital world" is carrying as much weight there as "overrated." His target is not classification. It's ontological classification specifically, which he defines as organising entities by their essences and possible relations, and — this is the key part — designing categories in advance to cover cases that haven't arrived yet. A library catalogue assumes that a new book's logical place already exists in the system before the book does. He starts by conceding the best case. The periodic table gets his vote for best classification ever. Organise elements by proton count and you get enormous descriptive and predictive value, and because you're organising actual things, it's about as close to essence as physical reality permits. And then he finds the flaw, which I love. Look at the noble gases. Helium is no more essentially a gas than mercury is essentially a liquid. Helium is a gas at most temperatures, and the chemists who classified it couldn't get it cold enough to see otherwise. So they took a property that holds at room temperature, absolutely unrelated to essence, and put it at the centre of the category. And it's stayed there ever since, a frozen accident everyone has got used to. His point: if a near-perfect scheme over physical essence still contains context errors, what chance does a domain with no essence at all have? Then the libraries. Take the religion block of the Dewey Decimal Classification, the D-D-C — the two hundreds. Nine subdivisions in total: natural theology, Bible, Christian theology, Christian moral and devotional theology, Christian orders and local church, Christian social theology, Christian church history, Christian sects and denominations, and then two-ninety, "other religions." Eight of the nine are Christian or Christian-adjacent, and everything else on earth is in the ninth. That one's easy to laugh off as period bias. The Library of Congress example is harder, and it's the one that does the work. The L-C's top-level categories under History run D-A Great Britain, D-B Austria, D-C France, D-D Germany, D-F Greece, D-G Italy — and then, presented as co-equal with those, D-R the Balkan Peninsula, D-S Asia, and D-T Africa. These are all top-level categories, all presented as co-equal. The Balkan Peninsula and Asia sit at the same level. Shirky asks what's being optimised, and rules out the obvious answers. It isn't geography. It isn't population. It isn't regional economic output. The Library of Congress has a staff of people who do nothing but think about categorisation all day long. What's being optimised is the number of books on the shelf. And then the line that carries the whole essay: the essence of a book isn't the ideas it contains. The essence of a book is "book." A book can be about several things at once, but the physical bound object can only be in one place, so it has to be declared to be about one main thing, regardless of its actual contents. Which brings the punchline. In the digital world, there is no shelf. The physical constraint that produced hierarchical classification is simply gone. And here's his parable. Yahoo, faced with organising the web with no physical constraints at all, hired a professional ontologist and built a hierarchy. Go to their Entertainment category and you'll find Books and Literature marked with an at-sign, which means the category isn't really there. It's a convenience link. To see those links you have to go where they really are, under Humanities. And when you get to Literature, booksellers aren't really there either, because they're a commercial service, so they're really in Business. Given the freedom to put anything anywhere, Yahoo added the shelf back. They couldn't imagine organisation without it, so they reinstated the constraint and then told users which location was real. Now, Shirky's argument is regularly over-read into "ontologies are useless," and he does not say that. But be careful with the part people quote back as his checklist, because it's weaker than the version they repeat. He does not offer a test. What he offers, in his own words, is a partial list of characteristics that help make it work — nine of them, in two groups. On the domain: a small corpus, formal categories, stable entities, restricted entities, clear edges. On the participants: expert catalogers, an authoritative source of judgment, coordinated users, expert users. The more of those characteristics that are true, he says, the better a fit ontology is likely to be. That's a description, not a decision procedure. Inverting it into a go/no-go screen for a project is my framing, and I'll go on using it, but it isn't his. His worked positive examples are two: the periodic table, and the D-S-M-Four, the fourth edition of the psychiatrists' Diagnostic and Statistical Manual, where the American Psychiatric Association is the authority that says what symptoms add up to what, and cataloguers and users are both expert. That people infrastructure, he says, is a big part of what makes it work, and it's expensive. Turn the list over and you get the negative version — large corpus, no formal categories, unstable entities, no clear edges, uncoordinated and amateur users — which, as he observes, is an almost perfect description of the Web. There's a whole half of that essay I've left out, and leaving it out is what turns him into a demolition man. The full title is "Ontology Is Overrated: Categories, Links, and Tags," and the links and the tags are a proposal. His way in is a question about merging. Suppose you merged your personal library into the Library of Congress. Do you and the Librarian first have to sit down and reconcile your scheme with theirs? Of course not. They take your books and ignore your categories, because every book carries an I-S-B-N, an International Standard Book Number. The merge happens at the level of the globally unique item, not the category. The presence of unique labels, he says, means that merging libraries doesn't require merging classification schemes. And on the web everything already carries such a label, because a web address is exactly that — so anybody can hang labels on those pointers without anybody first agreeing on a scheme. That is a tag. Tags, in his phrase, are important mainly for what they leave out: forgoing formal classification is precisely what lets them produce organisational value at vanishingly small cost. He quotes the man who built the social bookmarking service delicious — each individual categorisation scheme is worth less than a professional one, but there are many, many more of them. He also has the reply to the obvious objection, which is that free-form labels need a thesaurus to tidy them up — that if you say "movies" and I say "film" and somebody else says "cinema," we should all be collapsed together. His answer is that those words encode different things, and that with genuinely contested vocabulary all the signal loss is in the collapse, not in the expansion. Which is Doctorow's seventh obstacle arriving from the other direction, as a design principle rather than a complaint. One last piece of his, because it hands the next critic their whole argument. Why do we know a sport utility vehicle is a light truck rather than a car? Because the government says it is. Shirky calls that voodoo categorisation, where acting on the model changes the world: when the government says it's a truck, it is a truck, by definition. He raises it as a limit — most of the world, he says, is not amenable to voodoo, and nobody can settle whether Buffy the Vampire Slayer is science fiction because there's no authority and no force to apply. Read it in the other direction, which is my move and not his: where the authority and the force do exist, naming really does change the thing named. Which is the third critique, and the one I think is unanswered. Geoffrey Bowker and Susan Leigh Star, "Sorting Things Out: Classification and Its Consequences," published in nineteen ninety-nine. It's not a technical book. It's an ethnographic study, and in the words of Terrence Brooks, reviewing it, the effect of its examples is to illustrate how values, policies and modes of practice become embedded in large information systems and become expressed in classification systems. I owe you a warning about how I know this, and it's an awkward one to give in a lecture about hidden provenance. I have not read the book. What follows comes from Brooks's review, from Stefan Helmreich's review essay, and from Eric Nehrlich's reading notes, so when I quote, I am quoting them quoting it. Take the next few minutes as second-hand, and go to the primary text before you build on it. Brooks's verdict on what the book achieves is worth having up front: it is the accomplishment of this book, he writes, to recognise classification itself as an object of study, as a vehicle for ethnography. Their first case is the International Classification of Diseases, the I-C-D, which sounds like the most neutral document imaginable. A list of diseases. Developing it took many years and there are still many disagreements over it. Tropical countries believe tropical diseases are grossly under-represented compared with rich-world diseases like cancer and heart disease. And there's a detail I find unforgettable: in Japan a heart attack is considered a low-status way of dying, so death certificates will often list a stroke as the cause instead, which skews the results when they're compared with other nations. Their central concept is torque. Stefan Helmreich's review essay gives the definition: the process that unfolds when the time of the body and of its multiple identities cannot be aligned with the time of the classification system. Individual biographies, he writes, are twisted into tortured shapes where the scheme and everyday life don't line up. Note both halves of that. A person's own account of themselves fails to align with the system, and it is the person who gets twisted. The system wins. Their extreme cases make it unmistakable. Tuberculosis patients, wholly dependent on a doctor's diagnosis for their own status. And race classification under apartheid in South Africa, where citizens were tossed back and forth between being classified as White, then Coloured, then back again, and where a government declaration of your race forced you to change residence, job, and family. Helmreich walks one life through that machinery, and the detail is the argument. A South African boy, born of an Indian father and an African mother. At birth he is classified Asian, because a rule under the Group Areas Act has children living with the father's side. Then he reaches the age of majority, and a different statute takes over — the Population Registration Act, under which he is expected to follow the station of the parent of lower racial status. So he becomes African, and is expected to shift his associations accordingly. Nothing about him has changed. Two laws disagreed about which parent counted, and the thing that moved was the person. Helmreich adds that the torsion could go further still: conceivably this man could later have tried, with some difficulty, to pass as Coloured, by cultivating work and residence associations with the appropriate people and avoiding the street-level classificatory challenge that might come from a police officer. A life spent managing a category. I read that as a category acting on a person rather than describing one — but that reading is my gloss, not theirs. Then they do something cleverer, which is to show the same mechanism in an ordinary case with no villain. Nurses were asked to describe what they do so that it could be classified, standardised, and put into a billing system, and the result was the Nursing Interventions Classification, the N-I-C. Some nurses cheered that they would now be in the system. Others were aghast that they would no longer be free to do what they thought was right. Both reactions were correct. What the case makes evident is the tension between the system and the local adaptations you have to make for the system to function at all, and that tension runs through the whole book. Which gives the companion lesson, the one I'd write on the wall: there will always be elements that don't slot neatly into a category, and that is a fact about the inadequacy of the system, not about the element. Their example is the platypus. Their last point is about invisibility. Classification systems tend to get black-boxed, in the sense of Bruno Latour — that's his term, and it's the right one. They become infrastructure, and infrastructure is by definition the thing you stop noticing. We forget how much effort went into creating them, and the political and ethical issues bound up in their creation. What's left looks like a description of how things simply are. Put the cases and the black-boxing together and you get the book's actual argument, which is not the one people expect from a critique. It is not that classification is avoidable. It is that classification is consequential and usually invisible, so the only real choice available to you is whether the consequences get examined. Now, why do I say this critique is unanswered? Because everything the technical field built answers a different question. Consistency checking, profile validation, alignment, provenance tracking — all of it measures the model against the world. Is the model coherent, does it fit the data, does it agree with that other model. Torque runs the other way. It's the world being measured against the model, and a person losing. There's no reasoner for that. There's no lint rule. The only mechanism that addresses it is a human process: someone in the room asking, when a real case doesn't fit a category, who bears the cost. Helmreich puts the sharp end of it in a line — there is no experience of torque for those in power. So if the cost lands on someone who isn't in the room, the modelling decision is a policy decision, and it should be made by people with the authority to make policy. The book's fair criticism comes from Brooks, the reviewer I quoted earlier. He objects that the examples pile on without advancing the argument, that the next step nobody has taken is establishing which classification characteristics are associated with which human, linguistic, or institutional ones, and that the book is strangely disconnected from the large body of empirical, cognitive research on classification that already exists. All fair. Take the case studies as demonstrations, not as a predictive theory. Before I sum up, the defence, because I've given you three critics and no reply, and that isn't a fair hearing. Two years before the essay I just walked you through, in November two thousand and three, Shirky wrote a companion piece arguing that the Semantic Web is a machine for creating syllogisms, that it will therefore improve all the areas of your life where you currently use syllogisms, which is to say almost nowhere, and that it requires too much coordination and too much energy to effect in the real world. In two thousand and fifteen Dave McComb of the firm Semantic Arts published a direct rebuttal, and his central move is a joke at Shirky's expense: that entire article, he points out, is itself a syllogism. Making it a premise of your argument that a thing will fail because nobody argues in that style is an awkward place to be standing. But notice what McComb does not claim. He does not claim the critics have been answered. His own sentence is that we still have a long way to go to staunch the critics, which is an unusually honest thing for a defender to write. So let me say what I think each critique actually establishes, since that's the thing a listener needs. Doctorow establishes that declared metadata is unreliable in proportion to the incentive to misdeclare, and that no interface improvement fixes it. What he does not establish is that curated metadata fails. The successes the retrospectives themselves name — schema dot org, knowledge graphs, Wikidata, DBpedia, the biomedical ontologies — are governed, funded, bounded domains where somebody has an incentive to curate, and a curator on a funded biomedical ontology is not your clueless aunt filling in a form. His obstacles are about hand-annotating the open web, and the field moved off the open web: from annotating it by hand, to publishing linked datasets, to modelling a single enterprise. The retrospectives draw a sequencing lesson out of that, and it is about incentives rather than semantics. A committee that specifies before anybody ships produces artifacts nobody adopts. Schema dot org worked partly because the search engines gave publishers a direct reason to comply: it was started by Google, Bing and Yahoo with the express purpose of delivering better search results. Then look at what its team are careful to state on their website — that they are not attempting to create a universal ontology. So the vocabulary that actually got adopted answers Doctorow by giving the people doing the declaring something they want, and answers Shirky by declining, in writing, to attempt the thing he says can't be done. My reading is that those are one decision rather than two. Shirky establishes that designing categories in advance is the wrong tool for a large, ill-defined corpus with uncoordinated, amateur users. What he does not establish, and does not claim, is that it's the wrong tool everywhere. The periodic table and the D-S-M-Four both concede the point. Bowker and Star establish that categories act on the people they classify, and that this is invisible by design. What they do not offer is a way to tell in advance which categories will torque, which is exactly what an engineer would want. Their contribution is to make you look, and looking turns out to be most of it. So, the claim to keep from this lecture. Three critiques, three different layers. Incentives, applicability, and consequences. The incentive argument has an answer: prefer observed signal over declared signal, and audit who benefits from lying. The applicability argument has an answer: Shirky's characteristics, honestly applied before you start. The third, the consequences, is the one the field still has not answered. There is no technical reply to torque, only the discipline of asking, out loud and in the room, who pays when a real case doesn't fit. Next time, the company that took this word, kept about half of it, and built the most commercially successful thing anyone has ever called an ontology.