Conjectures

Content provenance, attribution, and quality in times of AI

How institutions and social interactions should change in light of AI

This essay looks at how society, our institutions, and individual attitudes should evolve to address the increasing diffusion of AI. Because this truly may be a profoundly important technology — and as a cautious optimist about its prospects, I want to engage in a critical discussion that will allow us to capture the best of its promise while avoiding the worst.

What prompts this question? The diffusion of AI throughout the world may deeply change our world — our relationships and the dynamics therein. As a professional software developer, I ponder the changes AI will bring. Not simply in speeding up routine workflows, but in potentially obviating the need for a large number of professionals. So there is a personal economic motivation behind this concern.

Second is the related, but broader, concern: of all the activities that were performed (exclusively) by humans, which will still be performed by humans? This question is both about the loss of economic value experienced by supplanted professionals, and about the almost-spiritual loss of meaning and purpose.

A typical reaction to the unsettling breadth of these changes is to pre-emptively reject them. This is a very safe and predictable human response — to prefer the status quo — and it still beats yearning for the golden yesteryears. But ultimately this broad-brush notion conflates the many benefits and risks. These risks vary in severity, and in turn can be fixed. But we must approach them precisely.

How we approach these questions is grounded in how we project AI will diffuse throughout society and the economy, and the impact it will have. Profoundly transformative technology always incites a broad range of reactions. On the edges are utopian and dystopian narratives that hold AI to be a near-supernatural force. But such prophetic notions do not generally translate into critical discussions. Critical discussions are what actually allow us to capture the good and avoid the bad. (This relates to Venkatesh Rao's notion of epic times as inescapable flows in time; the dystopian and utopian narratives become a kind of escape, which may be useful in its allegorical value, but must become more critical if it is to helpfully serve the trajectory of AI.)

Concerns

Concerns and criticism are where that critical discussion begins. Each objection locates a seam between the technology and the institutions it is arriving into; a problem, in the Popperian sense, to be resolved. The objections are not all of one kind. Some are about what AI produces, some about what it displaces, and some about what it gives bad actors leverage to do. They do not all resolve the same way, and it is worth taking them separately.

Nothing interesting is produced

The first criticism is that there is nothing interesting produced by AI. Not in the sense that nothing interesting is produced — which is trivially false — but rather that what is produced is simply a reflection of what was in the AI's training regime, so nothing truly novel is produced.

I think at some level this is how humans work, and arguably how all knowledge grows. Of course the volume of data AI trains on is vast; but that is a matter of degree and not of kind. Synthesis is genuinely useful, and arguably a lot of what humans do.

And even if we hold that nothing new is being generated, the act of having explored a part of the extremely large manifold — through the human's understanding, intuition and beliefs — is still valuable. This is what expositors and science popularisers do; the value is in the selection and the route taken through the material, not in the novelty of the material itself.

This criticism is often really about something else: the injustice whereby human art and output is used to train these models, and these humans have not been compensated for their work. That should be addressed directly, via the existing protections that owners of content have. No one should support AI developed via illegal means — i.e. by stealing books, or by not attributing them. But some of this is fair use, and arguably rightfully so. Where existing property rights do not cover these cases, they must evolve; and laws must evolve with them.

The human contributed nothing

The second concern is that all the interestingness is in the AI, and that the human contributed nothing.

This criticism arises from a confusion about the production process. Thinking and writing are not monolithic processes, but rather a collection of co-operating ones. Writing breaks down into: grand ideation, finding supporting arguments, finding evidence for those arguments, constructing a holistic flow, drafting, transitions, word choice. These are separate skills and competencies. And finally, the amount of time you can spend on each will translate into better output.

The grand idea is the core of the work. So if one uses AI for that, it is truly worth reflecting on in what sense the human authored the content. However, if AI was used for an auxiliary task — to develop an argument, or to find examples — it is not exactly the case that the human did not author the text; but nor did they fully author it either. It is one thing to express the blub version of an idea, and quite another to fully develop it. And we already have ways of generalising this: primary and secondary researchers on a paper. Where we land on the distinction will depend on the domain, but in general we will need a richer vocabulary to deal with it.

There is also a case on the other side. If AI helps us better express ourselves — to overcome some of the gaps in our skills — this is overall a good thing. Someone who has something worth saying and cannot yet say it well is not served by a norm that treats assistance as contamination.

I think the real, or perhaps an important but adjacent, psychological force here is how we think about human authorship in near-metaphysical terms. That there is something unique about humans and our product, and that this unique humanness must be preserved. I have sympathy for that. But sympathy does not tell us where the uniqueness lives, and locating it in authorship in particular is what gets us into trouble.

Perhaps the critique raises that something is lost — the small edit, the word choice, the transition you didn't make. Is control over authorship lost? Agreed, in part; and this points to the last ingredient: time. Authorship was partly constituted by the time a human spent iterating on each component, and that is precisely what AI removes.

Domains and norms

How much any of this matters comes down to how much we value AI output in different endeavours. In some areas — creative endeavours, music, art — humans may prefer human output; i.e. for things grounded in preference at the outset. But these positions are untenable in mathematics, engineering and pharmacology. We will not reject theorems, proofs, systems or compounds because AI produced them. We may reject visual art and music. But even for subjective and aesthetic domains there is a kind of objectivity to music — whether it sounds good. There is a measure of goodness that is inherent in the music, and it is independent of its authorship. So I suspect that, unless users explicitly choose to filter for a label such as human-produced, the quality of AI output will likely be equally preferable. Some users will choose to filter and some won't; and like a classical liberal, I would not take this as a moral choice, but as different operators making different choices.

There is also the question of how social norms should evolve. Once people have got out their most dramatic reactions to AI, there can be a normalisation. And then groups of norms will emerge.

In some, the exchange of ideas and prose is all that matters, independent of authorship. In professional mathematics, engineering and pharmaceuticals, whether I wrote something or AI wrote it will not matter, because the route to the solution does not matter. The product is evaluated in an "objective" environment — it can be grounded in some part of reality.

So we will likely have a rich ecosystem of norms. And depending on the broader context, the same activity may be acceptably performed with varying degrees of assistance from AI. Professional mathematics may be solely concerned with results — but even there the notion of attribution may change. And in the context of mathematics tests and examinations, access to calculators, let alone AI, may be restricted. The activity is the same; what the institution is measuring is not.

But what of writing, which lacks this objective measure? There too we will find different communities with varying degrees of AI acceptance. In many, getting your point and perspective across will be paramount — though these will still have to deal with the quality problem. In others, such as journalism, usage may be acceptable but require proper attribution.

Mechanisms

Norms only bind if something makes them checkable. So it is worth asking what the mechanisms actually are, and what each of them can and cannot do.

Attribution is the first, and it needs to become finer-grained. Self-attribution already carries weight: if someone says they authored a book and it turns out she had it ghost written, this is a problem today. We will need some generalisation of property rights, attribution and provenance. Attribution must cover what parts were driven by the human and what parts by the machine — the idea, the argument, the evidence, the prose. Not as a binary label on the finished work, but as a description of the division of labour. We do this already for papers; we can generalise it.

The second is identity. Here we have one mechanism, which is the Facebook model — you attach the human identity to the machine identity, or to the generated content. And such a medium will have inherent rate limits; generally, behaviour that is spammy in nature will have to be checked. A human identity is a scarce thing: you get one, and it accrues a history. So binding output to it imposes a cost on volume that a machine identity does not carry on its own. This tells you nothing about how something was made, only who stands behind it. But in many cases that is the more useful question.

The third is provenance. Again, we have ways of tracking this. Maybe users self-attribute, or you have a deep bibliography, or you show the record of how you engaged with the work. But I think at a fundamental level we will have to move away from this notion of human-authored content, because it will be very hard to distinguish the output of a hardworking adversary who wants to fool you. Provenance can show that a tool was used; it can never show that one wasn't.

Which leaves verification, and the quality side. Ultimately we will need to become more peer-to-peer. Again, this is not a new mechanism — peer review, code review, reputation among people who actually know each other's work. But these were built in a world where producing something plausible was expensive. They will have to carry far more weight now that it isn't.

Deception, and the quality of the commons

On the negative side there are two flavours. One is the breakout scenario: this is truly a new failure mode. The other is all the leverage that a deceptive human being gets to propagate deceit — spamming, phishing, misinformation campaigns.

The second flavour is something we have thought a lot about, and we have reasonably simple ways of addressing it, albeit social adoption may be slow. The breakout scenario is unique in that we have no prior institutional machinery for it at all; it is also the one about which I have the least useful to say here.

A third criticism is that AI is speeding up the enshittification of the internet. Again, this is a problem of quality — and the honest question is how we deal with low quality now. The answer is largely provenance and trust: we route around bad sources, and we rely on reputations built up over time. Those are the same mechanisms as above, which is why they matter. The difficulty is that they were calibrated to a world where producing plausible-looking content was itself a filter.

To simply say that we will stop using AI at large is not tenable. But in limited capacities, where trust can be strong, we may create spaces which are free of AI — so that we can focus on a limited activity where, for a multitude of reasons, that is useful. Examinations are one such space. There will be others.

The loss of human ability

Lastly is the concern about the loss of human ability — to think, to review, to judge — and the degradation of the human institutions built around human production.

There is a feeling amongst some that AI is inherently destructive. My take is that, like all tools, it will enhance some of our natural abilities, and in doing so will let the underlying capability soften. This is the standard trade-off of tool use: you lose some acuity in the specific skill, and you gain in capability or in time. What is different here is the breadth of what is being offloaded.

What are these natural capabilities? At the broadest level, all kinds of "active" thinking — when you consume some medium and must digest it. Books and written text are the typical medium. But this is a spectrum and not a binary. At the other end of the spectrum is short-form video; and yet even there, a short, incisive video can make you mentally spin. So it is less about individual messages than about the broader practice of how we engage with a medium — and that observation applies to AI as much as to video.

Perhaps, then, it is worth asking what we will relinquish control over. At an implicit level, anything we do not understand is already beyond our control. Hence mission- and safety-critical systems must be understood: even if the machines make all the actual code and config changes, the human must sign off. This requirement will hold in safety-critical systems, though perhaps for only a small fraction of people.

But what about when our livelihood or our lives are not at stake? Should we give up thinking? No. Examination is itself a good — it is how we come to hold a view rather than merely carry one. And it is the capability everything else in this essay quietly depends on. The norms, the attribution, the peer-to-peer verification all assume a population that can still tell good work from bad. If that erodes, the mechanisms erode with it. One answer is deliberate practice: periodically working without the tool, sampling the scenarios, in the way one does chaos testing on a system. But the fuller answer belongs to a separate essay on effective usage.

Closing

Running underneath most of these concerns is a reach for essentialism — a wish that there be some property of human production which AI cannot touch, so that the question can be settled once rather than negotiated case by case. I think that desire is itself part of what needs to be addressed. It is not, as such, a criticism of AI. It is a call to evaluate our institutions and norms in the face of a transformative technology: to say what attribution should record, what identity should carry, what provenance can and cannot show, and what verification will demand of us.

Those are answerable questions. The essentialist one is not.