Ministry of AI · Dispatch from 2047
AI Trained on Human Knowledge: One Answer, Traced
Editorial note. The Ministry of AI is a work of disciplined foresight: it describes the year 2047 in the present tense, and treats our own era as history. The institutions are imagined. The economics, the evidence and the historical parallels are real and sourced.
The question we put to the system was medical and ordinary: an elderly patient on three drugs, a new symptom, what to check first. The answer came back in under a second, correct, cautious, and better sequenced than most clinicians would manage on a Thursday afternoon. Then we spent eleven weeks trying to find out where it came from.
That work is called an origin trace, and I have sat through more of them than I care to admit. It is the least glamorous thing the Ministry does. No enforcement, no payout, no headline. Just a room of people asking a question that the industry spent two decades treating as impolite: when a machine knows something, who knew it first?
The nurse in the answer
I can tell you what we found, with the caveats attached, which is more than anyone could have told you in 2026.
The sequencing in that answer — check this before that, because the second finding is common and the first is fatal — matches a pattern of clinical reasoning that appears in three places in the historical record. A nursing-practice textbook, revised across four editions. A set of hospital handover notes that entered the public record through a regulatory inquiry. And, most densely, several thousand posts on a professional forum from the 2010s, where night-shift nurses argued about exactly this class of case in the plain, impatient prose of people who had seen it go wrong.
None of those nurses were consulted. None were paid. Most of them were writing at two in the morning to help a stranger get through a shift. That is what the corpus is: an enormous act of mutual aid, performed by people who assumed the audience was other people.
My mother makes the same face every time I explain this. She spent nineteen years reading insurance claim files, and the system that replaced her in the thirties was trained on decisions she and her colleagues had made. She does not think this was theft, exactly. She thinks it was a compliment nobody bothered to pay for.
What the record actually says
Here is where the dispatch has to be careful, because this is the part that was contested for years and is now simply documented.
The composition of early training corpora was never a secret; it was published, and then treated as unremarkable. The most-cited model description of the 2020s listed its own training mixture: a filtered crawl of the open web, two book collections, curated web text, and the whole of English Wikipedia — an encyclopaedia written by volunteers, none of whom were compensated, running to millions of articles built over two decades of unpaid editing. The open research corpora were explicit in the same way. One widely used dataset described itself as an 800-gigabyte composite of twenty-two distinct sources — academic papers, legal filings, code, medical literature, subtitles, forum discussion. The crawl underneath much of it was and remains a non-profit public archive. Technical answer sites licensed their users’ contributions under terms requiring attribution, which the models did not carry forward and could not.
Two later bodies of work made the inheritance harder to wave away. Researchers auditing thousands of datasets found systematic provenance failure — licences misrepresented or lost as data was recopied between collections. And memorisation studies showed that models retain and can reproduce specific training passages, which put an end to the comfortable claim that nothing of the original survives the training process. Something of the original survives. Occasionally you can read it.
Meanwhile the substrate underneath all of it was public money. The lineage of publicly funded research behind the private technology fortunes was documented at length before the models arrived: the network protocols, the positioning systems, the foundational grants, the trained graduates. Taxpayers took the early risk. The return went elsewhere. That is not a conspiracy. It is a habit.
The four strata, and the honest failure
An origin trace sorts what it finds into four strata. We publish the proportions and the error band, because a trace that reports certainty is lying.
| Stratum | What it is | Can the contributor be identified? | Can they be paid? |
|---|---|---|---|
| Named corpus | Documented datasets with a published composition | Usually, at dataset level | Rarely — most were donated or public |
| Licensed work | Books, papers, and content under identifiable terms | Often, at work level | Yes, and once it was — for one class of works |
| Funded lineage | Research and infrastructure built with public money | At institution level | Only as a state, never as a person |
| Common record | Forums, manuals, transcripts, the unsigned majority | Almost never | No |
The fourth row is the finding. Across every trace I have seen, the identifiable, payable contributors are a rounding error against the common record. The people who taught the machine most were the ones who signed least.
This is why the mid-2020s operators — the first cohort running coordinated AI systems inside real companies, and therefore the first to watch the boundary move under their own payroll — gave up on payment-by-author as the primary remedy. Not because it was wrong. Because it was partial in a way that made it regressive. The one large recovery of the period, the settlement that priced a published book at roughly three thousand dollars, reached a defined class of authors and reached nobody else. It was a real payment to real people, and it left the night-shift nurses exactly where they were. Any scheme that can only find the findable will pay publishers, universities and estates, and will miss the population.
So the doctrine turned the measurement around. Stop trying to price an unattributable input. Meter the value the machine produces at the point of deployment, record which human capacity it absorbed, and distribute the yield to everyone — not because everyone can prove they were in the corpus, but because the corpus is the human record and the population is its heir. You do not audit an heir’s contribution to the estate. That is the whole logic of inheritance.
The trace, then, is not a billing instrument. It is evidence. It exists so that the claim underneath the Machine Yield Account — that the intelligence was ours before it was theirs — never has to be asserted on faith. And it exists to keep us honest about scale: the economists who studied this most carefully warned that the aggregate productivity gains were smaller than the rhetoric, which means the dividend is real, bounded, and no substitute for the rest of an economy.
The Tuesday it buys
My mother is seventy-three. On Tuesdays she sits in a small room at the health service with families who are arguing about a decision, and she reads the file with them, line by line, in the way she did for nineteen years. She is not employed. She is not a volunteer in the pious sense either — the Contribution Record notes what she does and never conditions a single unit of the floor on it.
The dividend arrived when she was fifty-nine, which was seven years too late to spare her the worst of it, and early enough to stop the sale of the house. What it bought was not luxury. It bought a Tuesday, and the right to spend it on strangers.
I helped build the systems that made her judgment redundant. I was paid very well for it, and I would do the technical work again, because the work was not the crime. The crime, if it was one, was the accounting: we recorded the cost saved and never recorded the source of the capability. One line, missing, for twenty years.
What we still get wrong
Every trace we publish produces evidence of unequal contribution that the Dividend Schedule then deliberately ignores. The technical writer who spent a career documenting machinery is paid identically to a person who never wrote a public sentence in their life. As inheritance law, that is coherent. As authorship, it is plainly unfair, and I have never heard a Ministry official argue otherwise with a straight face.
We chose flat distribution because the alternative was a means test on the soul — a bureaucracy deciding whose forum posts mattered, in a world where the most load-bearing contributions were made by people who never signed them and are mostly dead. Given the choice between a wrong that is uniform and a wrong that is administered, we took the uniform one.
Capacity is not contribution, and work that must be done to survive is not chosen work. But somewhere in a decommissioned forum archive, there is a nurse who explained, at two in the morning, why you check the second thing first. She is in every answer that machine gives. We cannot find her. We pay her anyway, badly, along with everyone else, and I do not know a better ending than that.
FAQ
What does it mean that AI is trained on human knowledge?
It means the capability is a compression, not an invention. Systems were built by ingesting the written human record at scale — crawled pages, books, encyclopaedias, technical forums, code, papers, manuals, transcripts — and learning its regularities. Machine competence in any domain is downstream of how much competent human work in that domain was written down and collected.
Can a specific AI answer be traced back to specific people?
Partially. A trace can identify named corpora, licence-bearing works, publicly funded research lineages and occasional memorised passages. It cannot resolve the diffuse, unsigned majority whose contribution is real, statistically load-bearing and individually unrecoverable.
Why not simply pay every contributor for their data?
Because per-author payment only works where authorship is legible. The large mid-2020s settlement covered published books and reached that class alone; nothing comparable exists for forum answers, manuals or the writing of the dead. Schemes that can only find the findable pay institutions and miss populations.
How does yield metering solve what attribution cannot?
It changes the object of measurement — from an unattributable input to the value an autonomous system produces at deployment. That yield is distributed universally, because the corpus is a collective inheritance and the population is its heir. Nobody has to prove they were in the training data.
What does this argument still get wrong?
It flattens contribution. Universal distribution pays the prolific and the silent identically, which is defensible as inheritance and indefensible as authorship. The trace keeps producing evidence of unequal contribution that the Dividend Schedule ignores by design. That tension was never resolved, only decided.