The Record Is the Asset

The Sovereign Stack · Part Five
AI has made permission something machines can execute. Whether African makers see the money now turns on who writes the terms in a form a machine will honour.
August 30, 2026

“Get the right person in the room and they can tell you exactly where a record lives.”

Adaobi Orajiaku is describing Nigeria’s National Gallery of Art. Its documentation department maintains printed and handwritten catalogues recording how each work entered the national collection.

Orajiaku’s company, Atsur, has spent the past two years working with the Gallery to rebuild its ownership records. “It runs on filing systems, and often on the memory and relationships of the people who work there,” she told Guzangs. “Without them, the same information is much harder to reach.”

Inside the Gallery, that system works. Beyond it, the same knowledge becomes difficult to use. A buyer doing real diligence, Orajiaku says, “is not transacting on trust in an institution’s name. They are asking what the record is built on. Is there a document, a photograph, a bill of sale, something that exists independent of a person’s memory?”

Provenance that lives in recollection “is not something you can produce if a claim is challenged, and it will not survive being tested across two or three different legal jurisdictions,” she says. “That is the real test a record has to pass, and it gets harder, not easier, once an AI agent is the one deciding whether to transact.”

An AI agent cannot call the archivist who remembers which file completes the history. It needs an authority it can verify and terms it can parse.

Guinea has built one version of such a record for léppi. In 2025, the African Intellectual Property Organization, OAPI, registered Léppi de Guinée as the country’s first protected geographical indication for an artisanal textile. The registration names the Groupement Représentatif du Textile Léppi de Guinée as its holder and provides a basis for policing what may legitimately be sold under the protected name. Origin, production and authenticity are defined.

Its precision ends there. A geographical indication does not determine who owns the copyright in a photograph of the cloth, who controls a particular pattern or who may train a model on images of it. There is an authority for who may sell the cloth as the cloth. There is no equivalent record connecting its digital representations to the people entitled to authorise their use.

Type léppi de Guinée into an image generator and that limit appears immediately. The model can reach for indigo, cotton and the geometries associated with the cloth. It cannot show which part of what it learned came from a particular dyer or weaver, which patterns belong to a family rather than an individual maker, or who had the authority to permit a commercial imitation. It treats as a prompt what Guinea’s producers and legal institutions have spent years learning to treat as a protected name.

The record exists. The machine cannot read the part that matters.

Now follow one photograph of one cloth, because the chain is short and the blanks are the story. A weaver finishes a léppi and someone photographs it. Was she asked what would happen next? The photograph reaches a product page, a documentation project or a cousin’s feed. Asked? A crawler collects it. Asked? A dataset absorbs it. Asked? A model learns the relationship between the word léppi and those geometries. Asked? A generator produces a derivative for anyone who types three words, and somewhere a print-on-demand storefront sells the result.

In this common version of the chain, no right is asserted in a form a later system can reliably detect and act on. The person at the beginning is never asked because no field travels with the image in which her answer can be recorded.

Six steps from a woven léppi to a sold AI-generated derivative; at each step, the question “Asked?” goes unanswered because no machine-readable permission record travels with the image.

 

Rights that machines cannot read become difficult to enforce at machine speed.

Line the pieces up. Guinea holds a legal record: a regional registration defining the authentic cloth and naming the producer body that holds the designation. Nigeria holds an institutional record: catalogues and memory, authoritative inside the building and illegible outside it. The AI economy has produced a market willing to pay for documented, licensable material.

Three working systems. Nothing connects them.

The permission field is missing

A necessary precision: a language dataset, a living artist’s painting, a museum object and a communal textile tradition are legally and economically different assets. The rights over them differ profoundly. A record does not create a right where the law recognises none. It makes an existing right, claim of authority or community protocol visible enough for another system to act on. A machine cannot negotiate a right it cannot identify.

Ownership, provenance and consent are routinely collapsed into one. The confusion is convenient for the people doing the collapsing. Ownership asks who holds a legal right in a work. Provenance asks where the material came from. Consent asks what the person with authority actually allowed.

A photograph can carry immaculate, cryptographically verified provenance and still say nothing about whether the textile inside it may be reproduced. The web blurred these distinctions for decades because the user’s principal action often appeared to be viewing or sharing. Systems that ingest, transform, generate and sell have made the underlying rights impossible to ignore.

Machine-readable instructions are already part of AI regulation. On 2 August 2026, the transparency duties in Article 50 of the European Union’s AI Act began to apply. Covered providers of generative systems must mark AI-generated content in machine-readable form, although providers of systems placed on the market before that date have until 2 December 2026 to comply.

C2PA is one of the standards positioned to carry provenance at scale. It is used in some cameras and major platforms and is under consideration in ISO’s standardisation process. Its steering committee includes some of the world’s largest technology and media companies, but no African organisation at that level.

A file can now carry, tamper-evidently, the statement “this image was generated by AI,” and several major platforms can read it. C2PA shows that machine-readable assertions can travel with content. It does not determine who may grant permission or make the instruction enforceable. That depends on who is entitled to write the terms and who agrees to honour them.

The international system is being lapped by the market, although the scope of its work matters. In May 2024, after a quarter-century of negotiations pushed hard by African delegations, WIPO members adopted a treaty on intellectual property, genetic resources and associated traditional knowledge. It does not cover textile designs or traditional cultural expressions generally, so it does not settle entitlement for léppi. Even within its narrower field, it needs 15 ratifications or accessions to enter into force. As of late August 2026, it has four: Malawi, Uganda, Albania and Peru.

The AI market has already created prices for data and creative content while public institutions have barely begun to decide how older forms of community authority will be recognised in comparable digital markets.

Scarcity is not power

What the market pays, and whom, is now a matter of record. News Corp licenses journalism to OpenAI in a deal reported at more than a quarter of a billion dollars over five years. Reddit disclosed $203 million in data-licensing arrangements with terms of two to three years in its IPO filing. Shutterstock reported $203.3 million in 2025 revenue from a segment that includes AI-training data alongside Giphy, Studios and other services. Warner settled with Suno and Udio, while Universal settled with Udio. Parts of the labels’ copyright fight have become a licensing market.

Why do those sellers get paid? Their content is valuable, but value alone is not enough. They also possess something much of Africa’s scarce cultural and linguistic material has not yet been given: aggregation, documentation, attribution, contractual control and terms a buyer can negotiate.

The saleable asset is permissioned data: the object plus its provenance, authority and terms. At machine speed, the record becomes part of the asset. Without it, cultural scarcity does not become economic scarcity. Scarcity has no price until the market can tell who controls it.

Chris Emezue has watched this from an unusual vantage point. The Nigerian AI researcher, based in Montréal, spent six years with Masakhane, the grassroots African language technology collective, before founding Lanfrica, a platform that catalogues the continent’s language resources. This year, his team worked with GSMA to audit African language datasets. The finding overturned one of the field’s starting assumptions.

“Just because you can find something doesn’t mean you can use it,” he says. “We thought that if a thing can be found, we improve its utility. But then we realised, no, there’s another blocker.”

The blocker is the licence, the small legal ledger that tells a stranger what a shared artefact may be used for. The Data Provenance Initiative’s audit of more than 1,800 training datasets found licence information missing from roughly 70 per cent of listings across major hosting platforms. The same research found that low-resource-language datasets had the least commercially usable coverage. Asked whether African datasets fare better, Emezue is direct: “It’s more or less the same.”

Every large, funded African dataset his team examined carried a licence, usually because a funder required one. The omissions sit in the long tail of community-built work, created by the people the paying market is least likely to find. This is how inequality is formalised through documentation: the funded get papered; the grassroots get skipped.

Even the presence of a licence does not mean the people represented in the dataset chose it. “It might require scrutiny into how the licence was chosen and whose interest it is communicating,” Emezue says.

Chris Emezue, founder of Lanfrica, presenting in Montréal in front of a slide about organising information.
Chris Emezue, founder of Lanfrica, presenting in Montréal. The slide behind him carries the argument: organising information is what makes it useful. Photograph courtesy of Chris Emezue.

The cost of a missing term is not hypothetical. Masakhane’s first landmark achievement, a sweeping machine-translation effort across dozens of African languages, was built partly on JW300, a widely used corpus derived from Jehovah’s Witness translations. A legal review found that the source site’s terms prohibited text and data mining. Permission was requested and declined, and the corpus was withdrawn. The models already built did not disappear, but the shared substrate for reproducing and extending them did.

“Decades of work that people could have done in building machine translation for African languages,” Emezue says of the work that source could have supported. “That’s all gone.”

A disclosure, because this essay ought to practise what it argues. An earlier draft carried a statistic placing Africa, Latin America and the Middle East at 9.1 per cent of the training-data market. Traced to its source before publication, the figure came from a commercial market report with an unpublished methodology. Its regional categories contained no Africa at all; the continent was bundled with the Middle East in a single segment worth 3.9 per cent.

The number could not say what it was being used to say, so it is not in this essay. That is the record layer working as it should. A claim whose provenance cannot be established is a claim a careful buyer cannot use. The rule does not stop applying because the claim is convenient.

The worst of both worlds

Here is where this essay expected to land, and where its best source refused to follow.

The argument seemed clean. In a market that pays, unreadable rights stop being a shield and become a tax. The responsible buyer leaves, and the culture is abandoned.

Emezue stopped the thesis mid-sentence. “There’s a presumption you’re making, that if it’s unreadable, they abandon it,” he said. “ChatGPT has taught us that just because it’s unreadable doesn’t mean it’s abandoned.”

The generative boom was built on the opposite premise: extract everything, scale fast, let the lawyers absorb what follows. The aftershock has been a retreat from permissive sharing. The internet stopped feeling like a safe place to share openly once extraction at scale became the norm.

A missing licence offers no safety. It sorts the buyers. The buyer who cares about permission leaves; the actor who does not care takes. Unreadability selects for the least accountable buyer, while the person at the source may not know which of the two reached the work.

“If it’s unreadable, not only can your work not be found,” Emezue says. “Your work can be extracted and stolen and you don’t even know, and there’s nothing you can do about it.”

Readable rights make either path more answerable. They let an honest buyer identify whom to approach, and they give a creator with an underlying legal claim evidence with which to fight a dishonest one. Without them, there is less prospect of a deal and less chance of an effective defence.

There is recourse at the extractive end, but it is priced for the organised. The $1.5 billion Anthropic settlement with authors, finally approved in July, concerned the company’s acquisition and retention of pirated books; the court had separately held that the act of training itself could be fair use. The settlement does not establish that every unlicensed training use must be paid. It does show what claimants with identifiable works and documented rights can recover when an underlying infringement is established.

Organised paper pays. Spotify negotiates with labels and publishers that aggregate music rights. Getty aggregates image rights. Reddit aggregates access to, and contractual control over, a vast body of human conversation, which is why it can sign data-licensing agreements with AI companies. 

What aggregates African cultural data at a comparable, machine-legible scale? Very little. “If you don’t have the bargaining chip, you’re basically just begging,” Emezue says. “If everything is so scattered, they can pick it one by one. It’s very hard to see the power when everything is scattered.” Scarcity alone does not create bargaining power. Organised scarcity does.

The terms have to travel

An agent that can book a hotel still cannot reliably license a cloth. Culture needs machine-readable, machine-actionable rights information. A statement buried in legislation or a handwritten catalogue may be clear to a human reader; an agent needs terms it can parse and route to the right authority.

The record has to say who controls the asset, whether it may be licensed, for what purpose, for how long and at what price. It must identify who gets paid, whether permission can be revoked and how that revocation travels downstream.

Some of the components already exist, several of them African-authored. The most quietly radical is NOODL, the Nwulite Obodo Open Data Licence, developed by Chijioke Okorie and Melissa Omino through a collaboration involving the University of Pretoria’s Data Science Law Lab, Strathmore University’s Centre for Intellectual Property and Information Technology Law, and other African dataset researchers. Its premise is that uniform open licences can reproduce inequity unless benefit-sharing and local interests are made explicit.

“We could have said, let’s use a Creative Commons licence, because that’s how it’s done,” Emezue says. “But those licences weren’t created for datasets. We don’t always have to wait for the West to bring out the solution for us.”

The same ethic runs through datasets built on what he calls data farming rather than data mining: cultivate, harvest and give back, instead of extracting and leaving the ground barren. NaijaVoices, the consent-built speech corpus he works on, was collected that way and released under CC BY-NC-SA 4.0, with separate terms for commercial use. The recent Yoruba-English corpus from LyngualLabs carries NOODL into the world’s data infrastructure.

Local Contexts offers evidence that community authority can be made machine-readable. The non-profit, incorporated under Navajo Nation law, created Traditional Knowledge Labels that let communities attach their own protocols and expectations to cultural material. A Label may indicate non-commercial use, community use only, cultural sensitivity or openness to commercialisation. It does not grant permission; it announces that the community expects to be part of the negotiation.

Each customised Label carries a permanent, machine-readable identifier, and an API can propagate changes downstream. The Labels are extra-legal, educational mechanisms, not licences. They do not manufacture rights that a legal system does not recognise. They make community protocols and approved uses visible inside digital systems: a yes, a no and a call us first, all parseable even when legal force must come from elsewhere.

That leaves an institutional problem. Communities need bodies capable of issuing and maintaining the instructions and, where enforceable rights exist, acting on them before the defaults harden.

Who owns the yes

At Atsur’s inspection table in Lagos, the work begins with the object. The artwork is physically examined and documented. The seller’s identity is verified against government ID. The resulting certificate, Orajiaku says, “carries forward a verified starting point, not a guess.”

The difficult cases begin where the paper ends. Living artists can attest for themselves. An estate must establish who has standing to speak before anything is recorded. For an older work with no written history, Orajiaku refuses to pretend.

“We cannot manufacture a provenance that was never written down,” she says. “What we can do is start the clock from the moment the work enters our system, and be honest that anything before that point is documentation, not proof.”

Her records do not yet carry the executable layer, and she names it plainly. “A person can read a certificate and infer what it means to buy or license the work. An agent cannot infer. It needs the terms expressed in a form it can parse directly: price, scope, expiry, and proof that the party granting the licence actually holds the rights.”

Then she draws the divide and the roadmap: “Provenance tells you what a thing is and where it has been. Licensing infrastructure has to tell a machine what it is allowed to do with it. The provenance layer had to come first. The licensing layer is next.”

The people building this architecture still have to decide who may answer. Who issues the authoritative record for a particular léppi image or pattern: the photographer, the dyer, a cooperative, the producer body behind the geographical indication, the state, or several together? If two people claim the ability to say yes, which one does the system trust? If consent changes, who ensures the change propagates?

Custody carries its own warning. ANKA provided commerce infrastructure to more than 20,000 African sellers. In October 2025, after the judicial liquidation of its French parent, it was acquired by New York-based Global Shop. The deal raised an immediate question about one of the continent’s largest structured stores of independent seller and product data: who controlled it now, and under what terms?

Public reporting established the change in ownership, not the precise disposition of the seller records. That uncertainty matters. If creators and institutions do not hold custody of their structured records, the models of the future may license structured representations of the continent from intermediaries abroad. The yes may be sold by someone other than the people to whom it belongs.

Unauthorised extraction is one danger. A second is hiding inside the compliant licensing market now taking shape. It may become very good at routing permission requests and paying rights holders while much of Africa remains invisible because no authoritative record tells a buyer who may answer.

The future could be more lawful and still extractive. Companies may pay whoever holds the cleanest record. The winners will not necessarily be those with the most valuable culture. They may be those with the most legible ownership.

The next fight over cultural ownership will not be over whether a machine recognises where something came from. It will be over whether that origin carries enforceable instructions with it.

A sovereign dataset does not merely know what something is. It knows who can say yes.

The Sovereign Stack is a continuing Guzangs research series on the data engines, machine-learning pipelines, and transactional architectures shaping the future of the African and diaspora creative economy.

Part One: Who Trains the Machines That See Africa?
Part Two: The Machine Can See You. It Still Can’t Pay You.
Part Three: The Index Is the Institution
Part Four: The Machine Has Started Shopping

Subscribe

This is what we publish. Every week.

Original reporting on the designers, institutions, and economies defining African creativity, delivered to your inbox.

A weekly letter on African fashion, art, and design, and the people making the work. Reported properly, and worth your time.

By signing up, I agree to the Terms of Use (including the dispute resolution procedures) and have reviewed the Privacy Notice.