In April 2024, Canada announced more than C$2.4 billion in targeted AI support. C$2 billion was allocated to launch an AI Compute Access Fund and the Canadian Sovereign Compute Strategy. That was the right move. It is also, on its own, not sovereignty.

Here is the uncomfortable sentence: a model is its training data. Compute is rented. Data is destiny. If Canadian models train on corpora licensed from American aggregators, negotiated under American contracts and on American terms, then “sovereign AI” risks becoming a data centre with a flag on it.

The window where Canada’s data is still Canada’s to direct is open right now.

The land grab is happening now.

Frontier labs are signing training-data deals at a pace that would have been difficult to imagine three years ago. News archives, forums, code repositories and reference works are being locked up, territory by territory, into multi-year arrangements.

Every Canadian corpus that gets signed into a foreign exclusive is data that may never train a Canadian model. This is not a future problem. The window where Canada’s data is still Canada’s to direct is open right now, and it is closing with every deal signed somewhere else.

What Canada actually holds.

Canada’s data assets are better than the conversation suggests:

  • French/English bilingual corpora — genuinely scarce at global scale, and exactly the kind of data multilingual models need.
  • Legal and civic data — CanLII and its provincial siblings are among the best-structured public legal corpora in the world.
  • Health data — fragmented across provinces, which everyone treats as a problem. It is also the opportunity: whoever assembles it with proper consent and governance holds something no lab can scrape.
  • Resource-industry operational data — forestry, mining and energy. Decades of records from industries where Canada is genuinely world-class.
  • Indigenous data — subject to OCAP principles: Ownership, Control, Access and Possession. Sovereignty here means respecting that governance, not extracting around it. Any Canadian data strategy that ignores this is not serious.

What Canada does not have.

A supply layer.

In the United States, an industry is forming around the unglamorous work of finding corpora, clearing rights, packaging them for model builders and moving them at speed. Canada has no equivalent whose full-time job is sourcing Canadian corpora, doing the rights diligence and keeping them available to Canadian builders.

The compute got billions. The data got a conversation.

The missing infrastructure

Discovery → rights → provenance → access

A sovereign data layer does not mean putting every record in a government warehouse. It means knowing what exists, who governs it, what permissions apply, how it can be used and how a Canadian builder can access it without starting from zero.

Three things have to happen.

  1. The labs need a Canadian data partner. Cohere is the obvious anchor: a Canadian frontier lab with enterprise and sovereign positioning, and a genuine need for differentiated corpora. The role does not exist yet on anyone’s org chart. It should.
  2. Funding has to treat data infrastructure like compute infrastructure. A fraction of the Sovereign Compute Strategy directed at corpus discovery, rights clearance and licensing could do more for actual sovereignty than another data centre. Data centres without Canadian data are just cold buildings.
  3. Corpus owners need to know what they are sitting on. Every industry association, every operator with twenty years of records, every institution with an archive: these are strategic assets now. Most owners still price them at zero. The first people to tell them otherwise, with real buyer demand behind it, will assemble the supply.

Act now.

The compute is funded. The talent is here. The data is leaving—deal by deal, exclusive by exclusive, to buyers who may not give it back.

Canada spent $2 billion making sure the machines can be on Canadian soil. It is time to make sure what runs on them is Canadian too.

Source note

The C$2.4 billion and C$2 billion figures refer to the Government of Canada’s Budget 2024 AI announcements. Read the official Budget 2024 backgrounder ↗

Views expressed here are Archive Signal’s editorial position, not a claim that a dataset or licensing program currently exists.