Home › Data for AI › Third-party data
🗄️ · Build

Third-party data

Licensed or purchased datasets, where the contract decides what you may build more than the content does.

In one line

With third-party data the licence is the product constraint, and the clause that bites is almost always the one about derived works.

ConceptWhat it is

Third-party data is anything obtained from outside the organisation under terms — a purchased dataset, a licensed reference corpus, an open dataset with conditions. It is often the fastest way to acquire coverage nobody internally has.

The technical integration is usually the easy part. What decides whether it can be used is the licence, and the clauses that matter are rarely about reading. They are about redistribution, about derived works, and about whether an embedding of the data counts as one.

How it worksThe mechanics

Before ingestion, the grant is read and recorded as structured fields rather than as a remembered impression: what may be used, for what purpose, for how long, and what happens on termination. That record travels with the corpus and is checked at the point of use, the same discipline consent requires.

Attribution and segregation follow from it. Where a licence requires credit, the obligation has to reach the surface a reader sees, which means it is a rendering concern and not only a metadata one. Where terms differ across sources, keeping them separable in the index is what makes it possible to remove one later without rebuilding everything.

At a glanceSee it

Third-party data diagram

The grant becomes structured fields, not a remembered impression. Separability is what lets one source be removed later without rebuilding the index.

When to use itWhere it fits

  • When coverage you need genuinely does not exist internally and building it is slower than buying it.
  • Reference data — codes, taxonomies, standards — where an authoritative external source is the point.
  • Benchmark and eval datasets, where a shared public set makes results comparable to other people's.
  • When a permissive open licence makes the terms question short and durable.

When NOT to use itLimits & anti-patterns

  • Where the licence forbids derived works and an embedding is plausibly one, unless that has been resolved explicitly.
  • When termination would require rebuilding the index and no plan exists for it.
  • For a demo or kit corpus, where a self-authored dataset avoids the negotiation entirely.
  • When the same coverage exists internally and the purchase is buying convenience at the cost of a new obligation.

Trade-offsAdvantages & costs

Advantages
  • Immediate coverage of a domain that would take a long time to assemble.
  • Often higher quality and better curated than anything assembled in a hurry internally.
  • A clear licence is a much stronger position than scraped content with unexamined terms.
  • Public benchmark sets make your results comparable to published work.
Trade-offs & costs
  • The licence constrains the product, and the constraint is discovered late if it is not read early.
  • Termination is an architectural event, because the data is in an index and possibly in weights.
  • Recurring cost that scales with usage or seats rather than with value delivered.
  • Attribution obligations reach the user interface, so they are a design requirement and not only a legal one.

ExampleIn the real world

A classification kit licenses an industry taxonomy to label incoming documents. The agreement permits internal use and prohibits redistribution. Publishing the eval results is fine; publishing the labelled corpus alongside them, which is what makes a kit reproducible, is not. The kit shipped with a self-authored corpus instead, and that decision was cheaper made at the start than after the index existed.

ToolsHow to implement it

  • A structured licence record per sourcegrant, purpose, term and termination as fields, checked at use.
  • Source segregation in the indexa metadata field that makes removing one provider a filter rather than a rebuild.
  • Hugging Face Datasetsfor open benchmark sets where the licence is stated and machine-readable.
  • A self-authored corpusMIT, fixed seed, byte-identical on rebuild, and no negotiation at all.

Cost & effortWhat it takes

Licence fees are the visible cost and are frequently the largest single line in a corpus budget. The hidden cost is the termination plan: data in an index and possibly in weights means an exit is engineering work, not a cancelled subscription. Read the derived-works clause before the first ingest.

What changedWhat changed here

RecentAuto-linked from the brief, not a rewrite of this page
  • Amazon blocks Meta AI agent from shopping on its platform 23 Sep · US frontier labs

    Amazon blocked Meta's AI agent Muse from shopping on its platform, an early concrete case of a retailer refusing agent traffic. If you're designing an agent that acts on third-party sites, assume platform-level blocking is a real failure mode and plan for it in your architecture.

  • Meta’s AI agent has been blocked from using Amazon.com 21 Sep · TechCrunch AI

    Amazon blocked Meta's Muse agent from shopping on its site, citing an unauthorized AI agent violating its Conditions of Use. If you're building agents that act on third-party sites on a user's behalf, platform terms — not just technical capability — are the binding constraint on what your agent can do.

  • Your AI agents can now control your Google Home devices 16 Sep · TechCrunch AI

    Google launched early access to an MCP server for Google Home, letting AI agents like Claude and ChatGPT control connected devices, review camera summaries, and query smart home activity in natural language. This is a concrete template for how a consumer hardware platform exposes itself to third-party agents — useful if you're designing tool surfaces for your own product.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning