An API is the freshest source and the least reproducible, so a corpus built from one has to be materialised before it can be measured against.
ConceptWhat it is
An API source is a live service — a CRM, a ticketing system, a product catalogue — read over the network at ingest time. It is attractive because it is current by construction and needs no export process.
It is also the source with the weakest guarantees. Rate limits shape what you can read, pagination hides changes that happen mid-read, and the provider can alter a field or deprecate an endpoint without warning. The corpus is never twice the same, and unless that is handled deliberately, neither is any number computed from it.
How it worksThe mechanics
Reads are materialised: the response is written to durable storage exactly as received, with a timestamp and the request that produced it, before anything transforms it. That raw landing zone is what makes the corpus reproducible and what makes a parsing bug fixable without re-reading the service.
Around that sit the operational realities. Pagination is made consistent with a cursor or a stable sort, so a record cannot be seen twice or missed while the underlying data shifts. Rate limits are respected with backoff. Partial failures are recorded rather than silently producing a shorter corpus, which is the failure mode that looks like success.
At a glanceSee it
Landing the raw response first is what makes a parsing bug fixable without re-reading the service, and what gives the corpus a version at all.
When to use itWhere it fits
- When the source system has no export and the API is the only supported route in.
- Where freshness is a genuine product requirement rather than an assumption.
- For enrichment against a reference service, where a small number of lookups add real context.
- When the provider offers a change feed, which turns an unreliable poll into an ordered stream.
When NOT to use itLimits & anti-patterns
- As a live read at query time when the answer must be reproducible or the latency budget is tight.
- Under an eval, without pinning to a landed version — otherwise the score measures the provider.
- Where rate limits make a full corpus read impractical and the partial read is treated as complete.
- When a bulk export exists and is simply less trouble, which is more often than teams check.
Trade-offsAdvantages & costs
Advantages
- Current by construction, with no export process to schedule or maintain.
- Usually returns structured data with a documented schema and real types.
- Incremental reads are natural when the API exposes modification timestamps.
- Often the only route into a SaaS system you do not control.
Trade-offs & costs
- Not reproducible unless landed, so an unpinned eval score is partly a measurement of the provider.
- Rate limits and pagination make a complete, consistent read genuinely hard.
- Schema and field semantics change under you with little notice.
- Partial failures shrink the corpus quietly, and no downstream check will notice a missing tenth.
ExampleIn the real world
A product-support corpus is built by paging a catalogue API sorted by last-modified. Items edited during the read shuffle position, so a few are read twice and a few are never read at all. The index looks fine and is complete-looking. The gap only surfaced when a customer asked about a product the assistant insisted did not exist.
ToolsHow to implement it
- A raw landing zone on object storagethe single change that makes an API source reproducible.
- Airbyte or Fivetranmanaged connectors that have already solved pagination and rate limits for common systems.
- Tenacity or equivalent backoffretries that respect the limit instead of amplifying against it.
- Cursor-based pagination where offeredthe only pagination that is correct while the data underneath is moving.
Cost & effortWhat it takes
Often free or metered per call, with the real constraint being rate limits rather than money. Landing raw responses costs negligible storage and repays itself the first time a parser needs fixing. Engineering effort is moderate and concentrated in the unglamorous parts — pagination, retries and partial-failure accounting.