API documentation
Load any licensed dataset in three lines of Python.
Quickstart
Install the SDK:
pip install corporaThen load a dataset you have licensed:
import corpora
corpora.api_key = "sk_live_..."
ds = corpora.load("pit-radio-1998")
print(ds.schema)Authentication
All requests require an API key passed in the Authorization header. Keys are created in your workspace settings.
curl https://api.corpora.ai/v1/datasets \
-H "Authorization: Bearer sk_live_..."Search datasets
GET /v1/datasets?q=<query> returns datasets matching your query, including ones you have not licensed.
results = corpora.search("legal text", fmt="Text")
for d in results:
print(d.id, d.price_chf, d.status)Streaming
Large datasets are streamed in batches. Batches are shuffled and de-duplicated, and names are removed where detected.
for batch in ds.stream(size=512, shuffle=True):
model.train(batch)Licences
Every response includes a licence object. Check licence.status before training.
| Status | Meaning |
|---|---|
commercial | Cleared for training and production. |
research_only | Non-commercial use. Honour system. |
pending | Seller is obtaining rights. Usable in the meantime. |
unknown | Status unknown to all parties. |
Rate limits
Explorer: 60 requests per minute. Team: 6,000. Enterprise: unlimited, subject to available data.
Errors
| Code | Meaning |
|---|---|
401 | Invalid API key. |
402 | Licence required for this dataset. |
410 | Dataset withdrawn. A similar one may be relisted shortly. |
451 | Unavailable for legal reasons in your jurisdiction. Try another region. |