AI training datasets
Ten million rights‑cleared image‑text pairs
Custom AI training datasets built from the Piemags archive. Public domain imagery, paired with our own captions, 30 keywords per image, dates and a subject taxonomy. Multimodal training data on every conceivable subject, cut to your specification.
- Source ref
- LC‑USW36‑301 · fsac.1a35373
- Rights
- Public domain No copyright asserted in the image.
- Caption
- A woman operates a hand drill on a Vultee “Vengeance” dive bomber at the Nashville Division of Vultee Aircraft Inc., Tennessee, in February 1943. Photographed on colour transparency by Alfred T. Palmer for the Office of War Information.
- Keywords
- vultee vengeance; a-31 vengeance; vultee aircraft; dive bomber; hand drill; power tool; aircraft factory; aircraft production; aircraft assembly; aviation manufacturing; factory floor; woman worker; women in wartime; war production; wartime industry; home front; second world war; alfred t palmer; office of war information; farm security administration; nashville division; industrial photography; documentary photography; colour transparency; production line; 1943; nashville; tennessee; united states; historical photograph
30 of 30 - Date
- 1943‑02 · month certain
- Creator
- Alfred T. Palmer
- Taxonomy
- Industry › Aviation manufacturing › Second World War › United States
- Provenance
- Library of Congress, Farm Security Administration and Office of War Information Collection 12002‑41. Public domain as a work of the United States federal government, 17 U.S.C. §105. Holding institution rights statement: “no known restrictions on publication”. Source catalogue record retained in full.
- Caption source
- Generated from the source record, human reviewed.
available to licence now
full Piemags archive
every image
source imagery
What you licence
The pixels are free. The description is not.
Every image in this archive is in the public domain, and we say so plainly. Anyone can go and find them. What almost nobody has is ten million of them gathered in one place, deduplicated, dated, classified and described in a consistent, machine‑readable form.
That description layer is the product. Each record carries a factual caption, a keyword set of thirty terms, a date with a confidence marker, a position in a subject taxonomy, and a provenance note recording where the image came from and the basis on which we assess it to be in the public domain.
Fields are configurable. If your training pipeline wants captions only, you get captions only. If it wants long‑form captions, short captions and a tag list as three separate columns, that is a delivery option, not a rebuild.
The honest question
Why pay for public domain images?
It is the first thing any serious buyer asks, so here is the answer without the marketing.
The large free image‑text corpora are built from scraped alt‑text or from synthetic captions written by a vision model. Both are cheap for a reason. Alt‑text is whatever a web author happened to type, and synthetic captions describe what is visible without knowing what it is. Neither can tell you that the aircraft is an F‑86A rather than an F‑86F, or that the photograph was taken at Suwon in 1952.
The authors of PD12M, a 12.4 million image public domain dataset with synthetic captions, say this themselves in their own paper.
On the record
“Synthetic captions lack the contextual metadata (artistic medium, historical details).”
PD12M authors · arXiv:2410.23144
That gap is exactly what Piemags fills. And the market has already priced it. Shutterstock’s data, distribution and services segment took $203.3 million in the 2025 financial year, up 16 per cent, and the company attributes the increase “primarily from the sale and delivery of metadata to new and existing customers”.
Shutterstock FY2025 results · investor.shutterstock.com
The build‑it‑yourself cost
What it would cost you to make this instead
Commercial human keywording services publish rates of roughly €0.98 to €1.47 per image for a written title plus a keyword set, which is the same job a Piemags record does. At ten million images that is a project measured in tens of millions and in years, not months. Most annotation vendors also cannot identify historical subject matter at all, which is the part that takes the time.
Rate source: MicrostockGuru published pricing, checked August 2026. Piemags pricing below is set well under that replacement cost.
Coverage
Every conceivable subject
The archive runs from the nineteenth century to the present day and is not organised around a single theme. If you need a dataset on one narrow subject, we cut it. If you need breadth, it is already there.
Not listed is not the same as not held. Tell us the subject and we will tell you honestly how many records we have, before you commit to anything.
Rights and provenance
Stated plainly, because your lawyers will read this page
A lot of this market is vague about rights. We would rather be exact, because being exact is what makes the licence worth having.
- The images Every image is in the public domain. Piemags asserts no copyright in the images themselves. Digitising a public domain image does not create a new copyright in it, and we do not claim one.
- The dataset Piemags Ltd is the maker of this database within the meaning of regulation 14 of the Copyright and Rights in Databases Regulations 1997, and asserts UK database right in it. Substantial investment has been made in obtaining, verifying and presenting the contents.
- The licence Access is granted under a written licence agreement. The licence governs your use of the dataset we supply. It does not, and cannot, restrict what anyone does with public domain images obtained elsewhere.
- What you pay for Curated access, bulk delivery, per‑record provenance documentation, and a contractual warranty and indemnity package. Not an assignment of copyright in the underlying images.
- Provenance Every record carries its source, the basis on which we assess it to be in the public domain, and the date of that assessment. Public domain status is assessed jurisdiction by jurisdiction. Tell us which territory you need cleared.
- Caption origin Captions and keywords in this collection were generated from source records with machine assistance and reviewed by hand. Provenance flags are supplied at record level so you always know which is which.
- Not covered Trade mark, personality, publicity and data protection rights in the subject matter of images are not covered by the licence and remain the licensee’s responsibility. We will not pretend otherwise.
- EU AI Act Under the Commission’s training‑content template, licensed content requires only confirmation that it was obtained under agreement with rights holders, where scraped content requires domain listings, crawler behaviour and collection dates. Our provenance records are built to support Article 53 disclosure.
Pricing
Indicative rates, published
Most sellers in this market will not put a number on a page. We will, because you deserve to know whether this is worth a conversation before you spend a fortnight arranging one.
| Licence | Scope | Term | Indicative |
|---|---|---|---|
| Academic and research | Up to 1 million pairs, non‑commercial, named institution, no model redistribution | Perpetual | from £30,000 |
| Subject cut | One subject area, volume to suit, commercial use, non‑exclusive | Perpetual | from £50,000 |
| Commercial, non‑exclusive | Full 10 million pairs, one model family, extension fee for further families | Perpetual training licence, 12 month data access | from £1,000,000 |
| Enterprise | Full 10 million pairs, all current and future models, format engineering, delivery to your bucket, warranty and indemnity, 24 month refresh | Perpetual | from £2,500,000 |
| Exclusive | Full or category exclusivity, first look at new material | Negotiated | On application |
Minimum engagement £50,000. An annual term licence with quarterly refresh is available in place of a perpetual licence at 35 to 40 per cent of the perpetual fee per year. Prices are indicative and exclude VAT. The commercial tier works out at roughly ten pence per image‑text pair, against a replacement cost of the order of a pound per pair if you built it yourself.
Process
How a dataset gets built
Four steps, and you can stop at any of them. Nothing is signed until step four.
-
Tell us the specification
Subject, volume, date range, resolution, field set and delivery format. If you are not sure, describe the model you are training and we will suggest a shape.
-
Sample and scoping note
We send a free sample cut to your specification, with a written note on how many records exist, what the metadata actually looks like and where the gaps are.
-
You evaluate
Run it. Benchmark it against what you already have. Come back with what needs changing. Most specifications change at this point and that is expected.
-
Licence and delivery
A written licence with warranty and indemnity, then bulk delivery to your storage in the format you asked for, with provenance records alongside.
Questions
Straight answers
Do you own the images?
No, and we do not claim to. Every image is in the public domain. What Piemags owns is the dataset built around them: the captions, keywords, dates, taxonomy and provenance records, and the compiled and indexed collection itself.
What exactly am I licensing, then?
Curated access to the dataset, bulk delivery in your format, per‑record provenance documentation, and a contractual warranty and indemnity package. You are not buying a copyright assignment in the images, because there is no copyright in them to assign.
Can I get a sample before committing?
Yes. Samples are supplied on request, cut to your specification rather than pulled off a shelf, under a non‑commercial evaluation licence. Email sales@piemags.com and say what you need.
What format do you deliver in?
CSV, JSONL or Parquet for the metadata, with images delivered as files or as resolvable URLs, to your object storage or ours. WebDataset shards and pre‑joined image‑text pair formats are available. Tell us what your loader expects.
Are the captions written by a human or by a machine?
Both, and every record says which. Captions are generated from the underlying archival source record, then reviewed by hand. That is a fundamentally different thing from a vision model guessing at a picture with no source material, and it is why the captions carry facts a model cannot see: unit designations, place names, dates and proper nouns.
Can I have exclusivity?
On the dataset, yes, in full or by category or territory, and it is priced accordingly. On the images, no. Nobody can be exclusive over the public domain and anyone who tells you otherwise is selling you something they do not have.
Does this help with EU AI Act obligations?
It should make them considerably easier. Under the European Commission’s public training‑content template, licensed content needs only confirmation that it was obtained under agreement with rights holders, where scraped content needs domain listings, crawler behaviour and collection dates. Our per‑record provenance is built with that disclosure in mind. It is not legal advice and your counsel will want to see the actual records, which we will supply.
I only want one narrow subject. Is that worth your time?
Yes. Subject cuts start at £50,000 and are often the more useful product. A tightly scoped, deeply described set on one subject will usually do more for a fine‑tune than ten million mixed records.
Who is Piemags?
Piemags Ltd is a UK company, registered in England and Wales, number 11786963, based in Bridgend, Wales. It is a specialist visual archive whose imagery is licensed through Alamy, Getty, Pond5, IMAGO and SuperStock. The dataset business is built on the same archive and the same metadata pipeline.
Start here
Tell us what you are training
One email, and a real person reads it. Say what subject you need, roughly what volume, and what your loader expects. We will come back with what we hold, a sample and a price.
- Enquiries
- sales@piemags.com
- Company
- Piemags Ltd
- Registered
- England and Wales, no. 11786963
- Based
- Bridgend, Wales, United Kingdom
- Available now
- 10,000,000+ image‑text pairs
Piemags Ltd, registered in England and Wales, company number 11786963. All images referenced on this page are in the public domain and Piemags asserts no copyright in them. Piemags asserts UK database right in the compiled dataset. Pricing is indicative, excludes VAT and does not constitute an offer. Nothing on this page is legal advice.