AI training datasets

Ten million rights‑cleared image‑text pairs

Custom AI training datasets built from the Piemags archive. Public domain imagery, paired with our own captions, 30 keywords per image, dates and a subject taxonomy. Multimodal training data on every conceivable subject, cut to your specification.

Record structure Live record
A woman operates a hand drill on a Vultee Vengeance dive bomber at the Nashville Division of Vultee Aircraft Inc., Tennessee, February 1943. Photographed by Alfred T. Palmer for the Office of War Information.
Source ref
LC‑USW36‑301 · fsac.1a35373
Rights
Public domain No copyright asserted in the image.
Caption
A woman operates a hand drill on a Vultee “Vengeance” dive bomber at the Nashville Division of Vultee Aircraft Inc., Tennessee, in February 1943. Photographed on colour transparency by Alfred T. Palmer for the Office of War Information.
Keywords
vultee vengeance; a-31 vengeance; vultee aircraft; dive bomber; hand drill; power tool; aircraft factory; aircraft production; aircraft assembly; aviation manufacturing; factory floor; woman worker; women in wartime; war production; wartime industry; home front; second world war; alfred t palmer; office of war information; farm security administration; nashville division; industrial photography; documentary photography; colour transparency; production line; 1943; nashville; tennessee; united states; historical photograph
30 of 30
Date
1943‑02 · month certain
Creator
Alfred T. Palmer
Taxonomy
Industry › Aviation manufacturing › Second World War › United States
Provenance
Library of Congress, Farm Security Administration and Office of War Information Collection 12002‑41. Public domain as a work of the United States federal government, 17 U.S.C. §105. Holding institution rights statement: “no known restrictions on publication”. Source catalogue record retained in full.
Caption source
Generated from the source record, human reviewed.
Live record from the Piemags Library of Congress holdings. Caption and keywords shown to the Piemags 30‑keyword specification. Field set is configurable per dataset.
10M+ Image‑text pairs
available to licence now
30M+ Images in the
full Piemags archive
30 Keywords on
every image
100% Public domain
source imagery

What you licence

The pixels are free. The description is not.

Every image in this archive is in the public domain, and we say so plainly. Anyone can go and find them. What almost nobody has is ten million of them gathered in one place, deduplicated, dated, classified and described in a consistent, machine‑readable form.

That description layer is the product. Each record carries a factual caption, a keyword set of thirty terms, a date with a confidence marker, a position in a subject taxonomy, and a provenance note recording where the image came from and the basis on which we assess it to be in the public domain.

Fields are configurable. If your training pipeline wants captions only, you get captions only. If it wants long‑form captions, short captions and a tag list as three separate columns, that is a delivery option, not a rebuild.

The honest question

Why pay for public domain images?

It is the first thing any serious buyer asks, so here is the answer without the marketing.

The large free image‑text corpora are built from scraped alt‑text or from synthetic captions written by a vision model. Both are cheap for a reason. Alt‑text is whatever a web author happened to type, and synthetic captions describe what is visible without knowing what it is. Neither can tell you that the aircraft is an F‑86A rather than an F‑86F, or that the photograph was taken at Suwon in 1952.

The authors of PD12M, a 12.4 million image public domain dataset with synthetic captions, say this themselves in their own paper.

On the record

“Synthetic captions lack the contextual metadata (artistic medium, historical details).”

PD12M authors · arXiv:2410.23144

That gap is exactly what Piemags fills. And the market has already priced it. Shutterstock’s data, distribution and services segment took $203.3 million in the 2025 financial year, up 16 per cent, and the company attributes the increase “primarily from the sale and delivery of metadata to new and existing customers”.

Shutterstock FY2025 results · investor.shutterstock.com

The build‑it‑yourself cost

What it would cost you to make this instead

Commercial human keywording services publish rates of roughly €0.98 to €1.47 per image for a written title plus a keyword set, which is the same job a Piemags record does. At ten million images that is a project measured in tens of millions and in years, not months. Most annotation vendors also cannot identify historical subject matter at all, which is the part that takes the time.

Rate source: MicrostockGuru published pricing, checked August 2026. Piemags pricing below is set well under that replacement cost.

Coverage

Every conceivable subject

The archive runs from the nineteenth century to the present day and is not organised around a single theme. If you need a dataset on one narrow subject, we cut it. If you need breadth, it is already there.

Aviation and aerospaceMilitary and civil aircraft, air‑to‑air, airfields, airshows, spaceflight and launch operations.
Military and conflictLand, sea and air operations, personnel, equipment, ceremonies and training, across more than a century.
Cities and architectureStreets, squares, buildings, interiors and construction, with street addresses recorded in the caption where the source gives them.
Transport and industryRailways, shipping, motor vehicles, factories, engineering works and infrastructure.
Portraits and daily lifeStudio and documentary portraiture, work, leisure, markets, schools and domestic scenes.
Art, prints and engravingsPaintings, drawings, prints, posters and book illustration, flagged as reproductions where they are reproductions.
Maps and chartsTopographic, nautical, military and thematic cartography.
Science and technologyInstruments, laboratories, medicine, computing and technical diagrams.
Natural worldLandscape, geology, flora and fauna, weather and environmental documentation.
Objects and artefactsCoins, ceramics, tools, textiles, furniture and museum objects photographed against plain grounds.
Events and ceremoniesState occasions, sport, protest, disasters and public gatherings.
Agriculture and rural lifeFarming, fishing, forestry, villages and working landscapes.

Not listed is not the same as not held. Tell us the subject and we will tell you honestly how many records we have, before you commit to anything.

Rights and provenance

Stated plainly, because your lawyers will read this page

A lot of this market is vague about rights. We would rather be exact, because being exact is what makes the licence worth having.

  • The images Every image is in the public domain. Piemags asserts no copyright in the images themselves. Digitising a public domain image does not create a new copyright in it, and we do not claim one.
  • The dataset Piemags Ltd is the maker of this database within the meaning of regulation 14 of the Copyright and Rights in Databases Regulations 1997, and asserts UK database right in it. Substantial investment has been made in obtaining, verifying and presenting the contents.
  • The licence Access is granted under a written licence agreement. The licence governs your use of the dataset we supply. It does not, and cannot, restrict what anyone does with public domain images obtained elsewhere.
  • What you pay for Curated access, bulk delivery, per‑record provenance documentation, and a contractual warranty and indemnity package. Not an assignment of copyright in the underlying images.
  • Provenance Every record carries its source, the basis on which we assess it to be in the public domain, and the date of that assessment. Public domain status is assessed jurisdiction by jurisdiction. Tell us which territory you need cleared.
  • Caption origin Captions and keywords in this collection were generated from source records with machine assistance and reviewed by hand. Provenance flags are supplied at record level so you always know which is which.
  • Not covered Trade mark, personality, publicity and data protection rights in the subject matter of images are not covered by the licence and remain the licensee’s responsibility. We will not pretend otherwise.
  • EU AI Act Under the Commission’s training‑content template, licensed content requires only confirmation that it was obtained under agreement with rights holders, where scraped content requires domain listings, crawler behaviour and collection dates. Our provenance records are built to support Article 53 disclosure.

Pricing

Indicative rates, published

Most sellers in this market will not put a number on a page. We will, because you deserve to know whether this is worth a conversation before you spend a fortnight arranging one.

Indicative dataset licence pricing
Licence Scope Term Indicative
Academic and research Up to 1 million pairs, non‑commercial, named institution, no model redistribution Perpetual from £30,000
Subject cut One subject area, volume to suit, commercial use, non‑exclusive Perpetual from £50,000
Commercial, non‑exclusive Full 10 million pairs, one model family, extension fee for further families Perpetual training licence, 12 month data access from £1,000,000
Enterprise Full 10 million pairs, all current and future models, format engineering, delivery to your bucket, warranty and indemnity, 24 month refresh Perpetual from £2,500,000
Exclusive Full or category exclusivity, first look at new material Negotiated On application

Minimum engagement £50,000. An annual term licence with quarterly refresh is available in place of a perpetual licence at 35 to 40 per cent of the perpetual fee per year. Prices are indicative and exclude VAT. The commercial tier works out at roughly ten pence per image‑text pair, against a replacement cost of the order of a pound per pair if you built it yourself.

Process

How a dataset gets built

Four steps, and you can stop at any of them. Nothing is signed until step four.

  1. Tell us the specification

    Subject, volume, date range, resolution, field set and delivery format. If you are not sure, describe the model you are training and we will suggest a shape.

  2. Sample and scoping note

    We send a free sample cut to your specification, with a written note on how many records exist, what the metadata actually looks like and where the gaps are.

  3. You evaluate

    Run it. Benchmark it against what you already have. Come back with what needs changing. Most specifications change at this point and that is expected.

  4. Licence and delivery

    A written licence with warranty and indemnity, then bulk delivery to your storage in the format you asked for, with provenance records alongside.

Questions

Straight answers

Do you own the images?

No, and we do not claim to. Every image is in the public domain. What Piemags owns is the dataset built around them: the captions, keywords, dates, taxonomy and provenance records, and the compiled and indexed collection itself.

What exactly am I licensing, then?

Curated access to the dataset, bulk delivery in your format, per‑record provenance documentation, and a contractual warranty and indemnity package. You are not buying a copyright assignment in the images, because there is no copyright in them to assign.

Can I get a sample before committing?

Yes. Samples are supplied on request, cut to your specification rather than pulled off a shelf, under a non‑commercial evaluation licence. Email sales@piemags.com and say what you need.

What format do you deliver in?

CSV, JSONL or Parquet for the metadata, with images delivered as files or as resolvable URLs, to your object storage or ours. WebDataset shards and pre‑joined image‑text pair formats are available. Tell us what your loader expects.

Are the captions written by a human or by a machine?

Both, and every record says which. Captions are generated from the underlying archival source record, then reviewed by hand. That is a fundamentally different thing from a vision model guessing at a picture with no source material, and it is why the captions carry facts a model cannot see: unit designations, place names, dates and proper nouns.

Can I have exclusivity?

On the dataset, yes, in full or by category or territory, and it is priced accordingly. On the images, no. Nobody can be exclusive over the public domain and anyone who tells you otherwise is selling you something they do not have.

Does this help with EU AI Act obligations?

It should make them considerably easier. Under the European Commission’s public training‑content template, licensed content needs only confirmation that it was obtained under agreement with rights holders, where scraped content needs domain listings, crawler behaviour and collection dates. Our per‑record provenance is built with that disclosure in mind. It is not legal advice and your counsel will want to see the actual records, which we will supply.

I only want one narrow subject. Is that worth your time?

Yes. Subject cuts start at £50,000 and are often the more useful product. A tightly scoped, deeply described set on one subject will usually do more for a fine‑tune than ten million mixed records.

Who is Piemags?

Piemags Ltd is a UK company, registered in England and Wales, number 11786963, based in Bridgend, Wales. It is a specialist visual archive whose imagery is licensed through Alamy, Getty, Pond5, IMAGO and SuperStock. The dataset business is built on the same archive and the same metadata pipeline.

Start here

Tell us what you are training

One email, and a real person reads it. Say what subject you need, roughly what volume, and what your loader expects. We will come back with what we hold, a sample and a price.

Enquiries
sales@piemags.com
Company
Piemags Ltd
Registered
England and Wales, no. 11786963
Based
Bridgend, Wales, United Kingdom
Available now
10,000,000+ image‑text pairs