UK Geothermal Digital Working Group UK Geothermal Data Catalogue

About this catalogue

What it is

An index of UK geothermal data: what exists, who holds it, whether you may use it, and how to get at it. It links to the holder and does not host the data. Subsurface records are harvested from BGS's own metadata and credited to them; the depth here is on the demand, infrastructure and constraints side, which nobody currently indexes.

What it is not

It is not a data store, not a replacement for the BGS platform, and not a statement that anything listed is fit for your purpose. Where a licence is unclear, the record says so rather than guessing.

How records are classified

Two independent axes. Category says what a dataset is. Project stage says when in a project you need it, and comes from the working group's stage matrix. A record carries both, and neither replaces the other. Every record records how it was classified and whether a person has ever checked it.

Licences

Licence is recorded as an object with the holder's exact wording and the source of the assertion, because licences are inconsistent at source rather than merely ambiguous. The same dataset is described as Open Government Licence by one publisher and “Licence: Not set” by another. Where no licence is stated, the record says so. It is never guessed.

Of 352 records, 195 are openly licensed, 35 are unclear and 92 have no licence statement at all.

Every record with a landing page carries the date its link was last checked and the method used. The distinction matters more than the count. An automated check proves a server answered a script. It does not prove a reader can open the page.

Durham Research Online is the case that made the point. Its old address answered an automated request with the repository's front page, while sending real browsers through a redirect that refused the connection. The automated check recorded a pass; a person with a phone got an error. The repository had moved host, and the record now points there.

So the numbers are stated as what they are. 246 of 352 records have a landing page that answered an automated request. 11 sit behind bot protection that refused the check, which is evidence of neither working nor broken. 95 have no check result, almost all of them platform layers that have no page of their own and inherit their parent platform's link. None of these has been opened in a browser by a person. Where a record links somewhere important to you, click it before you rely on it.

Known gaps in this build

Three fields are missing across the whole catalogue: the holder's own persistent identifier for a dataset, the keywords they publish with it, and the URL of the licence they apply. None of them changes what a dataset is or whether you may use it, and the licence itself is recorded in words on every entry. They matter for machine reuse rather than for reading, and they come back on the next harvest from the holders.

Standards

Records are modelled on DCAT-AP 3.0, with a GeoDCAT-AP crosswalk for geospatial records harvested from ISO 19115 sources such as the BGS GeoNetwork, which publishes to UK GEMINI 2.3. Every record page emits schema.org/Dataset JSON-LD, and the whole catalogue is available as a DCAT catalogue at catalog.jsonld.

Where a source states UNFC codes they are carried on the record. GRMS is an announced initiative rather than a published standard, so a slot is reserved for it and nothing is built against it yet.

How to cite

Each record page carries its own citation string, including the record id and the catalogue URL. Cite the holder first; this catalogue is the route, not the source.