Skip to main content

One post tagged with "extensions"

View All Tags

One Model to Generate Them All

· 21 min read
Seth Fitzsimmons
Independent Consultantseth@mojodna.net
Jennings Anderson
GeoInformation Scientist, Metajenningsa@meta.com

If you publish geospatial data, you've probably done this. You put a dataset on a data portal or shared it with colleagues, maybe as a Shapefile, a GeoPackage, or a GeoParquet file. You describe its columns somewhere else: a PDF data dictionary, an FGDC metadata record, a spreadsheet of codes. Someone asks what a code means, so you add a note to the description, and later you fix a column in the data without going back to the description. Now there are three versions of the truth: the data, the portal entry, and the document. They disagree.

You do have a model. You just wrote it in a place nothing else can read. The data dictionary describes the columns, but no validator checks the data against it, and no tool can look up which values are legal. Every copy of that knowledge gets maintained by hand, and hand-maintained copies drift.

We had the same problem at Overture, at a larger scale. The Overture Schema started as JSON Schema, hand-written as YAML, with reference docs written separately and validation that only ran on JSON. When the schema changed, the docs had to catch up by hand, and they didn't always. So we asked how we could write the model so that the data, the validation, and the reference material all derive from it and can't diverge. Our answer was code. Last month we published v2.0.0 of the schema as a Python library built on Pydantic. The model is the source, and the JSON Schema, the reference docs, and the validation checks are all generated from it.