Skip to main content

2 posts tagged with "schema"

View All Tags

One Model to Generate Them All

· 21 min read
Seth Fitzsimmons
Independent Consultantseth@mojodna.net
Jennings Anderson
GeoInformation Scientist, Metajenningsa@meta.com

If you publish geospatial data, you've probably done this. You put a dataset on a data portal or shared it with colleagues, maybe as a Shapefile, a GeoPackage, or a GeoParquet file. You describe its columns somewhere else: a PDF data dictionary, an FGDC metadata record, a spreadsheet of codes. Someone asks what a code means, so you add a note to the description, and later you fix a column in the data without going back to the description. Now there are three versions of the truth: the data, the portal entry, and the document. They disagree.

You do have a model. You just wrote it in a place nothing else can read. The data dictionary describes the columns, but no validator checks the data against it, and no tool can look up which values are legal. Every copy of that knowledge gets maintained by hand, and hand-maintained copies drift.

We had the same problem at Overture, at a larger scale. The Overture Schema started as JSON Schema, hand-written as YAML, with reference docs written separately and validation that only ran on JSON. When the schema changed, the docs had to catch up by hand, and they didn't always. So we asked how we could write the model so that the data, the validation, and the reference material all derive from it and can't diverge. Our answer was code. Last month we published v2.0.0 of the schema as a Python library built on Pydantic. The model is the source, and the JSON Schema, the reference docs, and the validation checks are all generated from it.

The Overture Schema Is a Library Now

· 9 min read
Dana Bauer
Technical Product Manager, Overturedana@overturemaps.org
Seth Fitzsimmons
Independent Consultantseth@mojodna.net
Victor Schappert
Principal Engineer, AWSschapper@amazon.com
Jennings Anderson
GeoInformation Scientist, Metajenningsa@meta.com
Roel Bollens
Technical Program Manager, TomTomroel.bollens@tomtom.com
Tristan Diet
Specifications Engineer, TomTomtristan.diet@tomtom.com

Two weeks ago we quietly published v2.0.0 of the Overture schema to PyPI. It used to be JSON Schema, lovingly handwritten in YAML to get around some of JSON’s rough edges. Now it’s a Python library you install and import, with Pydantic models you can inspect, validate data against, build on, and extend.

Try it out:

pip install overture-schema
>>> from overture.schema.places import Place
>>> sorted(Place.model_fields)

['addresses', 'basic_category', 'bbox', 'brand', 'confidence', 'emails', 'geometry', 'id', 'names', 'operating_status', 'phones', 'socials', 'sources', 'taxonomy', 'theme', 'type', 'version', 'websites']

>>> print(Place.model_fields["taxonomy"].description)

A structured representation of the place's category within the Overture taxonomy.
Provides the primary classification, full hierarchy path, and alternate categories.

You can ask a feature type what fields it has, read the documentation for anything, and validate your own data against a feature type or model. You can do that from the Python interpreter, your favorite IDE with type hints and completion, or at scale within a Spark job.

For some of you, our migration to Pydantic isn't a big deal. Maybe you'll notice that we made a few correctness fixes to the schema structure and improved our documentation. For others, this is a huge and welcome change. The schema has gone from a document you read to code you can build with.