How the MyWaypoint refinery transforms public information into governed, searchable, API-accessible data products.
Public information is abundant. It is also inconsistent, scattered, and difficult to use.
Publicly available information arrives in dozens of formats — CSV files from one source, JSON feeds from another, HTML tables embedded in web pages, PDFs containing structured information that cannot be queried. Different sources use incompatible schemas. No common standard applies across the landscape. Update cadences vary. Field names change without notice. Metadata is missing, incomplete, or inconsistently applied.
The result is a large volume of information that exists but cannot be easily assembled, compared, searched, or consumed programmatically. The data is there. The usability is not.
Multiple jurisdictions
Many jurisdictions and source types
Format variation
CSV, JSON, XML, APIs, PDFs
Schema inconsistency
Different field names, types, conventions
Irregular updates
Variable cadences, no standard
Missing metadata
Attribution incomplete or absent
Not machine-readable
Requires extraction before use
The MyWaypoint refinery was applied to public information sources using the same seven-stage pipeline it applies to any data environment.
Each stage addressed a specific problem in the public data landscape — inconsistent formats, missing attribution, lack of searchability, absence of governance.
The pipeline did not require the sources to be clean or consistent before ingestion. Fixing the inconsistency is what the refinery does.
Acquire
Collect from public sources — portals, registries, websites, files, and feeds across jurisdictions and source types.
Validate
Scan each incoming source for completeness, structural integrity, and format conformance. Flag records with missing fields, structural violations, or quality issues before they enter the pipeline.
Normalize
Convert CSV, JSON, XML, and API responses into a unified record schema. Standardize field names, value formats, and data types across sources that were never designed to be consistent with each other.
Enrich
Attach metadata — source attribution, collection categories, update timestamps, geographic context, and topic classification. Add semantic context that makes records findable and useful beyond exact-match search.
Govern
Record the full lineage of each dataset — source URL, acquisition date, processing history, version tracking, and update frequency. Apply documentation standards so every dataset has an accurate, maintainable record of what it is and where it came from.
Search
Index normalized, enriched records for full-text search, faceted filtering, and semantic retrieval. Every dataset becomes discoverable through keyword, category, and contextual search.
Serve
Deliver through the MyCleanData catalog — searchable interface, organized collections, structured API endpoints, and downloadable dataset files. The result is publicly accessible and operationally maintained.
MyCleanData is what the refinery produced. A publicly accessible data platform at mycleandata.ca, built entirely from the seven-stage pipeline.
It contains publicly available datasets that have been acquired, validated, normalized, enriched, governed, indexed, and made available for search, API access, and download. Each dataset carries source attribution, processing history, and documentation generated through the refinery — not written manually afterward.
The platform is operational. It is maintained through the same pipeline that built it. When sources update, the refinery runs. When new sources are added, they move through the same process. The governance does not degrade over time because it is built into the operational model, not applied once and abandoned.
Searchable catalog
Governed, publicly available datasets
All 7 stages
Full refinery pipeline applied to every dataset
Maintained
Continuously operated through the same pipeline
Each capability in MyCleanData emerged directly from the refinery pipeline. None of it was built separately.
Full-text and faceted search across every dataset. Possible because each record was normalized and indexed through the pipeline — not because search was added afterward.
Produced by: Normalize → Search stages
Curated groupings of related datasets organized by topic, geography, or data category. Collections emerged from the classification and enrichment applied during the Enrich stage.
Produced by: Enrich → Govern stages
Structured API endpoints delivering dataset content and metadata programmatically. Possible because normalization created consistent, predictable schemas that map directly to API fields.
Produced by: Normalize → Serve stages
Source attribution, lineage, and update history for every dataset. Governance records were produced by the pipeline, not written manually — they are a direct output of the Govern stage.
Produced by: Acquire → Govern stages
Dataset descriptions, metadata schemas, field definitions, and provenance notes generated from the refinery process. Documentation is a structural output — not content authored separately after the fact.
Produced by: Enrich → Govern stages
Tutorials and usage guides explaining how to access and work with the data. Built alongside the platform as part of the governance and documentation layer rather than as an afterthought.
Produced by: Govern → Serve stages
MyCleanData demonstrates what the refinery produces. Enterprise deployments apply the same process to private, internal, and custom data environments.
The platform is not different for enterprise use. The sources are. The governance controls, pipeline stages, and output formats are the same.
More datasets did not create more value. Consistently refined datasets did. The decisions that improved searchability, API reliability, and governance quality were quality decisions, not quantity decisions. Coverage without quality produces a large collection of unreliable information — not a useful platform.
At catalog scale, manually maintaining quality is not realistic. Governance that is embedded in the pipeline — lineage tracking, source attribution, processing logs — does not degrade as scale increases. Governance applied as a manual layer afterward does. The refinery model was designed so governance is a structural property of every output, not an activity layered on top of it.
Undocumented data is not useful data — it is data that requires additional work before it can be used. Building documentation generation into the Enrich and Govern stages meant that documentation existed from the moment a dataset was produced, not as a follow-up task. This made the platform navigable at launch and keeps it navigable as new datasets are added.
Semantic search, retrieval-augmented generation, and AI-assisted discovery all operate on the enriched, normalized content produced by the refinery — not on the raw source data. The quality of the source information that enters the search and retrieval layer directly determines the quality of what comes out. Refining before indexing is not optional for reliable AI retrieval. It is the foundation of it.
A catalog where 90% of datasets follow a consistent schema and 10% do not is less useful than the 90% figure suggests. Inconsistency in the final 10% breaks assumptions that developers and analysts depend on. Predictable schemas enable reliable integrations. Consistent governance enables trust. The refinery produces consistency systematically — it does not depend on individual discipline applied source by source.
About this case study
MyCleanData is our current public case study — a live demonstration of the refinery applied to publicly available information. Enterprise deployments involving private or proprietary data are governed by confidentiality agreements. Additional case studies from enterprise engagements will be published as they become available under appropriate disclosure.
MyCleanData is live at mycleandata.ca. The same refinery model is available for private, enterprise, and custom data environments.