Case Study Public Data Intelligence

Building MyCleanData

How the MyWaypoint refinery transforms public information into governed, searchable, API-accessible data products.

Platform deployment Public data sources All 7 refinery stages Live at mycleandata.ca ↗
01 — The Challenge

Public information is fragmented.

Public information is abundant. It is also inconsistent, scattered, and difficult to use.

Publicly available information arrives in dozens of formats — CSV files from one source, JSON feeds from another, HTML tables embedded in web pages, PDFs containing structured information that cannot be queried. Different sources use incompatible schemas. No common standard applies across the landscape. Update cadences vary. Field names change without notice. Metadata is missing, incomplete, or inconsistently applied.

The result is a large volume of information that exists but cannot be easily assembled, compared, searched, or consumed programmatically. The data is there. The usability is not.

Multiple jurisdictions

Many jurisdictions and source types

Format variation

CSV, JSON, XML, APIs, PDFs

Schema inconsistency

Different field names, types, conventions

Irregular updates

Variable cadences, no standard

Missing metadata

Attribution incomplete or absent

Not machine-readable

Requires extraction before use

02 — The Approach

Applying the refinery model.

The MyWaypoint refinery was applied to public information sources using the same seven-stage pipeline it applies to any data environment.

Each stage addressed a specific problem in the public data landscape — inconsistent formats, missing attribution, lack of searchability, absence of governance.

The pipeline did not require the sources to be clean or consistent before ingestion. Fixing the inconsistency is what the refinery does.

01

Acquire

Collect from public sources — portals, registries, websites, files, and feeds across jurisdictions and source types.

02

Validate

Scan each incoming source for completeness, structural integrity, and format conformance. Flag records with missing fields, structural violations, or quality issues before they enter the pipeline.

03

Normalize

Convert CSV, JSON, XML, and API responses into a unified record schema. Standardize field names, value formats, and data types across sources that were never designed to be consistent with each other.

04

Enrich

Attach metadata — source attribution, collection categories, update timestamps, geographic context, and topic classification. Add semantic context that makes records findable and useful beyond exact-match search.

05

Govern

Record the full lineage of each dataset — source URL, acquisition date, processing history, version tracking, and update frequency. Apply documentation standards so every dataset has an accurate, maintainable record of what it is and where it came from.

06

Search

Index normalized, enriched records for full-text search, faceted filtering, and semantic retrieval. Every dataset becomes discoverable through keyword, category, and contextual search.

07

Serve

Deliver through the MyCleanData catalog — searchable interface, organized collections, structured API endpoints, and downloadable dataset files. The result is publicly accessible and operationally maintained.

03 — The Result

The result: MyCleanData.

MyCleanData is what the refinery produced. A publicly accessible data platform at mycleandata.ca, built entirely from the seven-stage pipeline.

It contains publicly available datasets that have been acquired, validated, normalized, enriched, governed, indexed, and made available for search, API access, and download. Each dataset carries source attribution, processing history, and documentation generated through the refinery — not written manually afterward.

The platform is operational. It is maintained through the same pipeline that built it. When sources update, the refinery runs. When new sources are added, they move through the same process. The governance does not degrade over time because it is built into the operational model, not applied once and abandoned.

Searchable catalog

Governed, publicly available datasets

All 7 stages

Full refinery pipeline applied to every dataset

Maintained

Continuously operated through the same pipeline

04 — Outputs

What the refinery produced.

Each capability in MyCleanData emerged directly from the refinery pipeline. None of it was built separately.

Searchable Catalog

Full-text and faceted search across every dataset. Possible because each record was normalized and indexed through the pipeline — not because search was added afterward.

Produced by: Normalize → Search stages

Collections

Curated groupings of related datasets organized by topic, geography, or data category. Collections emerged from the classification and enrichment applied during the Enrich stage.

Produced by: Enrich → Govern stages

API Access

Structured API endpoints delivering dataset content and metadata programmatically. Possible because normalization created consistent, predictable schemas that map directly to API fields.

Produced by: Normalize → Serve stages

Governance Records

Source attribution, lineage, and update history for every dataset. Governance records were produced by the pipeline, not written manually — they are a direct output of the Govern stage.

Produced by: Acquire → Govern stages

Documentation

Dataset descriptions, metadata schemas, field definitions, and provenance notes generated from the refinery process. Documentation is a structural output — not content authored separately after the fact.

Produced by: Enrich → Govern stages

Learning Resources

Tutorials and usage guides explaining how to access and work with the data. Built alongside the platform as part of the governance and documentation layer rather than as an afterthought.

Produced by: Govern → Serve stages

05 — Enterprise Application

The same model applies to enterprise data.

MyCleanData demonstrates what the refinery produces. Enterprise deployments apply the same process to private, internal, and custom data environments.

MyCleanData — Public Deployment
Sources Publicly available information from many source types
Process All 7 refinery stages applied to each dataset
Output Public catalog, collections, APIs, governance records
Access Publicly available at mycleandata.ca
Enterprise Deployment — Your Data
Sources Internal documents, repositories, APIs, and custom data
Process Same 7 stages — same governance, same quality controls
Output Private knowledge base, internal APIs, governed data products
Access Controlled by your organization — not public

The platform is not different for enterprise use. The sources are. The governance controls, pipeline stages, and output formats are the same.

06 — Lessons

Refinement matters.

Data quality matters more than data volume.

More datasets did not create more value. Consistently refined datasets did. The decisions that improved searchability, API reliability, and governance quality were quality decisions, not quantity decisions. Coverage without quality produces a large collection of unreliable information — not a useful platform.

Governance enables trust at scale.

At catalog scale, manually maintaining quality is not realistic. Governance that is embedded in the pipeline — lineage tracking, source attribution, processing logs — does not degrade as scale increases. Governance applied as a manual layer afterward does. The refinery model was designed so governance is a structural property of every output, not an activity layered on top of it.

Documentation is part of the data product.

Undocumented data is not useful data — it is data that requires additional work before it can be used. Building documentation generation into the Enrich and Govern stages meant that documentation existed from the moment a dataset was produced, not as a follow-up task. This made the platform navigable at launch and keeps it navigable as new datasets are added.

AI performs better on refined information.

Semantic search, retrieval-augmented generation, and AI-assisted discovery all operate on the enriched, normalized content produced by the refinery — not on the raw source data. The quality of the source information that enters the search and retrieval layer directly determines the quality of what comes out. Refining before indexing is not optional for reliable AI retrieval. It is the foundation of it.

Consistency creates value.

A catalog where 90% of datasets follow a consistent schema and 10% do not is less useful than the 90% figure suggests. Inconsistency in the final 10% breaks assumptions that developers and analysts depend on. Predictable schemas enable reliable integrations. Consistent governance enables trust. The refinery produces consistency systematically — it does not depend on individual discipline applied source by source.

About this case study

MyCleanData is our current public case study — a live demonstration of the refinery applied to publicly available information. Enterprise deployments involving private or proprietary data are governed by confidentiality agreements. Additional case studies from enterprise engagements will be published as they become available under appropriate disclosure.

See what the refinery produces.

MyCleanData is live at mycleandata.ca. The same refinery model is available for private, enterprise, and custom data environments.