Platform Architecture

How the Data Refinery Works

Transform raw information into governed, searchable, API-ready, and AI-ready data products.

The refinery applies a consistent seven-stage process to any structured or unstructured information, regardless of source type or origin.

How information flows through the refinery

Input Sources

Documents & PDFs
Websites & Web content
APIs & Data feeds
Databases
Structured files (CSV, JSON, XML)
Enterprise repos & SharePoint

MyWaypoint Refinery

01
Acquire
02
Validate
03
Normalize
04
Enrich
05
Govern
06
Search
07
Serve

Data Products

Enterprise search
API endpoints
Knowledge bases
Analytics-ready datasets
AI & RAG grounding layer
Governed data products

Any source type.

The refinery accepts structured and unstructured information from any source. Internal or external. Uploaded or crawled. Proprietary or public.

Documents

  • PDF
  • DOCX / TXT
  • Email exports
  • Scanned documents

Web Content

  • Websites
  • Web crawling
  • Public portals
  • Knowledge sites

Structured Data

  • CSV / XLSX
  • JSON / XML
  • Flat files
  • Data exports

Systems & APIs

  • REST APIs
  • Databases
  • Internal repos
  • Data feeds

Enterprise Sources

  • SharePoint
  • Business systems
  • Knowledge bases
  • Private collections
Stages 01–07

The Refinery Pipeline

Every source flows through the same seven-stage process.
The same governance. The same quality controls.
Every time.

01
Acquire
02
Validate
03
Normalize
04
Enrich
05
Govern
06
Search
07
Serve
01

Acquire

Information enters the platform from any source. The pipeline accepts structured and unstructured inputs without requiring prior cleanup.

  • Files, APIs, databases
  • Websites and web crawls
  • Documents and exports
  • SharePoint and enterprise repos
Input
02

Validate

All incoming data enters a secure quarantine environment before processing. Nothing proceeds until it passes quality and security checks.

  • Security scanning
  • Structure validation
  • Completeness checks
  • Quality gates and review
Quality
03

Normalize

Inconsistent formats, schemas, naming conventions, and field structures are transformed into unified, predictable records.

  • Schema standardization
  • Format conversion
  • Deduplication
  • Field alignment
Structure
04

Enrich

Normalized records are enhanced with metadata, extracted entities, classifications, and AI-assisted contextual signals.

  • Metadata generation
  • Entity extraction
  • Classification & tagging
  • Semantic context
Context
05

Govern

Every record carries its full history. Source, transformation, access controls, and ownership are recorded and maintained across the entire data lifecycle.

  • Lineage tracking
  • Provenance recording
  • Access controls
  • Auditability
Governance
06

Search

Refined records are indexed for full-text search, semantic similarity, and vector retrieval — making them accessible to both humans and AI systems.

  • Full-text indexing
  • Vector embeddings
  • Semantic search
  • RAG-ready retrieval
Discovery
07

Serve

Governed, searchable data is delivered through the channels your teams and systems already use.

  • APIs and dataset downloads
  • Search interfaces
  • Dashboards and analytics
  • AI and RAG systems
Delivery

Raw information vs refined information

The same source data. Before and after the refinery pipeline.

Raw information

Inconsistent formats

Different schemas across sources. Same field, different names.

Missing or incomplete metadata

No context about what records mean or where they came from.

No lineage or attribution

Impossible to explain where an answer came from.

Difficult to search

Full-text search returns irrelevant results. Semantic search fails.

Not consumable via API

Applications cannot reliably integrate the data.

Refined information

Standardized schema

Consistent fields, types, and naming across all records.

Enriched metadata

Source attribution, classification, entities, and context attached to every record.

Full lineage and attribution

Every output can be traced back to its source and transformation history.

Indexed and searchable

Full-text, semantic, and vector search all return accurate, relevant results.

API-accessible

Predictable structure, consistent delivery, ready for integration.

Trust requires traceability.

Every stage of the refinery is recorded. Not just the outputs — the transformations, the source, the quality checks, and the decisions made along the way.

This is what makes governance more than a policy. When an output can be explained, challenged, or audited at any point in its lifecycle, it becomes trustworthy. When it can't, it isn't.

Governance built into the pipeline means every data product is inherently explainable — not just correct.

Source Attribution

Every record is linked to the source it came from, when it was acquired, and who authorized its ingestion.

Lineage

Transformation history is preserved. Any output can be traced back through each stage to its origin.

Auditability

Processing logs are retained. Every transformation is reviewable, reproducible, and defensible.

Access Controls

Role-based permissions, data ownership, and access policies are enforced throughout the pipeline and at delivery.

AI Readiness

AI performs better on refined information.

Most AI problems are data problems. The model is rarely the issue.
The information it works with usually is.

Without governance

Retrieval pulls inconsistent or contradictory records

AI fills gaps with plausible but unsupported content

Outputs cannot be attributed to a verifiable source

Search returns irrelevant results that degrade responses

High hallucination risk, low explainability

With refinement

Retrieval surfaces validated, consistent records

Responses are grounded in governed, attributed information

Every output traces back to a verifiable source

Semantic search returns contextually accurate results

Lower hallucination risk, higher explainability

The refinery builds the data layer that makes AI grounding, RAG systems, and enterprise search reliable — not just technically possible.

What the refinery produces

Governed data, ready for the channels your teams and systems already use.

Search

Enterprise discovery experience. Full-text and semantic search across all refined content, with source attribution on every result.

API

Machine-accessible information. Structured, consistent API delivery for applications, integrations, and automated workflows.

Knowledge Bases

Governed internal information environments. Searchable, attributed knowledge built from your own documents and repositories.

Data Products

Curated, download-ready datasets with standardized fields, metadata enhancements, and documentation for analytical use.

Analytics

Reporting and operational visibility built on governed data. Consistent metrics because the underlying records are consistent.

AI & RAG

Grounded context for enterprise AI systems. Vector embeddings, semantic retrieval, and governed source attribution for reliable AI responses.

A refinery in production

The same process, applied to public data at scale.

MyCleanData is the public data marketplace and API platform powered by the MyWaypoint refinery. It demonstrates the complete pipeline applied to publicly available data — acquired, validated, normalized, enriched, governed, indexed, and served.

The same refinery model applies to private, internal, and custom data environments. Your documents. Your repositories. Your data.

Searchable catalog

Governed public datasets

API-accessible

Every dataset delivered via structured API

Full lineage

Every record attributed to its source

Same pipeline

Available for private enterprise environments

Ready to apply the refinery to your data?

If you're evaluating a data refinery deployment, a knowledge base build, or an AI-ready data pipeline, we'd like to discuss what that looks like for your environment.