Skip to content
SCHEMAVORTEX
Home

Questions

Frequently asked questions

What SchemaVortex is, where your data lives, who runs it, and what happens if you stop using it.

The platform

What is SchemaVortex?

SchemaVortex is a data lakehouse platform that extracts data from operational databases, keeps a versioned history of it, and serves the result as access-controlled SQL views. It runs inside the customer's own Azure subscription rather than as a service data is sent to, and it is configured through a browser instead of by writing pipeline code.

What is SchemaVortex not?

It is not a reporting or visualization tool, and it does not replace the source systems it reads. It prepares governed, historized data and serves it over standard SQL; the charts and dashboards are still built in Power BI, Excel or whatever the business already uses.

Who operates SchemaVortex once it is set up?

The customer's own team. The platform is self-managing and runs in the customer's own Azure subscription, and after handover no standing vendor credentials remain in that tenant. If a support engineer ever needs access, it is through a temporary, scoped account the customer creates and revokes.

How is this different from building the pipelines ourselves?

Hand-built ingestion means writing and then maintaining extraction, historization, schema-change handling, access control, lineage and a catalog as separate pieces of work, usually with specialist engineers. SchemaVortex ships those as one platform: an Origin is registered, its tables are discovered automatically, and what comes out is already historized, access-controlled and traceable.

Data and storage

Where does our data physically live?

In the customer's own Azure subscription. Extraction reads the sources in place, the Vault is written to the customer's own Azure Data Lake Storage, and serving runs in the same subscription, billed to the customer directly by Microsoft. There is no vendor cloud for the data to travel to.

What format is the data stored in?

Apache Parquet. The Vault is a set of Parquet files in the customer's own Azure Data Lake Storage, readable by Power BI, Excel, Spark and any other Parquet-aware engine, with or without SchemaVortex in the picture.

How is the data queried?

Through Mart views, which are Azure Synapse Analytics serverless SQL views over the Vault. Anything that speaks SQL can read them: Power BI, Excel, Tableau, Qlik, Python or any SQL client. Serverless means the query engine is billed per query, with no cluster to size or keep running.

Is historical data kept, and can we query the past?

Yes. Versions are added rather than overwritten, so a record's earlier states stay available for as long as you choose to keep them. Each Vault table is exposed both as a Latest table, holding the current state for everyday reporting, and as a History table for audits and point-in-time questions. A faulty load can be reverted: the source's tables go back to a moment before it, and the history up to that moment stays.

How long is data kept?

For as long as you choose to keep it. The current state of every table stays available; how far back the history behind it reaches is a retention setting you control. Keeping personal data no longer than necessary is a GDPR requirement, which is why the depth of history is a setting rather than a fixed promise.

What happens to our data if we stop using SchemaVortex?

It stays where it already is. The Parquet files sit in storage the customer owns, in an open format, and remain readable by any Parquet-aware tool. There is no proprietary catalog holding the data, and nothing has to be exported or migrated out.

Connecting sources

Which source systems can be connected?

Built-in connectors cover SQL Server, PostgreSQL, MySQL, MariaDB, Oracle, SAP HANA, IBM i and Infor Data Lake, along with the ERP, CRM and line-of-business systems built on them, and the list continues to grow. Excel, CSV and Parquet files are imported as tables. REST APIs and anything else without a built-in connector are pushed in through the Producer SDK.

Do we have to map the schema by hand?

No. Registering an Origin makes SchemaVortex discover it: the tables, the columns, their types and their keys are read from the source itself. From there it is a matter of selecting which Source tables to track and choosing an Extraction strategy for each.

What happens when a source table changes shape?

The Vault's column set only ever grows. When a source adds a column or changes a column's type, the Vault table is updated with one edit, and the new shape is stored as a new versioned column. The original stays as it was and no longer loads. Records from before a new column existed show it as empty. Because existing columns are never altered or removed, the columns a report already reads do not change underneath it.

Governance and traceability

How is access to sensitive data controlled?

A Data Steward classifies the data that needs it. A classification labels a kind of sensitive data, and people and groups are cleared for it. A query returns classified data only to people cleared for every classification it carries and to the Data Stewards and Data Wardens of its source system. For everyone else, the query returns NULL or a masked value. Access is decided as each query runs, so there is no masked copy of the data to build or keep in sync. No table enters the Vault until a Data Warden approves it, and a Vault Manager can add a second set of classifications to what the Vault delivers. Every approval, gate change and membership change is recorded with who made it and when.

What does lineage cover?

Views and columns. Any column in a Mart view can be traced back through every view to the source column behind it, and the same links run forward, so the Mart views a change would affect can be listed before the change is made. Because the platform resolves every view itself, lineage is read from what is actually running rather than from documentation someone maintains.

Is a separate data catalog needed?

No, the catalog is part of the platform. One schema browser spans the Vault and the Mart, there is a built-in SQL editor for writing and editing views, lineage opens from any table or column, and tables, views and Vault columns carry notes and tags. Nothing has to be deployed alongside the lake or rescanned on a schedule.

Does the AI see our data?

AI Chat does not: it reads the schema and the metadata, never a record, and it runs on your own Azure OpenAI. The AI Assistant lets a coding agent on a user's own computer read the catalog and propose changes. Where the deployment switches data access on, the agent reads data through a temporary login. On that login every masked column and every column hidden from AI comes back empty, even where the user may read it. The login is valid for one hour, and every act is recorded in the AI Assistant log under the user's name.

Can one person change a Mart view or a Vault table alone?

Yes, if they hold the permission to apply the change. A deployment can require every Mart change to go through a Mart Plan. Changes can also be made as proposals. A Mart Plan, a Vault draft or a Sandbox draft describes the change, the platform checks it against the current state of the Mart or the Vault, and a person who holds the permission to apply it publishes or approves it under their own name, which the audit trail records. In the Mart the split is built in: a Mart Manager authors plans, and only a Mart Publisher can publish one. Whether the author and the approver have to be different people is otherwise decided by how the permissions are granted in the deployment.

What happens when two people change the same view?

Every entry in a Mart Plan remembers the version of the view it was written against. If the view has changed in the meantime, the entry shows as a conflict, the plan cannot be published until the conflict is resolved, and the check runs again during the publish itself, so a change is never overwritten silently. The same rule protects a Vault draft: if the table's columns changed after the draft was written, it has to be opened and saved again before it can be applied.

Can the AI Assistant make changes on its own?

No. It can write a Mart Plan, a Vault draft or a Sandbox draft, and it can never publish, approve or discard one; a person does that, under their own name. Every act the assistant performs is recorded in the AI Assistant log, under the person who was using it.