# SchemaVortex > SchemaVortex is a data lakehouse platform that extracts data from operational databases, keeps a versioned history of it, and serves the result as access-controlled SQL views. It runs inside the customer's own Azure subscription rather than as a service data is sent to. Published by Fizzcode Ltd., Budapest, Hungary. The platform is configured through a browser instead of by writing pipeline code, and after handover the customer's own team operates it: no standing vendor credentials remain in the customer's tenant. ## Key facts - Storage: Apache Parquet files in the customer's own Azure Data Lake Storage. This store is called the Vault, and its column set only ever grows: existing columns are never altered or removed. When a source changes shape, a person updates the Vault table with one edit and the new shape is stored as a new versioned column, so the columns a report already reads do not change. - Query surface: Mart views, which are Azure Synapse Analytics serverless SQL views over the Vault. Readable from Power BI, Excel, Tableau, Qlik, Python or any SQL client. Serverless means billing per query, with no cluster to size. - History: every Vault table is exposed as a Latest table holding the current state and a History table behind it, so point-in-time questions and audits are answerable. Versions are added rather than overwritten, and how far back the history reaches is a retention setting the customer controls. - Retention and point-in-time revert: the history changes only at its two ends, and the data between them never changes. Retention removes the loads older than the period set per source, 24 months unless set otherwise. At the other end, after a faulty load, a source's tables can be put back to an earlier point in time. The platform rebuilds them from the copies of the loads kept in the customer's storage, with no downtime and no new read from the source. The loads after the chosen moment are removed for good, and a table kept as a full snapshot cannot be reverted. - Sources: SQL Server, PostgreSQL, MySQL, MariaDB, Oracle, SAP HANA, IBM i (DB2 for i, formerly AS/400) and Infor Data Lake, plus the ERP, CRM and line-of-business systems built on them, are extracted directly. REST APIs and anything else without a built-in connector come in through the Producer SDK. Registering an Origin discovers its tables, columns, types and keys automatically. - Bring Your Own Data: Excel, CSV and Parquet files can be imported as tables, and the Producer SDK (a .NET library) pushes data from any other system. Imported tables go through the same approval, masks and Vault history as database sources. - Governance: every column passes four gates in order: AI Gate (hiding a column from AI), Governance Gate (classifications a Data Steward puts on data), Approval Gate (table and column approval) and Delivery Gate (a second set of classifications on Vault columns). A classification labels a kind of sensitive data, and a reader must be cleared for all of a column's classifications to read its value. A Data Steward holds the AI Gate and the Governance Gate, a Data Warden the Approval Gate and everything a Data Steward holds, and a Vault Manager the Delivery Gate. Every AI session reads classified data and data hidden from AI as NULL. Access is decided as each query runs: a query returns the values its reader may read and NULL or a masked value for the rest. Every approval, gate change and membership change is audited. - Change control (four eyes): Mart views, Vault tables and Sandbox views can be changed through proposals, called Mart Plans, Vault drafts and Sandbox drafts. A proposal is validated against the current state, then published or approved by a person with the right permission, under their own name, which the audit trail records. For people the proposal workflow is optional: a person with the right permission can change a Mart view or a Vault table directly, and a deployment can require that every Mart change goes through a Mart Plan. The AI Assistant can write a proposal and can never apply one. - The Mart: the reporting model as SQL views over the Vault and over other Mart views, grouped into Mart databases. A direct change is compiled against Synapse before it is saved, and a Mart Plan is compiled as a whole when it is published: one view that fails stops the whole publish before anything changes. Every change is recorded with its author and time, and an earlier version of a view, or every view as it stood at a chosen moment, can be restored through a Mart Plan. A view either runs its SQL when queried or stores its result as Parquet in the customer's storage, rebuilt when the data behind it changes. A view computed from protected columns stores its result only after a Data Warden approves them. Further Query endpoints serve the same views over the same data, so heavy reporting load can be kept apart. - Sandbox: every user can get a personal schema of SQL views over the Vault and the Mart. Sandbox views run live and store no copy, return only what the person querying may read, and are never read by a Mart view. Colleagues can propose changes as Sandbox drafts, which the view's owner approves. - Lineage: view-level and column-level, resolved from the views the platform actually runs rather than from maintained documentation. It traces backwards to source columns and forwards for impact analysis. - Catalog: built into the platform, spanning the Vault and the Mart, with a SQL editor, and notes and tags on tables, views and Vault columns. Nothing to deploy alongside the lake or rescan on a schedule. - AI: AI Chat answers questions about the catalog and writes SQL, reading metadata only, on the customer's own Azure OpenAI. The AI Assistant lets a coding agent on a person's own computer (Claude Code, Codex, Gemini CLI, GitHub Copilot) connect with a personal key and work under that person's permissions: it reads the catalog, proposes Mart Plans, Vault drafts, Sandbox drafts and developer notes that a person approves, and where data access is switched on it reads data through a temporary one-hour login on which every masked column and every column hidden from AI comes back empty. Every act is recorded in the AI Assistant log. - Leaving: the Parquet stays in customer-owned storage in an open format, readable by any Parquet-aware engine. There is no proprietary catalog to export from. ## Questions and answers The questions and answers of the FAQ page, in the same wording. ### The platform **What is SchemaVortex?** SchemaVortex is a data lakehouse platform that extracts data from operational databases, keeps a versioned history of it, and serves the result as access-controlled SQL views. It runs inside the customer's own Azure subscription rather than as a service data is sent to, and it is configured through a browser instead of by writing pipeline code. **What is SchemaVortex not?** It is not a reporting or visualization tool, and it does not replace the source systems it reads. It prepares governed, historized data and serves it over standard SQL; the charts and dashboards are still built in Power BI, Excel or whatever the business already uses. **Who operates SchemaVortex once it is set up?** The customer's own team. The platform is self-managing and runs in the customer's own Azure subscription, and after handover no standing vendor credentials remain in that tenant. If a support engineer ever needs access, it is through a temporary, scoped account the customer creates and revokes. **How is this different from building the pipelines ourselves?** Hand-built ingestion means writing and then maintaining extraction, historization, schema-change handling, access control, lineage and a catalog as separate pieces of work, usually with specialist engineers. SchemaVortex ships those as one platform: an Origin is registered, its tables are discovered automatically, and what comes out is already historized, access-controlled and traceable. ### Data and storage **Where does our data physically live?** In the customer's own Azure subscription. Extraction reads the sources in place, the Vault is written to the customer's own Azure Data Lake Storage, and serving runs in the same subscription, billed to the customer directly by Microsoft. There is no vendor cloud for the data to travel to. **What format is the data stored in?** Apache Parquet. The Vault is a set of Parquet files in the customer's own Azure Data Lake Storage, readable by Power BI, Excel, Spark and any other Parquet-aware engine, with or without SchemaVortex in the picture. **How is the data queried?** Through Mart views, which are Azure Synapse Analytics serverless SQL views over the Vault. Anything that speaks SQL can read them: Power BI, Excel, Tableau, Qlik, Python or any SQL client. Serverless means the query engine is billed per query, with no cluster to size or keep running. **Is historical data kept, and can we query the past?** Yes. Versions are added rather than overwritten, so a record's earlier states stay available for as long as you choose to keep them. Each Vault table is exposed both as a Latest table, holding the current state for everyday reporting, and as a History table for audits and point-in-time questions. A faulty load can be reverted: the source's tables go back to a moment before it, and the history up to that moment stays. **How long is data kept?** For as long as you choose to keep it. The current state of every table stays available; how far back the history behind it reaches is a retention setting you control. Keeping personal data no longer than necessary is a GDPR requirement, which is why the depth of history is a setting rather than a fixed promise. **What happens to our data if we stop using SchemaVortex?** It stays where it already is. The Parquet files sit in storage the customer owns, in an open format, and remain readable by any Parquet-aware tool. There is no proprietary catalog holding the data, and nothing has to be exported or migrated out. ### Connecting sources **Which source systems can be connected?** Built-in connectors cover SQL Server, PostgreSQL, MySQL, MariaDB, Oracle, SAP HANA, IBM i and Infor Data Lake, along with the ERP, CRM and line-of-business systems built on them, and the list continues to grow. Excel, CSV and Parquet files are imported as tables. REST APIs and anything else without a built-in connector are pushed in through the Producer SDK. **Do we have to map the schema by hand?** No. Registering an Origin makes SchemaVortex discover it: the tables, the columns, their types and their keys are read from the source itself. From there it is a matter of selecting which Source tables to track and choosing an Extraction strategy for each. **What happens when a source table changes shape?** The Vault's column set only ever grows. When a source adds a column or changes a column's type, the Vault table is updated with one edit, and the new shape is stored as a new versioned column. The original stays as it was and no longer loads. Records from before a new column existed show it as empty. Because existing columns are never altered or removed, the columns a report already reads do not change underneath it. ### Governance and traceability **How is access to sensitive data controlled?** A Data Steward classifies the data that needs it. A classification labels a kind of sensitive data, and people and groups are cleared for it. A query returns classified data only to people cleared for every classification it carries and to the Data Stewards and Data Wardens of its source system. For everyone else, the query returns NULL or a masked value. Access is decided as each query runs, so there is no masked copy of the data to build or keep in sync. No table enters the Vault until a Data Warden approves it, and a Vault Manager can add a second set of classifications to what the Vault delivers. Every approval, gate change and membership change is recorded with who made it and when. **What does lineage cover?** Views and columns. Any column in a Mart view can be traced back through every view to the source column behind it, and the same links run forward, so the Mart views a change would affect can be listed before the change is made. Because the platform resolves every view itself, lineage is read from what is actually running rather than from documentation someone maintains. **Is a separate data catalog needed?** No, the catalog is part of the platform. One schema browser spans the Vault and the Mart, there is a built-in SQL editor for writing and editing views, lineage opens from any table or column, and tables, views and Vault columns carry notes and tags. Nothing has to be deployed alongside the lake or rescanned on a schedule. **Does the AI see our data?** AI Chat does not: it reads the schema and the metadata, never a record, and it runs on your own Azure OpenAI. The AI Assistant lets a coding agent on a user's own computer read the catalog and propose changes. Where the deployment switches data access on, the agent reads data through a temporary login. On that login every masked column and every column hidden from AI comes back empty, even where the user may read it. The login is valid for one hour, and every act is recorded in the AI Assistant log under the user's name. **Can one person change a Mart view or a Vault table alone?** Yes, if they hold the permission to apply the change. A deployment can require every Mart change to go through a Mart Plan. Changes can also be made as proposals. A Mart Plan, a Vault draft or a Sandbox draft describes the change, the platform checks it against the current state of the Mart or the Vault, and a person who holds the permission to apply it publishes or approves it under their own name, which the audit trail records. In the Mart the split is built in: a Mart Manager authors plans, and only a Mart Publisher can publish one. Whether the author and the approver have to be different people is otherwise decided by how the permissions are granted in the deployment. **What happens when two people change the same view?** Every entry in a Mart Plan remembers the version of the view it was written against. If the view has changed in the meantime, the entry shows as a conflict, the plan cannot be published until the conflict is resolved, and the check runs again during the publish itself, so a change is never overwritten silently. The same rule protects a Vault draft: if the table's columns changed after the draft was written, it has to be opened and saved again before it can be applied. **Can the AI Assistant make changes on its own?** No. It can write a Mart Plan, a Vault draft or a Sandbox draft, and it can never publish, approve or discard one; a person does that, under their own name. Every act the assistant performs is recorded in the AI Assistant log, under the person who was using it. ## Pages - [Frequently asked questions](https://schemavortex.com/faq.html): direct answers on what the platform is and is not, where data physically lives, formats and query engine, source systems, schema changes, governance, lineage, and what happens when a customer stops using it. - [Home](https://schemavortex.com/): overview, platform summary, and the path from source to query. - [Storyboard](https://schemavortex.com/storyboard.html): animated cards, each one a problem and how it gets solved, in nine chapters from connecting a source to the audit trail. - [Governance](https://schemavortex.com/governance.html): classifications, the four gates, the roles and the audit. - [Four eyes](https://schemavortex.com/four-eyes.html): Mart Plans, Vault drafts and Sandbox drafts: proposals checked against the current state and applied by an entitled person under their own name. - [The Vault](https://schemavortex.com/vault.html): how history is retained and how schema changes in a source are absorbed. - [Lineage](https://schemavortex.com/lineage.html): tracing a column to its source and impact analysis before a change. - [The Catalog](https://schemavortex.com/catalog.html): schema browser, built-in SQL editor, annotations. - [Open by design](https://schemavortex.com/open-by-design.html): customer-owned Azure subscription, open Parquet, no data lock-in. - [Connect anything](https://schemavortex.com/connect.html): supported sources and the tools that consume the result. - [Bring Your Own Data](https://schemavortex.com/bring-your-own-data.html): importing Excel, CSV and Parquet files, and the Producer SDK for pushing data from any system. - [Retention and point-in-time revert](https://schemavortex.com/point-in-time-revert.html): the two ends of the history: retention of the oldest loads, and putting a source's tables back to an earlier point in time after a faulty load. - [The Mart](https://schemavortex.com/mart.html): the reporting model as SQL views over the Vault: checked before it goes live, every change on record, restores through a Mart Plan, and stored results rebuilt when the data changes. - [The Sandbox](https://schemavortex.com/sandbox.html): a personal SQL space for every user on the governed data, and drafts that the owner approves. - [AI](https://schemavortex.com/ai.html): AI Chat, metadata only on the customer's own Azure OpenAI, and the AI Assistant for coding agents, with bounded data access and the AI Gate. ## Listings - [Microsoft Marketplace](https://marketplace.microsoft.com/en-us/product/fizzcodekorlatoltfelelossegutarsasag1620904005110.schemavortex): the official listing on the Microsoft commercial marketplace. The offer is contact-based, not click-to-deploy. ## Other languages - [Magyar](https://schemavortex.com/hu/): the full site in Hungarian. - [Magyar llms.txt](https://schemavortex.com/hu/llms.txt): this file in Hungarian, covering the pages under /hu/. ## Legal - [Privacy notice](https://schemavortex.com/privacy.html): data controller details and GDPR information.