Move Every Legacy Data Workload to BigQuery
Aviato migrates the data platforms enterprises are still running today: SAS, Informatica, IBM DataStage, SSIS, Talend, Oracle ODI, Oracle PL/SQL, Teradata BTEQ, Alteryx, Qlik, DataFlux, COBOL batch and unmaintained Python, onto Google BigQuery.
We do not hand-rewrite thousands of jobs. Aviato conversion accelerators translate the bulk of your existing logic into readable, version-controlled BigQuery SQL and Python, then our Google Cloud data engineers review, test and productionise the output.
What is legacy ETL migration to BigQuery?
Legacy ETL migration to BigQuery is the process of converting proprietary ETL jobs, stored procedures and analytics scripts into open, version-controlled SQL and Python running natively on Google BigQuery. Aviato does this with automated code conversion plus engineer-led validation, converting the majority of a codebase mechanically and reserving human effort for the genuinely complex jobs.
Reference Architecture
Every source platform runs through the same conversion and validation pipeline before it lands on BigQuery.
Sources
- SAS
- Talend
- Qlik
- Alteryx
- DataStage
- Informatica
- Oracle ODI
- SSIS
- Oracle PL/SQL
- Teradata BTEQ
- COBOL
- DataFlux
- Python
Conversion
- 1 Inventory
- 2 Automated conversion
- 3 Engineer review
- 4 Dual-run validation
Target
BigQuery
- Dataform
- Cloud Composer
- Dataplex
- Looker
Which Platforms Do You Migrate to BigQuery?
Each card shows what we take from the source system and what it becomes on Google Cloud.
SAS
Base SAS, SAS DI Studio, macros, PROC SQL
We convert
DATA steps, macro logic, PROC SQL, scheduled batch jobs
Lands as
Python and BigQuery SQL, orchestrated in Cloud Composer
Informatica
PowerCenter, IICS
We convert
Mappings, workflows, sessions, parameter files
Lands as
Dataform or dbt models on BigQuery
IBM DataStage
Parallel jobs and sequences
We convert
Parallel jobs, sequences, transformer logic
Lands as
BigQuery SQL and Dataform, orchestrated in Cloud Composer
SSIS
SQL Server Integration Services
We convert
Packages, control flow, data flow, script tasks
Lands as
Dataform on BigQuery, with Dataflow for streaming paths
Talend
Data Integration and Big Data jobs
We convert
Jobs, tMap transformations, contexts, routines
Lands as
Python and BigQuery SQL in Cloud Composer
Oracle Data Integrator
ODI
We convert
Interfaces, knowledge modules, packages
Lands as
BigQuery SQL and Dataform
Oracle PL/SQL
Database-resident logic
We convert
Packages, procedures, functions, triggers, cursors
Lands as
BigQuery stored procedures and scripted SQL
Teradata
BTEQ, FastLoad, MultiLoad, TPT
We convert
BTEQ scripts, utilities, macros, Teradata SQL dialect
Lands as
BigQuery SQL, with BigQuery Migration Service for bulk translation
Alteryx
Designer and Server
We convert
Workflows, macros, analytic apps
Lands as
Python and BigQuery SQL
Qlik
QlikView, Qlik Sense
We convert
Load scripts, data models, set analysis logic
Lands as
BigQuery as the modelling layer, Looker for semantics and reporting
DataFlux
Data quality and standardisation
We convert
Data quality rules, standardisation, match logic
Lands as
Dataplex data quality rules and Dataform assertions
COBOL
Mainframe batch, JCL, VSAM copybooks
We convert
Batch programs, copybook layouts, job control
Lands as
Python or Java services on Cloud Run, data landed in BigQuery
Legacy Python
Unversioned pandas scripts and notebooks
We convert
Ad-hoc scripts, notebooks, brittle local pipelines
Lands as
Tested, packaged Python using BigQuery DataFrames and Cloud Run
Not on the list? The same approach applies to most proprietary ETL and analytics dialects. Tell us what you run and we will confirm coverage in the assessment.
Why Enterprises Move Off Legacy ETL
The licence cost is the smallest part of the problem. The larger cost is the shrinking pool of engineers who can maintain these platforms, the change-freeze culture that builds up around jobs nobody fully understands, and the inability to plug modern AI into a closed system.
Petabyte-scale queries, no capacity planning
BigQuery scans petabytes in seconds by fanning a query across thousands of slots on demand. Analysis that legacy ETL had to pre-aggregate overnight, because the platform could not hold the full dataset, runs interactively against the raw grain. Queries your team currently does not attempt become routine.
A materially lower run-rate
Three costs disappear at once: the proprietary licence, the infrastructure sized for peak rather than actual use, and the specialist contractors needed to keep an unsupported platform running. BigQuery bills for compute actually consumed, so idle capacity stops being something you pay for.
Open, readable code
Transformation logic becomes SQL and Python in Git, reviewable by any data engineer, rather than a binary artefact locked inside a proprietary repository. Code review, testing and CI apply to your data pipelines the same way they apply to the rest of your engineering.
Native AI on the data where it sits
Gemini models, vector search and BigQuery ML run inside SQL, with no export step and no separate inference platform to maintain. Closed ETL platforms cannot offer this, which is why AI initiatives built on top of them stall at the data movement stage.
One governance layer
Dataplex handles cataloguing, lineage and data quality across the whole estate rather than per-tool. Data quality rules that were buried inside individual ETL jobs become visible, testable policy.
A hiring pool that actually exists
SQL and Python engineers are available and affordable. SAS, DataStage and DataFlux specialists increasingly are not, and the ones who remain command a premium to maintain a platform nobody is investing in.
How an Aviato BigQuery Migration Works
Five stages, with the legacy platform authoritative until the last one.
Discovery and code inventory
We take a read-only copy of your job definitions, scripts and scheduler metadata, then parse the estate to produce a complete inventory: every job, its dependencies, its runtime, its consumers, and how complex it will be to convert. Dead jobs get identified here, and on most estates a share of the code turns out to be unused.
Automated conversion
Aviato conversion accelerators translate the inventoried code into BigQuery SQL and Python. Output is generated as a Git repository mirroring the structure of the source estate, so it can be reviewed job by job against the original.
Engineer-led review and refactor
Automated output is a starting point, not a deliverable. Our Google Cloud data engineers review every converted job, refactor patterns that translated literally rather than idiomatically, and rebuild the small number of jobs that are better redesigned than converted.
Dual-run validation
Legacy and BigQuery pipelines run in parallel against the same inputs. We reconcile outputs row by row and column by column until results match, and keep both running side by side until your team is satisfied.
Cutover and decommission
Consumers are repointed, the legacy platform moves to read-only, and licences are wound down on an agreed schedule. Runbooks, lineage documentation and handover training go with it.
What You Get
- ✓ A complete, searchable inventory of your legacy estate with complexity scoring
- ✓ Converted code in your Git repository, structured to mirror the source
- ✓ Dual-run reconciliation reports proving output parity
- ✓ A BigQuery target architecture with Dataform, Cloud Composer and Dataplex
- ✓ A Looker semantic layer where reporting is in scope
- ✓ Runbooks, lineage documentation and engineer handover training
- ✓ A decommissioning plan mapped to your licence renewal dates
Proof
Lendlease
Aviato replaced Lendlease's legacy data tooling with a modern Google Cloud data platform, reducing the cost of running it by many multiples.
Book a Migration Assessment
Submit your details and a Principal Google Cloud Data Architect will inventory your legacy estate, identify what converts automatically, and give you a dated, wave-by-wave migration plan.
Prefer to schedule directly? Book a 20-minute scoping call with our Lead Architect ↗
Registered office: 59 Parry St, Newcastle NSW 2300
Mailing address: 6 Read St, Bronte NSW, 2024
Email: hello@aviato.consulting • Phone: +61 2 8359 9507
Frequently Asked Questions
Common questions on conversion coverage, validation, timelines and licence decommissioning.
Which legacy platforms can Aviato migrate to BigQuery?
Aviato migrates SAS, Informatica PowerCenter and IICS, IBM DataStage, SSIS, Talend, Oracle Data Integrator, Oracle PL/SQL, Teradata BTEQ, Alteryx, Qlik, DataFlux, COBOL mainframe batch and legacy Python to Google BigQuery. The same conversion approach extends to most other proprietary ETL and analytics dialects.
How long does a legacy ETL migration to BigQuery take?
Timelines depend on estate size and the number of downstream consumers, not on the source technology. A single subject area is typically a matter of weeks; a full enterprise estate runs in phased waves over several months. You get a dated, wave-by-wave plan at the end of the assessment rather than an estimate up front.
Do you rewrite our code by hand?
No. We use automated conversion to translate the bulk of the estate, then our engineers review, refactor and test the output. Hand-rewriting an enterprise ETL estate is slow, expensive, and introduces defects the original code did not have.
How do we know the migrated pipelines produce the same results?
Dual-run validation. Legacy and BigQuery pipelines run in parallel on the same inputs, and we reconcile outputs row by row and column by column. Cutover happens only once results match and your team signs off.
Will the business be disrupted during the migration?
No. The legacy platform stays live and authoritative throughout. Reporting, dashboards and downstream systems continue running on it until each wave has been validated on BigQuery and formally cut over.
What happens to our SAS or Informatica licences?
They stay in place until the workloads that depend on them are cut over. We build the decommissioning plan around your renewal dates so licence savings land as early as possible.
What does the converted code actually look like?
Readable, version-controlled BigQuery SQL and Python in your own Git repository, organised to mirror the structure of your source estate so every converted job traces back to its original. There is no Aviato runtime, no proprietary wrapper, and no ongoing dependency on us.
Can you migrate the reporting layer as well as the pipelines?
Yes. Qlik and other reporting layers are migrated to Looker, with the semantic model rebuilt on BigQuery. This is normally scoped as a separate wave once the underlying data is in place.
How do you handle data quality rules built in DataFlux or inside ETL jobs?
Data quality logic is extracted and rebuilt as Dataplex data quality rules and Dataform assertions, so it runs as part of the pipeline and is visible in the same governance layer as the rest of your estate.
Is Google Cloud migration funding available?
Often, yes. As a Google Cloud Premier Partner, Aviato applies for migration funding programmes on your behalf. Eligibility depends on workload size and is confirmed in writing before work starts.
What access do you need to run the assessment?
Read-only access to job definitions, scripts and scheduler metadata. We do not need access to production data or customer records to inventory and convert an estate.
What if our estate includes something not on your list?
Tell us what it is. The conversion approach is dialect-driven rather than vendor-driven, so coverage extends well beyond the platforms listed here. The assessment confirms what converts automatically and what needs redesign.
Already on Snowflake? See our Snowflake to BigQuery cost guarantee
Talk to an architect who has done this before.
Bring your current setup and the outcome you need. You will get a view on the approach, the risks and roughly what it costs.
Straight to a senior GCP architect. No SDR, no slide deck.
Not ready to talk? See how we migrated Hapana off AWS →
Or call +61 2 8359 9507 · Hello@aviato.consulting