# What this story is about Perficient needed to move workloads that target Azure Synapse Analytics into Databricks Unity Catalog across 37 Data Factory instances. Manually editing hundreds of pipelines and thousands of activities would be slow and error-prone. The team built a pattern-based engine that treats ADF pipeline JSON like an abstract syntax tree (AST), performs deterministic normalization, then programmatically rewrites activities that reference Synapse into Databricks equivalents.
# Architecture in two phases The framework runs as a two-phase compiler orchestrated by a single parameterized notebook:
- Convert the REST API export form of pipeline JSON into the ADF authoring format expected by Git/CI.
- Four explicit transformations: strip ARM envelope fields (id, type, etag), convert snake_case keys to camelCase, nest activity-specific fields under typeProperties, and wrap top-level fields under a properties block.
- Once JSON is normalized, scan every activity, detect references to Synapse resources, classify the migration pattern, and generate replacement activities targeting Databricks.
# How the engine finds what to change
- datasets/*.json: keep datasets whose linkedServiceName references a Synapse linked service.
- pipelines/*.json: classify each activity against patterns that indicate read or write against Synapse.
This produces concrete mappings such as SYNAPSE_LS and SYNAPSE_DATASETS and cross-reference tables the migration engine uses to decide replacements.
# Activity classification patterns (examples) The engine checks five reference patterns per activity to decide how it touches Synapse. Examples described include:
- activity.outputs[].referenceName ∈ Synapse datasets — copy sink (write to Synapse)
- typeProperties.dataset.referenceName ∈ Synapse datasets — lookup or query against Synapse
Each match yields a migration target classification (read, write, query) and the engine chooses a corresponding Databricks replacement pattern.
# Secrets and connection parameters
# Benefits and practical outcome
- Eliminated manual JSON edits across 2,200+ files and 500+ affected activities.
- Consistent, repeatable transformations with deterministic rules for normalization and case handling.
- Auto-discovered scope reduced missed dependencies and made classification explicit before code generation.
# When to consider this approach
- You have many ADF factories or pipelines referencing Synapse and manual edits would be large-scale and error-prone.
- You can export ADF artifacts and perform offline transformations in a controlled CI/CD or notebook-driven workflow.
# Short implementation notes
- Implement normalization as a reversible, testable set of transformations.
- Maintain an explicit list of user-defined containers to skip key conversion.
# Bottom line A compiler-like, pattern-driven engine lets teams migrate large ADF estates to Databricks efficiently by normalizing REST exports, auto-discovering dependencies, classifying activity patterns, and generating replacements that preserve credentials and pipeline semantics.