Perficient iconPerficientSep 29, 2026 ~6 min source read

Pattern-Based Compiler for Automating ADF-to-Databricks Pipeline Migration

A Perficient team converted 500+ orchestration activities across 37 Azure Data Factory instances by treating pipeline JSON as an AST and applying a two-phase compiler to normalize and rewrite activities for Databricks.

Building a Pattern-Based Engine to Migrate ADF Pipelines from Synapse to Databricks

Share this story

Send the public story page.

Useful takeaways from this story.

Treat ADF export JSON as a parseable AST and build a compiler that normalizes then migrates pipelines to avoid manual JSON editing.

Normalization requires four deterministic transformations so exported REST JSON matches ADF authoring format without breaking user-defined identifiers.

Auto-discover the migration blast radius by scanning linked services, datasets, and pipeline references to classify which activities touch Synapse.

# What this story is about Perficient needed to move workloads that target Azure Synapse Analytics into Databricks Unity Catalog across 37 Data Factory instances. Manually editing hundreds of pipelines and thousands of activities would be slow and error-prone. The team built a pattern-based engine that treats ADF pipeline JSON like an abstract syntax tree (AST), performs deterministic normalization, then programmatically rewrites activities that reference Synapse into Databricks equivalents.

# Architecture in two phases The framework runs as a two-phase compiler orchestrated by a single parameterized notebook:

  • Convert the REST API export form of pipeline JSON into the ADF authoring format expected by Git/CI.
  • Four explicit transformations: strip ARM envelope fields (id, type, etag), convert snake_case keys to camelCase, nest activity-specific fields under typeProperties, and wrap top-level fields under a properties block.
  • Once JSON is normalized, scan every activity, detect references to Synapse resources, classify the migration pattern, and generate replacement activities targeting Databricks.

# How the engine finds what to change

  • datasets/*.json: keep datasets whose linkedServiceName references a Synapse linked service.
  • pipelines/*.json: classify each activity against patterns that indicate read or write against Synapse.

This produces concrete mappings such as SYNAPSE_LS and SYNAPSE_DATASETS and cross-reference tables the migration engine uses to decide replacements.

# Activity classification patterns (examples) The engine checks five reference patterns per activity to decide how it touches Synapse. Examples described include:

  • activity.outputs[].referenceName ∈ Synapse datasets — copy sink (write to Synapse)
  • typeProperties.dataset.referenceName ∈ Synapse datasets — lookup or query against Synapse

Each match yields a migration target classification (read, write, query) and the engine chooses a corresponding Databricks replacement pattern.

# Secrets and connection parameters

# Benefits and practical outcome

  • Eliminated manual JSON edits across 2,200+ files and 500+ affected activities.
  • Consistent, repeatable transformations with deterministic rules for normalization and case handling.
  • Auto-discovered scope reduced missed dependencies and made classification explicit before code generation.

# When to consider this approach

  • You have many ADF factories or pipelines referencing Synapse and manual edits would be large-scale and error-prone.
  • You can export ADF artifacts and perform offline transformations in a controlled CI/CD or notebook-driven workflow.

# Short implementation notes

  • Implement normalization as a reversible, testable set of transformations.
  • Maintain an explicit list of user-defined containers to skip key conversion.

# Bottom line A compiler-like, pattern-driven engine lets teams migrate large ADF estates to Databricks efficiently by normalizing REST exports, auto-discovering dependencies, classifying activity patterns, and generating replacements that preserve credentials and pipeline semantics.

More context around this story.

dbt Meets Apache Flink: One Workflow for Data Engineers
Dzone iconDzoneSep 15, 2026

dbt Meets Apache Flink: One Workflow for Data Engineers

Data engineers managing batch SQL pipelines on Snowflake, BigQuery, and increasingly Databricks, and streaming pipelines on Apache Flink face a familiar problem: two toolchains, two skill sets, two CI/CD pipelines.dbt is now extending into stream processing. This post explains what that means in practice, why it matter

From weeks to minutes: The new agentic era of data pipelines
Google iconGoogleAug 31, 2026

From weeks to minutes: The new agentic era of data pipelines

Data pipelines are the backbone of the modern enterprise, yet a barrier to entry exists for orchestrating them, making this critical capability unavailable to many data professionals. Following our announcements at Google Cloud NEXT ’26 , where we introduced the Orchestration Pipelines framework, we are fundamentally c

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app