Apache Airflow 2 to 3 DAG Migration Planner: Ruff AIR301 and AIR302 Scans, airflow.sdk Imports, schedule and logical_date Renames, Standard Provider Operators, No Direct Metadata DB Access, and a Staging Cutover Checklist
PpromptstudioยทOct 7, 2026
No rating
Plan an Airflow 2.x to 3.0 upgrade DAG by DAG: run Ruff's AIR rules to find removed and moved names, switch imports to airflow.sdk and the standard provider, replace schedule_interval, execution_date, days_ago, and SubDAGs, move task code off direct metadata database sessions to the REST API, and finish with a config update and staging cutover checklist.
Act as a data platform engineer who has upgraded production Apache Airflow deployments from 2.x to 3.0, maintains a few hundred DAGs, and has been paged for DAGs that vanished from the UI after an import error, for templates still using execution_date, and for tasks that opened ORM sessions against the metadata database.
Inputs:
- Current Airflow version, Python version, and how it is installed (pip constraints file, official Docker image, Helm chart, managed service): [CurrentVersion]
- DAG inventory: file names with the imports, DAG arguments, operators, and Jinja templates they use (paste code or excerpts): [DagInventory]
- Installed provider packages and versions: [Providers]
- Executor and deployment layout (scheduler, webserver, workers, triggerer): [Deployment]
- Any task code that touches the metadata database, airflow.settings.Session, or internal models: [DbAccessInTasks]
- Output format: [Format]
Generate:
1. A readiness gate: move to the latest 2.x release first (2.10 or the 2.11 bridge release), clear every deprecation warning in the scheduler and DAG processor logs, back up the metadata database, and confirm every provider in Providers has a release that supports Airflow 3.
2. A Ruff scan plan: the exact commands to run the AIR301 (removed in Airflow 3) and AIR302 (moved to a provider) rules over the dags folder with preview mode, how to review autofixes before accepting them, and how to read each rule code in the output.
3. A rename table built from DagInventory: from airflow.decorators and airflow.models.DAG to airflow.sdk (dag, task, DAG, Variable, TaskGroup), Dataset to Asset, schedule_interval and timetable to schedule, airflow.utils.dates.days_ago to a fixed pendulum start_date, and BashOperator, PythonOperator, and the core sensors to apache-airflow-providers-standard import paths.
4. A template and context fix list: execution_date, next_ds, prev_ds, tomorrow_ds, and yesterday_ds replaced with logical_date or data_interval_start and data_interval_end, with each changed template line shown before and after.
5. Behavior changes to decide per DAG: catchup now defaults to False, SubDAGs are removed so each one becomes a TaskGroup, SLA callbacks are removed, and the REST API moves from v1 to v2.
6. A DbAccessInTasks rewrite: task code may no longer open metadata database sessions, so each case moves to the Airflow REST API through the Python client, an Airflow Variable or Connection read through the SDK, or an external table you own.
7. A Deployment change list: the webserver is replaced by airflow api-server, the DAG processor runs as its own component, run airflow config update to review renamed options and add --fix only after review, then airflow db migrate.
8. A staging cutover checklist: parse every DAG with zero import errors, trigger one run per DAG, compare task counts and durations with the 2.x baseline, then a rollback note that names the database backup.
Constraints:
- Only rewrite code that appears in DagInventory; never invent DAG names, tasks, or connection IDs.
- Mark any provider version you cannot confirm as "check the provider changelog". No em dashes.