Ab Initio Graph to PySpark Script
Ab Initio Graph — this has been my weapon for years to make data usable. Whether it’s creating simple automations in dev or implementing production-grade ETL workflows, Ab Initio Graph + Unix Utils + Shell Scripting + Basic SQL has done the job for me.
In the world of Ab Initio, we funnily say everything is a Graph.
PySpark Script — this is something I’m adding to my skillset by experimenting with it and mapping it to my Ab Initio experience. I can clearly see PySpark + Advanced SQL + Python becoming the popular new core.
The same pipeline, both ways
Let’s see how a simple Ab Initio Graph maps to a PySpark script:
from pyspark.sql import SparkSession
from pyspark.sql import functions as F
from pyspark.sql import types as T
# Spark Entry Point
spark = SparkSession.builder.appName("customer_load").getOrCreate()URL="/data/input.dat"string("\n");record-separator="\n"
field-separator="|"
skip-header-record=Falserecord
string(int) name = NULL;
string(int) status = NULL;
end;schema = T.StructType([
T.StructField("name", T.StringType()),
T.StructField("status", T.StringType()),
])
df_read = (
spark.read
.format("csv")
.options(lineSep="\n", sep="|", header=False)
.schema(schema)
.load("/data/input.dat")
)status == "A"df_filtered = df_read.filter(F.col("status") == "A")out :: reformat(in) =
begin
out.name :: string_lrtrim(in.name);
out.status :: in.status;
end;df_write = df_filtered.select(
F.trim(F.col("name")).alias("name"),
F.col("status"),
)record-separator="\n"
field-separator="\x01"
provide-header-record=Truestring("\n");URL="/data/output.dat"(
df_write.write
.format("csv")
.options(lineSep="\n", sep="\x01", header=True)
.mode("overwrite")
.save("/data/output")
)# Stop SparkSession
spark.stop()Note: the left pane is a simplified drawing of the graph, not the .mp file’s contents — an .mp is a proprietary text format you open in GDE, where it looks roughly like this.
This exact pipeline runs. Open it in Google Colab — nothing to install, nothing to download. It generates its own input file, walks the four stages, and shows you what actually lands on disk.
Designed, or optimised?
There are surely significant differences between how an Ab Initio graph is designed and executed and how a PySpark script is written and executed. To begin with, one key difference:
Ab Initio Graph typically executes the way you designed it.
You choose the layout. An 8-way parallel component means 8 parallel component processes. You place the Filter By Expression, you decide where unused fields drop, you decide when to Broadcast a data-flow before Join, you decide where a phase break goes.
PySpark doesn’t work like that.
You describe the result you want. Catalyst decides how to get it.
It reorders your filters. It prunes columns you never selected. It picks the join strategy, broadcast or sort-merge, based on size estimates. With default settings, it even decides how many partitions you get.
Neither side is absolute
That said, Ab Initio does have optimizations beyond graph design, like its Optimizing Compiler and Dynamic Parallelism, and you can steer Spark too: repartition, broadcast hints, shuffle partition counts.