Do you remember those first few Ab Initio graphs?
The ones you must have designed in your enterprise project’s Dev environment, back in your initial days — just to get a feel for GDE and the core Ab Initio components like Co>Op and EME.
They were surely not production-ready graphs. But they were enough to teach you how data flows — and how to design and run the basic unit of every Ab Initio-based data pipeline, the .mp file.
Following the same approach, we are going to write a PySpark script that translates one of those early graphs — but this time in an open environment, since PySpark is open source, unlike Ab Initio.
The graph we’re translating is about as simple as it gets:
Read Text File → Filter → Reformat → Write Text File
One small caveat
Back in your initial days
Test data was something you made by hand — straight on the Unix box, in vi. A handful of rows, typed out, saved, and the graph run again.
Here
The notebook generates it for you with Python’s Faker library — seeded, so you get exactly the same rows every run, and the same ones I did.
Watching the data
And do you remember the Watchers?
The ones you put on each flow, to sneak into the record-stream and see what was actually coming out of a component.
We’ll be doing something very similar in PySpark — df.show(5). Drop it after any step, look at five rows, and carry on to the next one. Same habit, and the notebook uses it after each transformation.
Same instinct, better tooling. Open it and run it top to bottom.