Read this once. Every example in the Playground works the same way — this is that way.
It runs in your browser
Every example has a ▶ Try it Live button. It opens a notebook in Google Colab — Google’s free, browser-based notebook service. You need a free Google account. That’s it. No install, no setup, no Spark on your laptop.
Hit Runtime → Run all and watch it go, top to bottom.
Spark itself gets pip installed in the first cell, inside Colab’s throwaway runtime. That’s why the first run takes a minute.
It’s your own sandbox
Edit anything. Break anything. Nothing you do touches these notes or my repo — you’re read-only there.
Want to keep your changes? File → Save a copy in Drive (or in GitHub), into your own account.
The data comes with it
You never download anything by hand. Each example builds its own data in the notebook, one of three ways:
1Inline
spark.createDataFrame(data, schema)
Rows typed straight into the notebook — nothing on disk. Ab Initio's Create Data. Used when the source doesn't matter, only the rows.
2Generated file
plain Python + Faker
A real file on disk, seeded so you get exactly the same rows I do. For read and write examples, where a file is the whole point.
3Downloaded
!wget … /datasets/
A fixed dataset pulled from my notebooks repo. For when the exact data matters — a deliberately messy file, or two tables that must line up for a join.
Files live under /content:
- inputs →
/content/ravi-writes/data/input/ - outputs →
/content/ravi-writes/data/output/
/contentis temporaryColab wipes it when the runtime recycles. Fine for a demo — don’t leave anything there you want to keep.
Peeking at the files
A Colab cell doubles as a terminal. Prefix a line with ! to run a unix command:
!ls -la /content/ravi-writes/data/output/customers_active
!head -5 /content/ravi-writes/data/input/customers.dat
!wc -l /content/ravi-writes/data/input/customers.datWorth doing. What Spark actually writes to disk is usually more interesting than you’d expect.