From raw taxi files to a managed Spark job
An independent project that moves NYC taxi data through local Spark exploration, cloud storage, and a parameterised Dataproc job.
Explore the decisionsData engineering · Analytics · BI
I'm Aman Gupta. This is where I turn data engineering work into clear system designs, practical guides, and honest lessons.
Selected work
Case studies and reference architectures focused on constraints, trade-offs, and repair—not tool lists.
An independent project that moves NYC taxi data through local Spark exploration, cloud storage, and a parameterised Dataproc job.
Explore the decisionsA reference architecture for measuring freshness, completeness, and publication state—not just whether scheduled tasks ran.
Explore the decisionsSelected writing
Notes about building pipelines, defining trustworthy metrics, and operating data products in the real world.
A dlt + Iceberg demo showing “metadata over compute”: table history lives in portable metadata files, not in any single catalog or engine.
Ontology engineering encodes expert knowledge into machine-readable structures. It is "data modeling for meaning," ensuring all systems see data the same way.

A non-linear path
My engineering path changed disciplines, but not its underlying concern: understand the system, respect its constraints, and make its behaviour legible.
More about me