Large-Scale Data Engineering
Package Network Fingerprint & Temperature Intelligence
Large-scale Azure Databricks/PySpark pipeline that derives package-network fingerprint and temperature intelligence from multi-billion-row operational source data.
- Native PySpark transformations with partition-aware joins, window operations, column pruning, and selective caching—no Python UDFs
- Approximately 4.5 billion source rows producing ~200 million Fingerprint records and ~22 million Temperature outputs with a normalized score from cold 0 to hot 1
- Row-count reconciliation, duplicate detection, and null/range validation