Hao Sun

Data science master’s student at UNSW, working toward data engineering. I like the part of data work most people skip: getting it collected, cleaned, stored and moving on a schedule, so the analysis on top can be trusted.

Education

2025–2027

Master of Data Science

UNSW Sydney, finishing July 2027

2019–2023

Bachelor of Computing Science (Honours)

University of Technology Sydney, majoring in data analytics and AI

Projects

2026

Sydney bus reliability

Python, DuckDB, Parquet, Streamlit

A pipeline that polls Transport for NSW’s live GTFS-Realtime feed every two minutes and works out when Sydney buses actually arrive. The first 20 hours produced about 490,000 arrivals.

What I found and how it works
17.8% early 71.7% on time 10.5% late

About 72% of buses fall inside the official window: no more than 1 minute early or 5 minutes late. That window makes early buses look more common than late ones. With the same 3-minute cut-off on both sides, buses are late about four times as often as they are early.

The feed only gives predictions, never actual arrivals, so I take the last prediction made within 150 seconds of the stop as the arrival time. Raw snapshots are kept untouched, so the whole dataset can be rebuilt if that logic changes. A quality report checks for collection gaps and how close to arrival the predictions were made (median 42 seconds).

Figures are provisional while collection continues.

2026

Warehouse pipeline with dbt and Airflow

Snowflake, dbt Core, Apache Airflow, Astronomer Cosmos

Turns Snowflake’s TPC-H order data into a tested fact table, scheduled to run daily in Airflow.

How it’s built

Models are layered: staging views clean and rename the raw tables, intermediate models join orders to line items and calculate discounted amounts, and a mart model produces fct_orders, ready for BI. dbt tests, including referential checks, run on every build.

Cosmos runs each dbt model as its own Airflow task, so a failure shows exactly where it happened. It connects to Snowflake with key-pair authentication and runs locally in Docker through the Astronomer CLI.

2026

Sprout Isle

React, TypeScript, GitHub Actions

A habit tracker where each habit you keep grows a pixel island. It includes an Insights tab that sets habits against a daily mood and energy check-in.

The analytics side

Insights has a daily completion chart with a 7-day rolling mean, a habit-by-day grid, and phi correlations between each pair of habits over 60 days. The mood view estimates how much higher you rate your mood on days you did each habit, with 95% confidence intervals, and plots completion against mood with Pearson’s r.

Every chart has a table view, and the data exports to CSV with one row per habit per day, ready to load into pandas. Each push runs the tests and deploys to GitHub Pages.

2026

Golf performance tracker

Python, SQLite

A command-line app for logging rounds hole by hole and following scores against par over time. I built it to practise relational schema design and SQL.

Experience

2025

Data entry supervisor

Randstad, for the 2025 federal election

Supervised the team entering voter and election data. Kept entry accurate and confidential, and managed the workflow so the team met fixed election deadlines.

2024

Scrutiny officer

NSW Electoral Commission

Entered polling data in the EMA system and cross-checked it against voter records, keeping the count accurate under tight election timelines.

2021

Web and business development

EWE Group, a fulfilment and 3PL company

Built, designed and maintained the company website. SEO changes lifted clicks and visits by 25%. Researched start-ups likely to need fulfilment services.

Other work

2020–2025

Front end team member, Woolworths

2019

BBBS education program, Hunter Valley Grammar School

Skills

Code
Python, SQL, R, TypeScript
Data
dbt, Snowflake, Airflow, DuckDB, Parquet, SQLite, pandas
Analysis
Tableau, Power BI, Streamlit
Tools
Git, GitHub Actions, Docker, React
Languages
English and Chinese, both fluent