Skip to content

Navigation Menu

Sign in
Sign up
@Smars-Bin-Hu
Smars-Bin-Hu
Follow

Smars Hu Smars-Bin-Hu

🎯
Focusing
Data Engineer | MSc in Big Data Analytics @ Trent University, CA

Block or report Smars-Bin-Hu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Smars-Bin-Hu /README.md

👋 Hi there

I’m Smars Hu. A data engineer & minimalist based in Toronto 🇨🇦

GIF

God help those who help themselves

自救者,人恒救之

🔨 3 Years of Experience in Data Engineering

💼 Data Engineer @ ScotiaBank, Toronto, Canada🇨🇦

🎓 Master of Science in Big Data Analytics @ Trent University, Ontario, Canada 🇨🇦

📝 My website

📞 Book a meeting with me

🚀 Certificates

associate-badge-de 0_atCqIOGA5HoUwVl0

🔨 Projects

Simulated an enterprise-level on-premise self-managed big data distributed cluster using Docker containers. Integrated components include Hadoop, Zookeeper, Spark, Hive, MySQL, Airflow, Prometheus, ClickHouse, and Power BI. Developed a data warehouse for an e-commerce backend based on dimensional modeling theory and built a BI analytics system for reporting and data analysis.

HTML tutorial

Reproduced a modern enterprise-grade Azure cloud data engineering architecture widely adopted in North America. Leveraged technologies such as Databricks, PySpark, ADLS Gen2, Unity Catalog, Delta Lake, Power BI, and Azure Data Factory (ADF) to develop cloud-native data pipelines on Azure and perform exploratory data analysis (EDA).

HTML tutorial

💻 Tech Stack

☘️ Languages

Python SQL Java Scala R

☘️ Distributed Computation & Data Warehouse

Apache Hadoop Apache Spark Apache Hive Delta Lake Apache ZooKeeper

☘️ Streaming & Lakehouse Architecture

Apache Flink Apache Kafka Strcutured Streaming

☘️ Data Engineering Practices

☘️ Databases: OLAP, OLTP & NoSQL

MySQL Oracle (OLTP)

ClickHouse Static Badge (OLAP)

Redis Elastic Search Kibana (NoSQL/Search)

☘️ Cloud-Native Data Engineering, Containerization & Platform Tools

Azure Databricks (Synapse, ADLS Gen2, Databricks, Data Factory)

AWS (S3, Lambda)

☘️ DevOps & Monitoring:

Docker Kubernetes Prometheus Grafana

GitLab CI GitLab CI Bitbucket

☘️ Basic Tools

CHAT GPT Claude 3.7 DeepSeek Cursor (AI)

Linux Git Apache Maven Anaconda (OS, Version Control, API, Dev environment)

Jira Markdown (Project Management, Doc)

🌐 Social

https://www.linkedin.com/in/smars-hu/ https://www.youtube.com/@smars_hu https://www.instagram.com/smars.hu/

Pinned Loading

  1. EComDWH-BatchDataProcessingPlatform EComDWH-BatchDataProcessingPlatform Public

    This project aims to build an enterprise-grade offline data warehouse solution based on e-commerce platform order data.

    Python 173 23

  2. azure-cloud-datapipeline-EDA azure-cloud-datapipeline-EDA Public

    A cloud-native data pipeline and visualization project analyzing Formula 1 racing data using Azure, Databricks, Delta Lake, Tableau, and Python for insightful EDA and interactive dashboards.

    Jupyter Notebook 98 10

AltStyle によって変換されたページ (->オリジナル) /