Ivan Ong at a Databricks event

Hello there! I’m Ivan.

I’m a data engineer with 7 years of experience across technology and financial services, building reliable data platforms that power analytics, machine learning and AI.

Today at Applied Materials, I build data foundations for MLOps and AI in a complex global supply chain environment.

Previously, I worked across data engineering, analytics and technical advisory, with a focus on turning data into measurable business outcomes:

  • Databricks — Helped drive platform adoption and consumption through technical advisory, architecture validation and hands-on AI workshops.
  • Google · via JLL — Improved data latency and reduced infrastructure costs across pipelines and data platforms supporting global real estate analytics.
  • HSBC — Built data pipelines, dashboards and predictive analytics that supported revenue growth, operational efficiency and cost reduction.

I enjoy working at the intersection of data engineering, AI and business problems, especially where better data can make complex systems easier to understand and operate.

THE JOURNEY

Where I’ve worked

Technical advisory, data platforms and applied analytics.

Jun 2026 — Present

Data Engineer (Data and AI Platforms)

Applied Materials

Building data foundations for MLOps and AI, connecting reliable pipelines with platforms that support machine learning and AI applications.

Jul 2025 — Jan 2026

Senior Solutions Engineer

Databricks

Technical advisor in pre-sales, helping customers design and adopt data and AI solutions on Databricks.

  • Validated customer system designs and advised on technical architecture.
  • Guided customers in optimizing Databricks workloads and platform usage.
  • Led workshops teaching customers to build AI use cases with Databricks.

Aug 2022 — Jun 2025

Data Engineer

JLL· Supporting Google REWS

Vendor Data Engineer through JLL, building data foundations for Google Real Estate & Workplace Services (REWS).

  • Built and maintained the warehouse and data models powering global real estate portfolio analytics.
  • Developed ingestion and transformation pipelines with validation and data quality checks for reliable reporting.
  • Partnered with application, analytics and business teams to model assets, space utilization and lease metrics across regions.

Oct 2019 — Aug 2022

Data & Analytics

HSBC

Assistant Vice President, Data & Analytics
Apr 2021 — Aug 2022
Data Analyst
Oct 2019 — Apr 2021

Combined data engineering, analytics and predictive modeling across two roles.

  • Built data pipelines and transformations to support analysis and reporting.
  • Designed dashboards and analyzed data to inform business decisions.
  • Developed predictive models using Python.

THE TOOLBOX

Skills & capabilities

Data engineering depth. Practical AI. Business context.

Programming Languages & Libraries

The tools I use to query, transform and work with data.

  • SQL
  • Python
  • PySpark
  • R
  • pandas

Data Platforms & Cloud

Cloud warehouses and distributed data platforms.

  • Databricks
  • Apache Spark
  • Google Cloud
  • AWS
  • Big Query

Data Engineering

Reliable pipelines, well-defined models and efficient workloads.

  • ETL / ELT pipelines
  • Data Orchestration
  • Data Modeling
  • System Design
  • Data Quality & Validation

Machine Learning & AI

From predictive models to practical AI applications.

  • Predictive models
  • MLOps
  • LLM integration
  • Structured extraction
  • Prompt iteration

Engineering & Delivery

Tested code, integration and repeatable delivery.

  • GIT
  • REST APIs
  • Automated testing
  • GitHub Actions

How I work

Business problem framing
Translate business requirements into practical data models, pipelines and AI use cases.
Technical advisory
Validate system designs, explain trade-offs and guide platform optimization.
Customer enablement
Make complex ideas usable through workshops and collaboration with technical and business teams.

CONTINUOUS LEARNING

Certifications

Credentials in cloud architecture, data engineering and machine learning.

SELECTED WORK

Projects

Applied AI and reliable data foundations.

APPLIED AI01

Email Intelligence Pipeline

AI pipeline prototype

Turning unstructured emails into validated data, ready to query.

  • Python
  • Gemini
  • BigQuery
  • FastAPI
  • Pydantic
  • pytest
Read case study

The problem

Emails contain useful business context, but free-form text is difficult to classify, search and use in downstream workflows.

Architecture

  1. Email files
  2. Gemini extraction
  3. Pydantic validation
  4. BigQuery
  5. FastAPI

Engineering choices

  • Extracts intent, entities, sender details and summaries into a validated schema.
  • Checks content-based identifiers before processing and uses parameterized API queries.
  • Includes mocked database and API tests, CI checks and a tool for comparing prompt revisions.

Outcome

Connects LLM extraction to warehouse storage and an API with intent filtering, making email content available as structured records.

Scope & trade-offs

Prompt comparison is qualitative; a scored evaluation dataset remains future work.

DATA ENGINEERING02

HDB Resale Data Pipeline

Open-data engineering exercise

From historical housing transactions to traceable, quality-checked datasets.

  • Python
  • pandas
  • Jupyter
  • pytest
  • GitHub Actions
Read case study

The problem

Historical resale files vary in schema and require explicit rules for missing fields, duplicates and unusual prices before they can support consistent analysis.

Architecture

  1. Public CSVs
  2. Schema normalization
  3. Profiling & validation
  4. Cleaning
  5. Datasets & reports

Engineering choices

  • Preserves source-file lineage while reconciling historical schemas.
  • Separates rejected records from price anomalies retained for review.
  • Uses deterministic transformations, edge-case tests and notebook execution in CI.

Outcome

Produces cleaned, transformed and SHA-256-hashed outputs for 2012–2016 transactions, with separate rejection and anomaly reports.

Scope & trade-offs

Compact identifiers can collapse distinct transactions. Collision reports make that trade-off visible; the cleaned dataset remains available before identifier-level deduplication.