
Harvard University
Fundamentals of TinyML
4.567 ratingsFocusing on the basics of machine learning and embedded systems, such as smartphones, this course will introduce you to the “language” of TinyML.
This course is designed for data engineers, analytics engineers, data platform engineers, and data architects who work with data lakes and want to modernize their data infrastructure. It's also valuable for software engineers transitioning into data roles and technical leads evaluating Apache Iceberg for their data.
By the end of this course, you will be able to:
- Build and configure an Apache Iceberg lakehouse using catalogs, object storage, and query engines like Spark and Trino
- Design optimal table structures using hidden partitioning, sort orders, and column metrics to maximize query performance
- Migrate existing data from Hive tables, Parquet files, CSV, and databases into Iceberg using snapshot, migrate, and reserialization approaches
- Implement production workflows using Write-Audit-Publish for validation, branching for testing, and rollback for recovery
- Evolve table schemas and partition specifications without downtime or rewriting data
- Execute maintenance operations including data file compaction, metadata compaction, and snapshot expiration
- Configure write strategies (merge-on-read vs copy-on-write) and distribution modes for different workload requirements
- Manage concurrent operations and avoid conflicts in multi-writer scenarios
To be successful in this course, you should have:
- Working knowledge of SQL and relational database concepts (tables, schemas, queries)
- Basic understanding of data engineering concepts including ETL/ELT, data warehouses, and data lakes
- Familiarity with command-line interfaces and Docker for running the course environment
- Comfort reading and understanding code examples in Python/PySpark (code is provided; you don't need to write from scratch)
- Experience with Apache Spark or distributed computing is helpful but not required—core concepts are explained throughout the course
Apache Iceberg, Iceberg, Apache, and the Apache feather logo are either registered trademarks or trademarks of The Apache Software Foundation. No endorsement by The Apache Software Foundation is implied by the use of these marks.
Some points apply to every course of this kind; see how we rank.
Snowflake is in Tier 3: good universities, respected companies, nonprofits and well-known teachers of our institution ranking (74/100).

Harvard University
Focusing on the basics of machine learning and embedded systems, such as smartphones, this course will introduce you to the “language” of TinyML.

Harvard University
Learn the concepts and techniques that make up the foundation of data science and machine learning.

Harvard University
An introduction to the intellectual enterprises of computer science and the art of programming.

Harvard University
Build a movie recommendation system and learn the science behind one of the most popular and successful data science techniques.

Massachusetts Institute of Technology
Instructors: Esther Duflo and Sara Ellison View the complete course: https://ocw.mit.edu/courses/14-310x-data-analysis-for-social-scientists-spring-2023 This course introduces methods for harnessing data to answer questions of cultural, social, economic, and policy interest. We will start with esse…

Harvard University
An introduction to programming using Python, a popular language for general-purpose programming, data science, web programming, and more.