Weeks 1-2 (9/24, 9/29) [Slides]
Course Overview
Introduction to big-data management: frameworks, data management, analytics, machine learning, etc. The classes focus on frameworks and data management. What is big data. Frameworks and cloud computing. OLTP vs OLAP vs BigData. Big data frameworks. Data cleaning.
Weeks 2-3 (10/1, 10/6, 10/8) [Slides]
Relational Databases: SQL refresher, relational model, xml, json, semi-structured data, RDBMS, AWS RDS.
Database System Concepts by Avi Silberschatz, Henry F. Korth, and S. Sudarshan
Database Management Systems by Johannes Gehrke and Raghu Ramakrishnan
Database Systems: The Complete Book by Héctor García-Molina, Jeffrey Ullman, and Jennifer Widom
Week 4 (10/13, 10/15) [Slides]
Limitations of RDBMSs and motivation for NoSQL; Intro to Cassandra and MongoDB [less cassandra, more sync policies]
Non-RDBMS: Key-value stores, distributed storage, NoSQL storage: Column Stores (C-store, HBase, Cassandra)
Good summary of NoSQL material: http://www.christof-strauch.de/nosqldbs.pdf
http://cassandra.apache.org/
Stonebraker, Mike, Daniel J. Abadi, Adam Batkin, Xuedong Chen, Mitch Cherniack, Miguel Ferreira, Edmond Lau et al. "C-store: a column-oriented DBMS." In Proceedings of the 31st international conference on Very large data bases, pp. 553-564. VLDB Endowment, 2005.
DeCandia, Giuseppe, Deniz Hastorun, Madan Jampani, Gunavardhan Kakulapati, Avinash Lakshman, Alex Pilchin, Swaminathan Sivasubramanian, Peter Vosshall, and Werner Vogels. "Dynamo: amazon's highly available key-value store." In ACM SIGOPS operating systems review, vol. 41, no. 6, pp. 205-220. ACM, 2007.
George, Lars. HBase: the definitive guide: random access to your planet-size data. " O'Reilly Media, Inc.", 2011.
Week 5
MongoDB (cont'd) (10/20)
Review session (10/22) [Review Slides]
Week 6
Midterm (10/27)
Data Management for AI and Vector Stores (10/29) [Slides]
Survey of vector database management systems. JJ Pan, J Wang, G Li - arXiv preprint arXiv:2310.14021, 2023 - arxiv.org
https://learn.microsoft.com/en-us/data-engineering/playbook/solutions/vector-database/
https://www.mongodb.com/resources/basics/databases/vector-databases
Week 7
Data Management for AI and Vector Stores (cont'd) (11/03)
Graph databases, Neo4j (11/5) [Slides]
Week 8 (11/10,11/12) [Slides]
Intro to Big data frameworks : Distributed file systems with focus on HDFS, MapReduce, Hadoop, and Spark.
Chapter 2 of Mining of Massive Datasets by Jure Leskovec, Anand Rajaraman, and Jeff Ullman.
Dean, Jeffrey, and Sanjay Ghemawat. "MapReduce: simplified data processing on large clusters." Communications of the ACM 51, no. 1 (2008): 107-113.
Shvachko, Konstantin, Hairong Kuang, Sanjay Radia, and Robert Chansler. "The hadoop distributed file system." In Mass storage systems and technologies (MSST), 2010 IEEE 26th symposium on, pp. 1-10.
Learning Spark: Lightning-Fast Big Data Analysis by Andy Konwinski, Holden Karau, Matei Zaharia, and Patrick Wendell
Zaharia, Matei et al. "Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing." In Proceedings of the 9th USENIX conference on Networked Systems Design and Implementation, 2012.
Texera and workflow systems, add slides about it
Week 9
No class, instructor at conference (11/17)
Spark practice (11/19) [MapReduce Slides, Spark Slides]
Week 10 (11/24)
Final review [Slides] [slides][notebook]
Week 11 (12/01, 12/03)
Spark deployment (AWS vs. Databricks) and applications
[Lakehouse and Column Stores] [Slides]
Streaming data querying and Flink [Slides]
What Is a Lakehouse: https://www.databricks.com/blog/what-is-data-lakehouse
What Is a Lakehouse: https://aws.amazon.com/what-is/data-lakehouse/
Learn Flink: https://nightlies.apache.org/flink/flink-docs-stable/docs/learn-flink/overview/
FINAL (TBD)