Course Hours: Mon 13.00-15.00, Wed 13.00-15.00
Room: H.208
Weekly Lab: Fri 17.00-19.00, Room H.208
Office Hours: Available upon request
Big Data requires the storage, organization, and processing of data at a scale and efficiency -typically of heterogeneous nature and in streaming flow- that go well beyond the capabilities of conventional information technologies. Such requirements have been first introduced for processing the web, and they are today a common place in many industries. In this respect many traditional assumptions break, new query and programming interfaces are required (Map/Reduce), and new computing models will emerge (Cloud Computing). This course aims to introduce parallel/distributed data processing using the MapReduce (M/R) paradigm and provide insights for developing applications on top of the Hadoop platform.
Big data raises also new challenges in data mining. Given the scale and speed of data that needs to be processed as well the variety of parameters to be taken into account, state of the art machine learning algorithms working offline and expecting homogeneous and clean data are also challenged. There is on ongoing effort to design Big Data Mining algorithms accommodating a parallel/distributed or even a streaming evaluation. Of course such kind of incremental, partial evaluation impacts the quality of obtained statistical models and thus algorithms compromise between quality of the learning and computation time. The course will adopt an algorithmic viewpoint: data mining is about applying algorithms to data, rather than using data to “train” a machine-learning engine of some sort.
The course will consist of lectures based both on textbook material (freely-available for download on the Web) and scientific papers. It will also include programming assignments that will provide students with hands-on experience on building data-intensive applications using existing Big Data tools and platforms. The intended audience of this course is MSc and PhD students but also practitioners who plan to design or develop state-of-the-art algorithms available today for Big Data analysis.
Course material, assignments and announcements are posted on eLearn.
21/09/2026: Course Overview
23/09/2026: Scalable Data Analytics (Spark 4)
28/09/2026: Scalable Data Analytics (Spark 4)
30/09/2026: Scalable Data Analytics (Spark 4) — Assignment 1 Announcement
05/10/2026: Finding Similar Items (LSH → vector search)
07/10/2026: Finding Similar Items (LSH → vector search)
Lab 1 (09/10): From MapReduce to Spark
12/10/2026: Massive Data Processing (Spark SQL, lakehouse)
14/10/2026: Massive Data Processing (Spark SQL, lakehouse)
16/10/2026: Assignment 1 Due
19/10/2026: Extracting Association Rules (+FP-Growth) — Assignment 2 Announcement
21/10/2026: Extracting Association Rules (+FP-Growth)
Lab 2 (23/10): DataFrames & Spark SQL
26/10/2026: Streaming Analytics (sketches + Structured Streaming) — Papers Discussion
28/10/2026: Streaming Analytics (sketches + Structured Streaming) — Αργία (Holiday)
02/11/2026: Schema Discovery — Project Discussion
04/11/2026: Schema Discovery
Lab 3 (06/11): Structured Streaming & Kafka — Assignment 2 Due
09/11/2026: Semantic Summaries
11/11/2026: Semantic Summaries — Αργία (Holiday)
16/11/2026: Property Graphs (GQL, SQL/PGQ, PG-Schema)
18/11/2026: Property Graphs (GQL, SQL/PGQ, PG-Schema)
Lab 4 (20/11): Hands-On Schema Discovery
23/11/2026: Data Ethics (EU AI Act, foundation models)
25/11/2026: Data Ethics (EU AI Act, foundation models)
30/11/2026: Property Graphs Partitioning
02/12/2026: LLMs for Big Data Management
07/12/2026: Quantum Data Management
09/12/2026: Quantum Data Management
14/12/2026: Project / Crash-course Presentations
16/12/2026: Project / Crash-course Presentations
Lab 5 (18/12): Hands-On Quantum