Introduction
When a data platform grows from a few gigabytes to terabytes or petabytes, storing the data is only one part of the problem.
A typical data lake can store huge amounts of data cheaply in systems such as Amazon S3 or HDFS. The real challenge is making that collection of files behave like a reliable analytical table.
Consider an e-commerce company receiving millions of orders every day. Its analytics team may want to answer questions such as:
How much revenue was generated last month?






