13.4 Big Data Frameworks
Big data describes problems where the data is too large, too fast, or too varied for a single machine to handle. The classic frame is the five Vs. This lesson connects them to architecture choices.
The Five Vs
| V | Meaning | Engineering Lever |
|---|---|---|
| Volume | Terabytes to petabytes | Distributed storage (S3, HDFS) |
| Velocity | Events per second to millions | Streams (Kafka, Kinesis, Pulsar) |
| Variety | Tables, JSON, logs, video | Lakehouse, schema registry |
| Veracity | Quality, trust, provenance | dbt tests, lineage tools |
| Value | Output that moves a metric | Use-case prioritisation |
Reference Architecture
Figure 5.11 - The classic batch-plus-stream Lambda topology.
Toolchain Map
| Layer | Open Source | AWS | GCP | Azure |
|---|---|---|---|---|
| Stream | Kafka, Pulsar, NiFi | Kinesis, MSK | Pub/Sub | Event Hubs |
| Batch ingest | Airbyte, Fivetran OSS | Glue, DMS | Datastream | ADF |
| Storage | HDFS, MinIO | S3 | GCS | ADLS |
| Compute | Spark, Flink, Trino | EMR, Glue, Athena | Dataproc, Dataflow | Synapse, Databricks |
| Serving | ClickHouse, Pinot, Druid | Redshift | BigQuery | Synapse SQL |
Discussion
Loading…