# Data Lake Platform Boilerplate

## Overview
Big data storage, data ingestion, ETL processing, analytics and data science workflows

## Tech Stack
- **Storage**: Apache Hadoop, AWS S3, Azure Data Lake
- **Processing**: Apache Spark, Apache Flink
- **Database**: Apache Hive, Delta Lake
- **Workflow**: Apache Airflow, Prefect
- **Analytics**: Jupyter, Apache Zeppelin
- **Streaming**: Apache Kafka, Kinesis

## Specialized Agents
- `data-lake-architect` - Data lake design and governance
- `etl-pipeline-engineer` - Data transformation workflows
- `big-data-processing-specialist` - Distributed data processing
- `data-catalog-manager` - Metadata management and discovery

## Key Features
- Scalable data ingestion
- Multi-format data storage
- ETL/ELT processing pipelines
- Data cataloging and governance
- Analytics and visualization
- Real-time streaming processing

## Data Sources
- Structured databases (SQL)
- Semi-structured data (JSON, XML)
- Unstructured data (logs, documents)
- Streaming data sources
- IoT sensor data
- Social media feeds

## Architecture Patterns
- Lambda architecture (batch + streaming)
- Data mesh architecture
- Lake house architecture
- Event-driven data processing

## Implementation Requirements
- Petabyte-scale data storage
- Distributed processing capabilities
- Data governance and lineage
- Security and access controls
- Cost optimization strategies
- Self-service analytics tools