Key Responsibilities
1. Data Pipeline Development & Engineering
- Design, develop, and maintain scalable ETL/ELT pipelines using Python, SQL, and GCP tools
- Build batch and real-time data processing workflows
- Automate data pipelines using Apache Airflow (Cloud Composer) for orchestration
2. Cloud Data Platform Management (GCP)
- Develop and manage data solutions using BigQuery, Cloud Storage, Dataflow, and Pub/Sub
- Implement scalable data warehousing and lakehouse architectures
- Monitor and optimize cloud infrastructure performance, reliability, and cost
3. Data Modeling & Transformation
- Design and implement logical and physical data models
- Perform data cleansing, transformation, and aggregation for analytics readiness
- Ensure high data quality, integrity, and consistency across systems [expertia.ai]
4. Analytics & Visualization Enablement
- Enable business insights through dashboards using Looker Studio / LookML
- Collaborate with analysts and business teams to define reporting requirements
- Translate data into actionable insights for decision-making
5. Collaboration & Stakeholder Engagement
- Work closely with data scientists, analysts, and business stakeholders to gather requirements
- Provide data platform support and resolve data-related issues
- Communicate technical solutions effectively to non-technical stakeholders
6. Data Governance, Security & Compliance
- Implement data governance, access control, and security best practices
- Ensure compliance with data privacy regulations and enterprise policies
7. Continuous Improvement & Optimization
- Identify and implement performance optimizations across pipelines and databases
- Introduce automation and best practices for DataOps and CI/CD
- Stay updated with evolving GCP and data engineering technologies
Key Skills & Competencies
Technical Skills
- Strong programming expertise in Python and SQL
- Hands-on experience with Google Cloud Platform (BigQuery, Dataflow, Pub/Sub, Cloud Storage)
- Workflow orchestration using Apache Airflow / Cloud Composer
- Experience with ETL/ELT pipeline design and data warehousing
- Visualization and reporting using Looker Studio / LookML
Functional & Process Skills
- Strong understanding of data lifecycle (ingestion transformation analytics)
- Expertise in data modeling, data integration, and pipeline optimization
- Familiarity with Agile / DevOps practices in data engineering
Soft Skills
- Strong analytical and problem-solving capabilities
- Excellent communication and stakeholder management skills
- Ability to work in cross-functional, global teams
Qualifications & Experience
- Bachelor’s / Master’s degree in Computer Science, Data Engineering, or related field
- 5+ years of experience in data engineering or similar roles
- Proven experience working with GCP-based data platforms
- Hands-on experience in Python, SQL, Airflow orchestration
Good to Have
- Experience with PySpark / Spark / Kafka (streaming pipelines)
- Knowledge of CI/CD, Docker, Kubernetes
- Exposure to AI/ML pipelines and advanced analytics
- Certifications such as Google Professional Data Engineer
Success Metrics (KPIs)
- Pipeline performance and reliability (uptime, latency)
- Data quality and accuracy metrics
- Time-to-deliver data for business insights
- Cost optimization on cloud data infrastructure
- Adoption and usability of dashboards (Looker Studio)