Bachelor's degree in Computer Science, Information Systems, Data Engineering, Statistics, or a related field; a Master's degree is an advantage.
Professional certifications in Cloud Data Engineering, Big Data Engineering, or related technologies are preferred.
Minimum 1 year of hands-on experience in Data Engineering, Data Warehousing, or a related role.
Proven experience designing, developing, and maintaining production-grade data pipelines in cloud or hybrid environments.
Strong proficiency in SQL and data modeling techniques, including dimensional, relational, Data Vault, or equivalent methodologies.
Experience designing and implementing ETL/ELT pipelines for batch and streaming data processing.
Hands-on experience with cloud-based data warehouses, data lakes, or lakehouse platforms on AWS, Microsoft Azure, or Google Cloud Platform (GCP).
Familiarity with modern data storage technologies, including Parquet, ORC, Iceberg, Delta Lake, Hudi, and data partitioning strategies.
Experience using data processing frameworks such as Apache Spark, distributed SQL engines, or similar technologies.
Hands-on experience with Apache Airflow or similar workflow orchestration tools.
Good understanding of CI/CD practices for data engineering, including automated testing, deployment, and environment promotion.
Strong knowledge of data quality management, data validation, monitoring, and data reliability best practices.
Understanding of data governance principles, including metadata management, data lineage, data ownership, data contracts, and data security.
Familiarity with cloud infrastructure monitoring, performance optimization, and cost management for data workloads.
Strong analytical, troubleshooting, debugging, and problem-solving skills, particularly in resolving production issues.
Excellent communication skills with the ability to explain technical concepts to both technical and non-technical stakeholders.
Strong collaboration skills and the ability to work effectively with cross-functional teams, including IT, Business, and external partners.
Demonstrated ownership, accountability, and commitment to delivering reliable, scalable, and high-quality data solutions.
Experience in the telecommunications, media, or subscription-based industry is an advantage.
Experience working with customer, billing, subscription, or network operations data is preferred.
Job Responsibilities:
Design, develop, implement, and maintain scalable data pipelines to support enterprise data integration, transformation, and analytics.
Build, optimize, and manage ETL/ELT processes for both batch and real-time data workloads.
Design and maintain data models, database schemas, and storage structures that support business intelligence and analytical requirements.
Develop and optimize SQL queries to ensure efficient data processing and high-performance data retrieval.
Build and maintain cloud-based data warehouse, data lake, or lakehouse solutions using industry best practices.
Develop and manage data orchestration workflows using Apache Airflow or equivalent scheduling tools.
Implement data validation, monitoring, and quality assurance processes to ensure data accuracy, consistency, and reliability.
Establish and maintain data governance standards, including metadata management, data lineage, ownership, and data security.
Monitor, troubleshoot, and resolve production issues affecting data pipelines, data platforms, and processing workflows.
Optimize data processing performance, cloud resource utilization, and operational costs for data engineering workloads.
Collaborate with Data Analysts, Data Scientists, Software Engineers, Business Teams, and other stakeholders to understand data requirements and deliver scalable data solutions.
Support CI/CD implementation for data engineering projects to enable reliable deployment and continuous delivery.
Develop and maintain technical documentation for data architecture, pipelines, workflows, and operational procedures.
Ensure compliance with data governance, privacy, security, and organizational standards throughout the data lifecycle.
Continuously identify opportunities to improve data platform performance, scalability, automation, and operational efficiency.