

Search by job, company or skills

We are looking for a Python Developer with strong PySpark experience to design, build, and maintain scalable data pipelines and analytics solutions. You will work closely with data engineers, analysts, and business stakeholders to transform large datasets into reliable, high-performance data products using Python, PySpark, and modern big-data platforms.
Key Responsibilities
· Design, develop, and maintain ETL/ELT pipelines using Python and PySpark on platforms such as Databricks, Spark on Kubernetes, or cloud-native Spark services.
· Write clean, efficient, and reusable PySpark code for data ingestion, transformation, aggregation, and quality checks at scale.
· Optimize Spark jobs for performance (partitioning, caching, broadcast joins, skew handling) and cost in cloud environments (AWS/Azure/GCP).
· Develop and maintain SQL queries, views, and data models to support analytics, reporting, and machine learning use cases.
· Implement data quality checks, monitoring, and alerting for pipelines (e.g., using Great Expectations, custom checks, or platform-native tools).
· Collaborate with cross-functional teams to understand business requirements and translate them into technical designs and data solutions.
· Participate in code reviews, enforce coding standards, and contribute to shared libraries and frameworks for data engineering.
· Troubleshoot and resolve production issues, perform root-cause analysis, and implement long-term fixes.
· Document data pipelines, data models, and operational procedures clearly and concisely.
Required Skills & Qualifications
· 8+ years of professional software/data engineering experience, with at least 2+ years focused on Python + PySpark.
· Strong proficiency in Python (data structures, OOP, packaging, testing, virtual environments).
· Hands-on experience with PySpark and Apache Spark concepts (RDDs, DataFrames, Spark SQL, Catalyst optimizer).
· Solid SQL skills, including complex queries, window functions, and performance tuning.
· Experience building and maintaining ETL/ELT pipelines on big-data platforms (e.g., Databricks, EMR, Dataproc, Synapse).
· Familiarity with at least one major cloud platform: AWS, Azure, or GCP (storage, compute, IAM, networking basics).
· Experience with version control (Git), CI/CD pipelines, and basic DevOps practices for data projects.
· Good understanding of data modeling, data warehousing concepts, and dimensional modeling.
· Strong problem-solving, debugging, and analytical skills.
· Bachelor's degree in Computer Science, Engineering, Mathematics, or a related field (or equivalent practical experience).
Srihire (Srihire Consulting Services Private Limited) is an India-based staffing and recruitment company that offers IT and non‑IT staffing, RPO, MSP and staff‑augmentation services
Job ID: 150485067