T&T - Engineering - Consultant - AWS SAS Data engineer - Multiple Location
- Job requisition ID : 111928
- Location: Ahmedabad
- Entity: Deloitte Touche Tohmatsu India LLP
The Team
Deloitte's Technology & Transformation practice helps organizations uncover and unlock the value buried within vast amounts of data. Our global network provides strategic guidance and implementation services to help companies manage data from disparate sources and convert it into accurate, actionable information. Our practice addresses the continuum of opportunities in cloud, next-generation analytics, and advanced integration technologies to drive insight-driven advantages.
Your work profile
The Data Engineer is responsible for designing, developing, and maintaining scalable data pipelines and robust data architecture within the AWS ecosystem. This role focuses on building high-performance data processing solutions, executing complex data migrations, and ensuring data integrity across diverse cloud-based environments. The engineer will bridge the gap between raw data and actionable insights through advanced processing and architectural design.
Key Skills & Competencies
- Design, develop, and maintain scalable ETL/ELT pipelines using PySpark for large-scale data processing.
- Build and optimize complex data workflows and transformation logic to handle diverse datasets.
- Execute end-to-end data migration strategies, moving data from legacy systems (such as SAS environments) to modern cloud architectures.
- Design and implement robust data models and schemas to support advanced analytics and business intelligence.
- Manage data ingestion, cleansing, validation, and orchestration processes across various formats (Parquet, Avro, JSON, CSV, etc.).
- Leverage AI-assisted coding tools (such as GitHub Copilot) to enhance productivity and ensure high code quality.
- Architect and implement data solutions utilizing AWS components (e.g., AWS Glue, EMR, S3, Redshift, Lambda, Kinesis).
- Ensure strict adherence to cloud security best practices, including identity management, encryption, and data protection.
- Optimize the performance, reliability, and cost-efficiency of cloud data workloads.
- Troubleshoot and proactively resolve data pipeline failures and production issues to ensure system stability.
- Collaborate with cross-functional teams to design and deploy scalable, cloud-native data architectures.
Success Metrics
- Core Data Processing: Strong hands-on experience with PySpark and distributed computing frameworks.
- Legacy Integration: Proven experience working with SAS for data manipulation and statistical analysis.
- Cloud Ecosystem: Deep expertise in AWS services (S3, Glue, EMR, Redshift, Athena, etc.).
- Data Migration: Expert knowledge of data migration patterns, schema mapping, and validation techniques.
- Design Patterns: Strong understanding of Data Modeling, Star/Snowflake schemas, and Data Lakehouse architecture.
- Technical Fundamentals: Proficiency in SQL, Python, and advanced data structures.
- DevOps & Automation: Proficiency in DevOps practices and CI/CD pipelines using AWS Code Pipeline or GitHub Actions.
- Orchestration: Experience with workflow orchestration tools like Apache Airflow or AWS Step Functions.
- Streaming: Knowledge of real-time data processing using AWS Kinesis or Kafka.
- Compliance: Experience working within regulated environments with stringent data privacy and security standards.
- Strong communication and collaboration skills for working within cross-functional, agile teams.
- Ability to thrive in a fast-paced environment with evolving technological requirements.
- Proactive mindset with meticulous attention to detail regarding data quality and integrity.
- A continuous learning attitude, particularly regarding emerging AI-driven data technologies.
Preferred Qualifications
Education - B. tech / BE in Computer Science or Information Technology
Location and way of working
- Location: Any location in India (Preferably Pune)
- Work Model: Hybrid – Mandatory 3 days