Company Profile:
Founded in 1976, CGI is among the largest independent IT and business consulting services firms in the world. With 94,000 consultants and professionals across the globe, CGI delivers an end-to-end portfolio of capabilities, from strategic IT and business consulting to systems integration, managed IT and business process services and intellectual property solutions. CGI works with clients through a local relationship model complemented by a global delivery network that helps clients digitally transform their organizations and accelerate results. CGI Fiscal 2024 reported revenue is CA$14.68 billion and CGI shares are listed on the TSX (GIB.A) and the NYSE (GIB). Learn more at cgi.com.
Your future duties and responsibilities
Job Title: Data Engineer (Azure & Databricks)
Position: Lead Analyst
Experience: 10–12 Years
Category: Software Development / Engineering
Shift: General Shift
Work Mode: Work from Office / Hybrid (as per business requirement)
Primary Location: India, Karnataka, Bangalore
CGI Office: Bangalore
Employment Type: Full Time
Required qualifications to be successful in this role
About the Role :
We are seeking a highly experienced and technically hands on Lead Data Engineer to design, build, and optimize our next generation Enterprise Data Warehouse and data Lakehouse solutions. You will be responsible for leading complex data integration initiatives, migrating legacy workloads, and building highly scalable ETL/ELT pipelines using Python, Databricks, and the Azure Cloud ecosystem.
As a technical leader, you will collaborate closely with enterprise architects, business stakeholders, and data scientists to ensure our data infrastructure is robust, secure, and aligned with business objectives.
Mandatory Key Responsibilities :
. Architecture & Design: Design and implement modern, scalable Data Warehousing and Data Lakehouse architectures on Microsoft Azure using Databricks and Delta Lake.
. ETL/ELT Development: Architect, build, and maintain highly efficient data pipelines to ingest, clean, transform, and load massive volumes of data from various sources (relational, non relational, APIs, flat files).
. Databricks & PySpark Mastery: Write optimized, production grade PySpark and Python code in Azure Databricks for complex data transformations and advanced analytics.
. Legacy Modernization: Lead the migration of legacy Enterprise Data Warehouse (EDW) workloads (e.g., traditional SQL stored procedures, Unix scripts, legacy ETL tools) to a modern Azure/Databricks cloud stack.
. Azure Infrastructure: Utilize and orchestrate Azure native services including Azure Data Factory (ADF), Azure Data Lake Storage (ADLS Gen2), Azure Key Vault, and Azure SQL/Synapse.
. Performance Tuning: Monitor, troubleshoot, and optimize Spark clusters, SQL queries, and ETL jobs for maximum speed, efficiency, and cost effectiveness.
. Leadership & Mentorship: Lead a team of data engineers, conduct code reviews, enforce engineering best practices (CI/CD, automated testing), and guide the team on technical roadblocks.
Required Skills & Qualifications
. Experience: 10 to 12 years of core experience in Data Engineering, Business Intelligence, or Enterprise Data Warehousing.
. Data Warehousing: Deep understanding of data modeling concepts (Kimball, Inmon, Data Vault), schema design (Star/Snowflake), and historical data tracking (SCD Types 1, 2, 3).
. ETL/ELT: Extensive background in designing enterprise grade data integration frameworks and orchestrating complex workflows.
. Python: Strong programming proficiency in Python (Pandas, core Python, OOP) for data manipulation, automation, and API integration.
. Databricks & Spark: 4+ years of hands on experience with Databricks, Apache Spark (PySpark & Spark SQL), and Delta Lake architecture.
. Azure Cloud: Proven expertise in the Azure Data ecosystem, specifically Azure Data Factory (ADF) and ADLS Gen2.
. Advanced SQL: Expert level SQL skills, capable of writing and tuning highly complex queries, window functions, and migrating legacy stored procedures.
Preferred Skills
. Experience migrating legacy on premise EDW environments (SQL Server, Teradata, Oracle, Unix Shell scripting) to the cloud.
. Experience with CI/CD pipelines and DevOps practices using Azure DevOps, Git, or GitHub Actions.
. Understanding of Data Governance, Data Quality frameworks, and compliance standards (e.g., RBAC, Unity Catalog in Databricks).
Certifications: Azure Data Engineer Associate (DP 203) or Databricks Certified Data Engineer Professional
#LI-SS65
Together, as owners, let’s turn meaningful insights into action.
Life at CGI is rooted in ownership, teamwork, respect and belonging. Here, you’ll reach your full potential because…
You are invited to be an owner from day 1 as we work together to bring our Dream to life. That’s why we call ourselves CGI Partners rather than employees. We benefit from our collective success and actively shape our company’s strategy and direction.
Your work creates value. You’ll develop innovative solutions and build relationships with teammates and clients while accessing global capabilities to scale your ideas, embrace new opportunities, and benefit from expansive industry and technology expertise.
You’ll shape your career by joining a company built to grow and last. You’ll be supported by leaders who care about your health and well-being and provide you with opportunities to deepen your skills and broaden your horizons.
That same commitment to fairness extends to how we use technology. To support our recruitment team, AI tools may be used to help assess applications though they never replace human judgement. All hiring decisions remain entirely in the hands of our recruitment professionals.
Come join our team—one of the largest IT and business consulting services firms in the world.