Company Profile:
We’re Hiring at CGI for our GCC - Right Here in Hyderabad!
Join us at the intersection of technology, finance, and innovation. You will be working to support one of the top-tier financial institutions in the U.S. You’ll help shape digital solutions for a global enterprise—from the ground up.
This is more than a job. It’s your opportunity to:
- Work on cutting-edge technologies
- Collaborate with global teams
- Build a career with purpose and impact
Ready to build the future of banking? Let’s talk.
Position Description:
Job Title: Senior Software Engineer – Data Expert
Position: Data Expert
Experience: 6-8 years
Category: Software Development/ Engineering
Shift: General
Main location: India, Telangana, Hyderabad
Position ID: J0726-0241
Employment Type: Full Time
Job Overview:
CGI is looking for an experienced Data Engineer with strong hands-on expertise in Python, PySpark, and Data Engineering practices, along with the ability to work effectively between business stakeholders and development teams.
The ideal candidate will be responsible for understanding business requirements, translating them into scalable data solutions, and collaborating with developers and other technical teams to ensure successful delivery. The role requires a combination of strong technical knowledge, analytical thinking, communication skills, and an understanding of modern data engineering practices.
Key Responsibilities
- Design, develop, test, and maintain scalable data pipelines and data processing solutions using Python and PySpark.
- Develop efficient data transformation, cleansing, aggregation, and processing workflows for large and complex datasets.
- Apply industry-standard Data Engineering best practices across data ingestion, transformation, validation, quality, and optimization.
- Analyze business requirements and translate them into clear technical and data requirements.
- Act as a bridge between business stakeholders, data teams, and software development teams to ensure requirements are clearly understood and delivered.
- Collaborate with business analysts, product owners, developers, architects, QA teams, and other stakeholders throughout the project lifecycle.
- Understand existing business processes and identify opportunities to improve them through data-driven solutions.
- Design and implement reliable and scalable ETL/ELT pipelines.
- Work with structured and unstructured data from multiple sources and develop appropriate data processing solutions.
- Optimize PySpark jobs and data pipelines for performance, scalability, and reliability.
- Perform data validation, quality checks, reconciliation, and troubleshooting to ensure accuracy and consistency.
- Investigate data issues, identify root causes, and implement appropriate technical solutions.
- Participate in technical discussions, solution design, code reviews, and architecture discussions.
- Ensure data solutions are developed in accordance with organizational standards, security requirements, and data governance practices.
- Support production data pipelines and resolve technical issues within agreed SLAs.
- Create and maintain technical documentation, data flow diagrams, mapping documents, and process documentation.
- Participate in Agile/Scrum ceremonies, including sprint planning, daily stand-ups, backlog refinement, sprint reviews, and retrospectives.
- Provide technical guidance and collaborate with developers to ensure successful implementation of data solutions.
Required Technical Skills
Python
- Strong hands-on experience with Python for data engineering and data processing.
- Good understanding of Python programming concepts, functions, modules, exception handling, and object-oriented programming.
- Experience developing reusable and maintainable Python-based data processing components.
- Ability to troubleshoot and optimize Python applications and data processing scripts.
PySpark / Apache Spark
- Strong hands-on experience with PySpark.
- Good understanding of Spark architecture and distributed data processing.
- Experience developing PySpark jobs for large-scale data transformation and processing.
- Knowledge of Spark DataFrames, SQL, transformations, actions, joins, aggregations, and partitioning.
- Experience troubleshooting and optimizing Spark jobs for performance and scalability.
Data Engineering
- Strong understanding of data engineering concepts and best practices.
- Experience developing ETL/ELT pipelines.
- Understanding of data ingestion, transformation, validation, reconciliation, and data quality processes.
- Good knowledge of data modeling and database concepts.
- Experience working with large datasets and complex data processing requirements.
- Understanding of batch and, preferably, real-time data processing concepts.
- Strong SQL skills and experience working with relational databases.
- Understanding of data warehouse, data lake, and modern data platform concepts.
Business & Technical Collaboration
- Ability to understand business requirements and convert them into technical specifications and data solutions.
- Work closely with business stakeholders to clarify requirements, priorities, dependencies, and expected outcomes.
- Collaborate with developers to explain business requirements and ensure technical implementation aligns with business objectives.
- Identify gaps or ambiguities in requirements and proactively work with stakeholders to resolve them.
- Communicate technical concepts clearly to both technical and non-technical audiences.
- Coordinate across multiple teams to track requirements, dependencies, risks, and delivery milestones.
- Support user acceptance testing and business validation of data solutions.
- Demonstrate strong stakeholder management and interpersonal skills.
Key Responsibilities Across the Delivery Lifecycle
- Requirement gathering and analysis.
- Data profiling and source-system analysis.
- Technical solution design.
- Data pipeline development.
- Data transformation and processing using Python/PySpark.
- Data quality and validation.
- Performance optimization.
- Production support and troubleshooting.
- Business stakeholder communication.
- Technical documentation.
- Code reviews and development best practices.
- Continuous improvement of data engineering processes.
Good to Have
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Knowledge of Kafka or other messaging/event-streaming technologies.
- Experience with data lake technologies such as Delta Lake, Apache Iceberg, or Hudi.
- Knowledge of Hadoop ecosystem technologies such as HDFS, Hive, or Impala.
- Experience with workflow orchestration tools such as Apache Airflow.
- Experience with Docker, Kubernetes, or containerized data applications.
- Knowledge of CI/CD practices and tools.
- Exposure to data governance, metadata management, and data security.
- Experience working with Agile/Scrum methodologies.
Soft Skills
- Strong analytical and problem-solving abilities.
- Excellent communication and stakeholder management skills.
- Ability to work effectively between business and technical/development teams.
- Strong documentation and presentation skills.
- Ability to work independently as well as collaboratively in a team environment.
- Proactive approach to identifying and resolving technical and business challenges.
- Strong ownership and accountability for deliverables.
- Ability to manage multiple priorities in a fast-paced environment.
About CGI :
CGI is an equal opportunity employer. In addition, CGI is committed to providing accommodation for people with disabilities in accordance with provincial legislation. Please let us know if you require reasonable accommodation due to a disability during any aspect of the recruitment process and we will work with you to address your needs.