Design, build, and operate structured, semi-structured, and unstructured data stores supporting real-time and batch analytics workloads within your assigned scope.
Recommend and apply appropriate distributed data and computation technologies for assigned use cases, balancing performance, scalability, and maintainability.
Implement secure, reliable access patterns for connecting to multiple internal and external data sources.
Apply encryption and other controls to protect data at rest in line with security standards.
Build and maintain real-time analytics pipelines using distributed processing frameworks such as Spark and Presto/Trino.
Containerize and deploy data services using Kubernetes and related orchestration tooling.
Design and implement data replication strategies to support active/active data center resilience.
Design, implement, and test backup and recovery procedures for owned data platforms.
Own code quality, testing, and documentation for data engineering components within your scope.
Collaborate with engineering, security, and platform teams to ensure data solutions meet reliability and operational standards.