Job purpose
Seeking a hands-on Data Engineer to design, build, and automate data pipelines that bring new datasets onto the platform quickly and reliably as they are identified by the investment team
Key priority is core data engineering: ingesting, transforming, and organizing data so it lands in the right place, in the right shape. Over time, the role will extend into the firm's AI roadmap — preparing large document sets for retrieval-augmented generation (RAG), managing vector databases, and enabling effective querying across integrated datasets to generate insight and reporting.
This is not a fixed-scope project role. The engineer will operate as a responsive, standing capability — picking up new pipeline builds as analysts surface new data requirements, while maintaining and improving what is already in place.
Key responsibilities
- Pipeline development — Design, build, and maintain robust data pipelines to ingest new datasets onto the firm's data platform as demand arises from the investment and analyst teams.
- Automation — Automate ingestion, transformation, and data-quality workflows to reduce manual effort and speed up onboarding of new sources.
- Data platform — Structure and organize datasets so they can be integrated, joined, and queried effectively across the platform.
- AI enablement (roadmap) — Process large volumes of documents into formats suitable for AI consumption; store and index content in vector databases; support retrieval workflows (RAG) that surface the right information reliably.
- Insight & reporting (roadmap) — Enable effective querying across multiple integrated datasets to combine information and support sensible, decision-ready reporting.
- Stakeholder collaboration — Work directly with the hiring manager and analysts to scope new requirements, priorities the pipeline backlog, and respond quickly as new needs emerge.
Key competencies
Required Skills
- 4–5+ years of hands-on data engineering experience, ideally within financial services or data intensive environments.
- Strong experience building and orchestrating data pipelines with Apache Airflow.
- Solid working knowledge of Apache Spark for large-scale data processing.
- Strong Python and SQL skills, with a focus on clean, maintainable, production-quality code.
- Experience designing data models and organizing datasets for downstream integration and querying.
- Track record of automating data workflows and building for reliability, monitoring, and data quality.
- Comfortable operating with loosely defined, evolving scope — able to self-organize, scope work with stakeholders, and deliver iteratively.
Nice to Have
- Exposure to AI/LLM engineering: document processing at scale, embeddings, and vector databases (e.g. pgvector, Pinecone, Weaviate, or similar).
- Experience building or supporting retrieval-augmented generation (RAG) workflows.
- Familiarity with cloud data platforms and modern data stack tooling.
- Prior experience supporting investment management, hedge fund, or capital markets data environments.
- Experience working as a dedicated remote resource embedded with a client team.