You will build the storage and retrieval backbone of Aiva: the context graph itself, the ingestion pipelines that populate it, and the query layer that serves it to agents in real time. This is a deep backend role for an engineer who cares about data modeling, query performance, and correctness under scale — and who wants their work to sit on the critical path of every agent interaction across the Aivar platform.
Key Responsibilities
-
Implement the context graph data layer: schema, entity and relationship management, versioning, and time-aware (temporal) memory records.
- Build streaming and batch ingestion pipelines that capture events from Convogent conversations, Velogent process runs, and customer systems of record — with idempotency, ordering, and replay support.
- Develop entity resolution and deduplication: matching customers, accounts, and interactions across channels into a single coherent graph identity.
- Implement the hybrid retrieval service — graph traversal, vector k-NN, and metadata filtering — with caching, pagination, and strict latency budgets.
- Build and maintain Aiva's public APIs and internal SDKs (REST/gRPC, and MCP-compatible tool surfaces) with robust authentication, rate limiting, and multi-tenant scoping.
- Own performance engineering: query profiling, index design, load testing, and capacity/cost tuning across graph and vector stores.
- Implement data lifecycle mechanics: TTL and decay policies, archival, tenant-scoped deletion (right-to-be-forgotten), and audit logging in line with AI Gateway governance policies.
- Write thorough tests, runbooks, and operational dashboards; participate in on-call for the platform.
Must-Have Qualifications
-
6+ years of backend engineering in Python, Go, or Java/Kotlin, building services in production at meaningful scale.
- Hands-on production experience with at least one graph database (Neptune, Neo4j, ArangoDB) or a demonstrated ability to model highly connected data in relational/NoSQL stores.
- Practical experience with search/vector retrieval: OpenSearch/Elasticsearch, pgvector, or a dedicated vector database, including index tuning and relevance debugging.
- Strong data pipeline skills: Kafka/Kinesis or equivalent, exactly-once/idempotent processing patterns, and schema evolution.
- Solid AWS experience (EKS or Lambda, DynamoDB, S3, IAM) and infrastructure-as-code (Terraform/CDK).
- Rigor around API design, testing, observability (metrics, tracing), and secure multi-tenant service patterns.
Nice-to-Have
-
Experience with entity resolution, identity graphs, CDPs, or master data management.
- Exposure to RAG systems or LLM-serving infrastructure.
- Familiarity with Gremlin/openCypher/SPARQL query languages.
- Experience meeting compliance requirements (data residency, PII handling) in BFSI, healthcare, or telecom.