About the opportunity
Ericsson’s Service Architecture team, within Service Management in the Operate Service Line, seeks an AI & Cloud Architect to define the strategy for an enterprise AI platform at global telecom scale. You will shape secure, reliable cloud-native capabilities that automate and augment managed service operations across ML/LLM platforms, MLOps, vector databases, knowledge graphs, RAG, and Agentic AI. As technical authority, you will connect engineering, product, enterprise architecture, and senior leadership, guiding teams from PoC to production using TOGAF-aligned practices.
What you will do
-
Own architecture for data ingestion, feature stores, model training, inference, retrieval, graph context, and agent orchestration.
-
Define reusable reference architectures for LLM, RAG, and multi-agent solutions.
-
Lead architecture reviews, ADRs, technology radar updates, and build-versus-buy decisions.
-
Design multi-region ML inference on AWS, Azure, and/or GCP, including GPU scaling, versioning, and safe deployments.
-
Govern MLOps for versioning, registries, validation, retraining, and release promotion.
-
Set patterns for self-hosted models using vLLM/TGI/Triton and commercial models through secure API gateways.
-
Govern model selection, fine-tuning, quantisation, batching, latency, accuracy, cost, security, and data residency.
-
Establish vector database standards using Pinecone, pgvector, Weaviate, Milvus, or equivalent, including embeddings, hybrid search, re-ranking, privacy, and index lifecycle.
-
Design property-graph and RDF/ontology architecture using Neo4j, Neptune, Cypher, or SPARQL for graph-enhanced reasoning.
-
Architect agent memory, tools, guardrails, human escalation, recovery, evaluation, and golden datasets; integrate agents with OSS/BSS, APIs, and event streams.
-
Set standards for IAM, network security, PII redaction, audit logging, GDPR/AI Act, OpenTelemetry, SRE, and GPU FinOps.
You will bring
-
8–15 years of software/systems engineering experience, including 4+ years as an architect or principal engineer.
-
Bachelor’s/master’s degree in computer science, Software Engineering, AI/ML, or equivalent; advanced degree preferred.
-
Proven production delivery of large-scale AI/ML systems with vector retrieval, knowledge graphs, and Agentic AI/LLM systems using open-source and licensed models.
-
Strong cloud-native experience on AWS, Azure, or GCP; multi-cloud is an advantage.
-
TOGAF Foundation/Practitioner or equivalent enterprise architecture practice.
-
Expert Python and proficient PySpark; experience with Spark, SQL, Kafka/Flink, Delta Lake/Iceberg, and distributed ML.
-
Knowledge of supervised, unsupervised, deep learning, reinforcement learning, and sequence/time-series models, including validation and bias assessment.
-
Experience with Terraform/Pulumi, Kubeflow, MLflow, Weights & Biases, LangGraph/AutoGen/CrewAI/Semantic Kernel, OpenTelemetry, Prometheus, and Grafana.
-
Strong security, compliance, observability, and reliability skills.