Be at the Forefront of the Agentic AI Revolution
At Skan AI, you'll be part of the team pioneering the context engine for human and agentic execution, bringing context from enterprise operators, systems, and processes to power how the world's largest organizations execute their most complex, mission-critical work.
Why Join Skan AI
We're in hyper-growth mode at exactly the right moment in history. As enterprises race to adopt agentic AI, we're uniquely positioned to deliver the clear signal they desperately need: a platform that trains and grounds AI Agents in trillions of real execution signals, enabling reliable, compliant automation of their most complex processes.
Backed by Dell Technologies Capital and other leading investors, we're the only company that can bridge the gap between AI's promise and enterprise reality, making us perfectly positioned to define the agentic era for modern enterprises.
Our diverse, collaborative team of 250+ innovators is solving category-defining challenges at the intersection of AI, process intelligence, and enterprise work. Diverse perspectives fuel breakthrough thinking, cross-functional collaboration is the norm, and our work directly transforms how Fortune 500 companies operate. We are shaping the future of work itself.
We are seeking a Senior Database Administrator (DBA) with deep, hands-on expertise across PostgreSQL, MongoDB, and Databricks to own the health, performance, scalability, and reliability of our large-scale data environments. This is a highly technical role for someone who thrives on operating databases at very high data volumes and who takes pride in building systems that are fast, resilient, and cost-efficient.
The primary focus of this role is administration and platform engineering: configuring, setting up, and continuously optimizing database platforms; instrumenting comprehensive monitoring and metrics; right-sizing environments in response to changing demand; and standing up schemas and configurations aligned to evolving requirements. As a secondary but important focus, the role supports data analysis activities, troubleshoots production data issues, plans capacity, performs maintenance, and takes proactive action to prevent outages before they occur.
The ideal candidate is technically outstanding across all three platforms, highly collaborative, and genuinely eager to learn and grow as our data estate and technology stack evolve.
-
Provision and configure PostgreSQL, MongoDB, and Databricks environments from the ground up, applying tuned parameters, storage layouts, and connection strategies appropriate to very large data volumes.
-
Optimize database configuration on an ongoing basis — memory, buffers, connection pooling, indexing strategy, partitioning/sharding, vacuum/compaction, query planning, and workload isolation — to sustain performance as volumes grow.
-
Stand up schemas and configurations based on functional and non-functional requirements gathered from engineering, analytics, and product stakeholders, balancing normalization, access patterns, and performance.
-
Manage Databricks clusters, workspaces, and compute policies, tuning Spark configurations, autoscaling behavior, and storage (Delta Lake) layouts for efficient processing of very large datasets.
-
Design and implement monitoring across all three platforms — throughput, latency, query performance, replication lag, lock contention, storage utilization, and resource saturation.
-
Establish meaningful metrics, dashboards, and alerting thresholds that surface issues early and drive proactive intervention rather than reactive firefighting.
-
Define and track service-level objectives (SLOs) and key operational indicators, and report on the health and performance of the data estate.
-
Resize and scale environments — vertically and horizontally — based on real-time and forecasted demand, ensuring performance during peaks while controlling cost during troughs.
-
Implement and tune autoscaling, sharding, replication, and partitioning strategies to accommodate sustained growth in data volume and concurrency.
-
Continuously evaluate cost-to-performance trade-offs across cloud storage and compute tiers.
-
Troubleshoot and resolve production data issues, including slow queries, data anomalies, replication problems, contention, and platform incidents — often under time pressure.
-
Monitor for performance degradation and take proactive action to prevent outages before they impact users, based on data-driven signals and early warning indicators.
-
Participate in incident response and on-call rotations, conduct root-cause analysis, and drive durable fixes and preventative measures.
-
Plan for capacity upgrades, forecasting growth in storage, compute, and I/O, and scheduling upgrades ahead of demand.
-
Perform routine and preventative maintenance — patching, version upgrades, index maintenance, vacuum/compaction, backup verification, and configuration hygiene — with minimal disruption.
-
Maintain runbooks, operational documentation, and disaster-recovery procedures.
-
Support data analysis activities by assisting analysts and engineers with query optimization, data access, and structuring data for efficient analytical workloads.
-
Partner with analytics teams to ensure Databricks and downstream platforms deliver reliable, performant access to large datasets.
A defining requirement of this role is the ability to architect data platforms that remain high-performing, highly available, and recoverable at very high data volumes.
-
Architect high-performance database designs capable of sustaining very high transaction and query volumes without degradation.
-
Design high-availability (HA) architectures — replication, clustering, failover, and multi-zone/multi-region topologies — to meet demanding uptime targets.
-
Design and validate disaster-recovery (DR) strategies, including backup/restore, replication, failover runbooks, and defined RPO/RTO targets, with regular DR testing.
-
Translate business and technical requirements into scalable, resilient, and cost-effective platform designs across PostgreSQL, MongoDB, and Databricks.
-
Extensive hands-on database administration experience, with senior-level depth in PostgreSQL, MongoDB, and Databricks.
-
Demonstrated experience scaling data environments for very large data volumes, including sharding, partitioning, replication, and performance tuning at scale.
-
Deep expertise in database configuration, optimization, indexing, and query performance tuning.
-
Proven experience designing and operating high-availability and disaster-recovery solutions for high-volume databases.
-
Strong skills in monitoring and observability — instrumenting metrics, dashboards, and alerting to drive proactive operations.
-
Experience with capacity planning, maintenance, patching, and upgrade management in production environments.
-
Solid command of SQL and a scripting/automation language (e.g., Python, Bash) for operational tooling and automation.
-
Working knowledge of cloud data platforms, storage/compute scaling, and Delta Lake / Spark concepts within Databricks.
-
Strong troubleshooting and root-cause analysis skills, with the ability to stay calm and effective during production incidents.
-
Experience with infrastructure-as-code and automation (e.g., Terraform, Ansible) and CI/CD for database changes.
-
Familiarity with containerization and orchestration (Docker, Kubernetes) for database and data workloads.
-
Exposure to data governance, security, encryption, and compliance requirements at scale.
-
Relevant certifications across PostgreSQL, MongoDB, Databricks, or major cloud providers.
-
Prior experience mentoring engineers or leading platform initiatives.
-
Technically savvy — genuinely strong and current across all three platforms, with the instincts to diagnose and optimize complex systems.
-
Highly collaborative — works closely with engineering, analytics, and product teams and communicates clearly with both technical and non-technical audiences.
-
Eager to learn and grow — curious, adaptable, and motivated to expand skills as the data estate and technologies evolve.
-
Proactive and ownership-minded — anticipates problems, prevents outages, and takes end-to-end responsibility for platform health.
-
Detail-oriented and reliable — disciplined about maintenance, documentation, and operational rigor.
Skan AI is an equal opportunity employer committed to building a diverse, inclusive, and respectful workplace around the world. We do not discriminate based on race, color, religion or belief, sex (including pregnancy, sexual orientation, gender identity, or gender expression), national origin, ancestry, age, disability, medical condition, genetic information, marital or family status, military or veteran status, or any other characteristic protected by applicable laws in the locations where we operate.
We welcome people from all backgrounds and provide reasonable accommodations throughout the hiring process.