About us
Factori is a rapidly growing location and consumer intelligence company built for an AI-first world. Each day we ingest billions of privacy-compliant signals and maintain the world’s most comprehensive consumer, mobility and point-of-Interest database.
Our AI-driven insights help Fortune 500 brands, global media agencies and technology innovators decide where to invest, build and grow. Headquartered in New York City with hubs across APAC, EMEA and the Americas, we operate with the ownership mindset of a startup and the global footprint of an enterprise data leader.
About the Role:
We’re looking for a Data Engineer with 3–6 years of experience to design, build, and operate large-scale data platforms. You’ll develop high-performance batch and streaming pipelines, optimize distributed data processing systems, and contribute to the infrastructure that powers our AI and analytics products.
Key Responsibilities:
-
Build and maintain scalable batch and streaming data pipelines using Apache Spark and distributed data technologies.
-
Design, develop, and optimize ETL workflows with a focus on reliability, performance, and data quality.
-
Develop efficient SQL queries, data models, and storage solutions across systems such as Bigquery, PostgreSQL, Elasticsearch, DuckDB, and cloud native data platforms.
-
Build and maintain workflow orchestration using Airflow (or similar schedulers).
-
Monitor, troubleshoot, and optimize production data pipelines to ensure reliability and performance.
-
Collaborate with product and engineering teams to design and deliver resilient, production-grade data systems.
-
Leverage AI-assisted development tools to improve development speed, code quality, and engineering productivity.
What we are looking for :
3–6 years of experience building distributed data processing systems or large-scale data platforms.-
Strong understanding of Apache Spark, distributed computing concepts, ETL design, HDFS, and SQL.
-
Hands-on experience with relational and analytical databases such as PostgreSQL, Elasticsearch, DuckDB, or similar data storage technologies.
-
Proficiency in Java or Python with strong software engineering fundamentals and object-oriented programming skills.
-
Experience with Airflow (or similar workflow orchestration tools) and at least one public cloud platform (GCP preferred; AWS or Azure also welcome).
-
Familiarity with Linux environments, Bash scripting, Git, and cloud-native development practices.
-
Experience with data formats and table/storage technologies such as Parquet, Avro, ORC, Iceberg, or Delta Lake is a plus.
-
Hands-on experience using AI-assisted development tools such as GitHub Copilot, ChatGPT, Cursor, or Claude to accelerate development and improve code quality.
You’ll thrive here if you :
-
Take ownership and enjoy solving challenging engineering problems.
-
Build reliable, scalable, and maintainable software with a strong focus on quality.
-
Continuously learn new technologies and improve your engineering skills.
-
Communicate effectively and collaborate well across engineering and product teams.
-
Enjoy working in a fast-paced environment where your work has a direct impact on customers and the business.
Good to have
-
Experience with Kafka, Pub/Sub, or other streaming technologies.
-
Experience with Kubernetes, Docker, or infrastructure-as-code tools.
-
Exposure to data warehousing technologies such as BigQuery, Snowflake, or Redshift.
-
Understanding of CI/CD, observability, and production monitoring for data platforms.
Benefits :
-
Early impact – Help shape the tech stack and build products from the ground up.
-
Agile culture – Small teams, zero bureaucracy.
-
Great benefits – Group Health Insurance, daily breakfast, Friday team lunch, Fun O’Clock Fridays, and unlimited coffee, tea & snacks.
How to apply :
For more information visit - https://factori.ai