About Company : Sign3 Labs is an AI-native fraud and customer intelligence platform purpose-built for Indian Financial Services. We work with major banks, NBFCs, small finance banks, enterprises and FinTech across the country, helping them make risk decisions at onboarding, login and transaction. We are a venture-backed, revenue-generating company scaling commercially across India, with a growing team spanning engineering, data science, GTM and operations.
Role: Data Scientist Junior
Roles & Responsibilities:
Core modelling
- Binary classifiers under heavy class imbalance. Fraud base rates are low. You will build GBMs (XGBoost, LightGBM) and shallow neural models that have to perform at the low-FPR operating points clients actually care about — not at default thresholds.
Feature engineering across signal families. Device, behavioral, location, image, telecom and alternate-data signals, plus velocity aggregations across multiple time windows. You will learn to turn raw telemetry into features that survive adversarial drift.
Anomaly and unsupervised detection. Isolation Forest, auto encoders, PCA residuals, peer-group outliers — for the cold-start cases where labeled fraud is weeks away.
Graph features (growing scope).Contribute to our graph-based fraud and community-detection work. You will start with graph-derived tabular features and can grow deeper into graph modelling if the work suits you.
Productionization and monitoring
- Deployment. Package models for our real-time decisions path. Understand latency budgets and feature availability at scoring time (no leakage, no missing features in production).
Monitoring. PSI, CSI, score-distribution drift, feature drift, fraud-capture curves by client. Own the weekly model-health review for at least one deployed model within six months.
Threshold economics. Work with clients to translate fraud capture customer-friction trade-offs into operating points they can sign off on. This is where most fraud models actually fail, and where you will learn the most.
Collaboration
- With engineering, on feature pipelines, offline/online parity and real-time inference.
With product and clients, on what a model is actually allowed to flag, block or auto-approve. You will sit in review calls with bank fraud heads.
Required Skills
- 1–3 years of hands-on ML experience building supervised models on real data (not only coursework or Kaggle).
- Strong Python and SQL. Comfortable with pandas/Polar, scikit-learn, XGBoost or LightGBM, and writing non-trivial SQL against large tables.
- Statistical intuition. You know why accuracy is the wrong metric for fraud, can explain ROC vs PR curves, and understand calibration and threshold selection.
- Engineering hygiene. Git, code review, reproducible experiments, basic CI. Your models should run when someone else checks out your branch.
- Clear writing. You can write a one-page memo that a non-technical product manager and a fraud-ops lead both walk away understanding.
- Exposure to fraud, risk, credit, AML or payments data — at a bank, FinTech, card network or risk-tech vendor.
- Experience with imbalanced-class techniques beyond naive resampling (focal loss, cost-sensitive learning, threshold-moving, calibration).
- Experience with real-time feature serving, model monitoring or MLOps tooling (MLflow, Feast, SageMaker, Vertex, or equivalents).
- Exposure to graph data, NLP/NER, or geospatial features.
- A public repo, paper, competition placement or blog post that shows how you think.
What We Offer:
- Compensation benchmark to top-quartile Indian FinTech; ESOPs for all full-time employees with a clear vesting schedule.
- Health insurance for you and your family, and a meaningful learning & conference budget.
- Access to real, messy, high-volume Indian financial data and the compute to do something interesting with it.