ML Engineer - Tabular Data Experimentation
Actively Hiring
Full-time Posted 8 days ago
Responsibilities
- check_circle End-to-end ML on tabular and panel data: feature engineering, validation strategy including time-aware splits, gradient boosting and related methods, calibration, drift monitoring.
- check_circle Designing A/B tests that answer real questions: hypothesis design, power and MDE, handling peeking and multiple testing, variance reduction (CUPED and similar), interference and network effects.
- check_circle Extracting credible answers from offline data when online experimentation isn't feasible - DiD, synthetic control, IV, uplift modeling - and choosing the right method for the situation rather than the most fashionable one.
- check_circle Knowing when offline evaluation is sufficient and when it isn't, and being willing to defend that judgment.
- check_circle Designing evaluations for LLM-based components the team owns: offline metrics, online proxies, drift detection, power analysis, guardrail metrics - with the same rigor as classical ML.
- check_circle Translating business questions into ML formulations: metrics, loss, constraints, the trade-offs that actually matter to the product.
Basic qualifications
- 5+ years in applied ML with real product impact, not Kaggle-only.
- Deep working knowledge of tabular and panel data: temporal leakage, non-stationarity, the ways these problems differ from i.i.d.
- Solid statistics and experimentation - comfortable designing a test from scratch, and equally comfortable pushing back when someone wants to "just look at the p-value."
- Causal inference at the depth where you can pick a method, justify it, and explain what it doesn't tell you.
- Production-grade Python - maintainable code, not just research notebooks.
- SQL at production-analytics level; able to get to the right sample independently, including non-trivial joins, window functions, and reasoning about query plans.
- A clear view of when an LLM is the right tool for a problem and when a GBDT or classical method is, and the ability to defend either choice with evaluations and cost.
- Building or substantially reworking a production recommender - two-tower, GBDT ranking, rerankers, candidate generation and ranking at scale.
- LLM-augmented retrieval and ranking: semantic retrieval, LLM rerankers, embedding-based candidate generation.
- Causal ML tooling (DoubleML, EconML).
- MLflow, W&B, feature stores.
- LLM-as-judge methodology and a working understanding of its failure modes.
- Prompt optimization as an empirical practice (DSPy-style or hand-rolled).
Benefits
- check_circle Stock options grant (we’re a Silicon Valley Company)
- check_circle Competitive salary
- check_circle Medical insurance for you and 75% off for your relatives
- check_circle On-site position with 4 days at the office and 1 day WFH
- check_circle Budget for lunch
- check_circle Parking
- check_circle Multisport card
- check_circle Cheerful team spirit and fun office atmosphere
Tags & Focus Areas
Remote Machine Learning Ai
About TANGOME
Ready to Join the Team?
Apply once with DevFound — we route your profile to TANGOME and keep you posted on matching AI roles.