Sr Associate Data Scientist (Databricks Platform)
Actively Hiring
Full-time $84k - $140k Posted 6 months ago
Responsibilities
- check_circle Build, train, evaluate, and optimize machine‑learning models using Spark MLlib, Python, and cloud‑based toolchains
- check_circle Perform exploratory data analysis (EDA), statistical profiling, and feature engineering on large-scale datasets hosted in Databricks
- check_circle Implement and manage MLflow experiment tracking, model registry, versioning, and reproducibility workflows
- check_circle Contribute to model monitoring, performance tuning, drift detection, and continuous improvement.
- check_circle Develop notebooks, jobs, and workflows within Databricks for data preparation, model training, and batch/streaming inference
- check_circle Utilize Unity Catalog for secure, governed data access, lineage, and metadata management
- check_circle Work with Delta Lake (bronze/silver/gold layers) for scalable feature pipelines supporting both training and production
- check_circle Collaborate with Engineering to migrate workloads to Databricks and support transformations, optimizations, and cost‑efficient compute usage.
- check_circle Build reusable, production‑grade feature pipelines in PySpark and SQL
- check_circle Implement data validation, quality checks, and transformation logic consistent with enterprise guidelines
- check_circle Participate in design sessions for ingestion, medallion architecture workflows, and schema evolution
- check_circle Partner with Data Engineering, Analytics, Product, and SMEs to translate business problems into data‑driven solutions
- check_circle Document model assumptions, data transformations, evaluation metrics, and deployment patterns
Basic qualifications
- Bachelor’s degree in Data Science, Computer Science, Analytics, Math, Statistics, Engineering, or related field, or related experience
- Typically requires 2+ years of experience in applied ML, data science, or advanced analytics
- Hands-on experience with Python, PySpark, SQL, and Git-based workflows
- Practical exposure to cloud-based ML environments (preferably Databricks)
- Understanding of ML techniques such as regression, classification, clustering, time-series forecasting, and embeddings
- Ability to work with large, complex datasets
Preferred qualifications
- Experience with Databricks MLflow, model serving, and workflow orchestration
- Familiarity with Delta Lake storage formats, feature engineering at scale, and medallion architecture patterns
- Experience deploying models into production environments with monitoring and observability
Tags & Focus Areas
Remote Data Science Ai
About McKesson
McKesson is an impact-driven, Fortune 10 company that touches virtually every aspect of healthcare. We are known for delivering insights, products, and services that make quality care more accessible and affordable. Here, we focus on the health, happiness, and well-being of you and those we serve – we care.