D
DevOps/MLOps Engineer
Actively Hiring
Full-time $120k - $150k Posted 8 months ago
Role overview
You'll own the entire deployment pipeline and model serving infrastructure. This is a hybrid DevOps + MLOps role – you'll ensure our application deploys reliably AND that our AI models (both frontier and local) serve efficiently.
Our cost optimization strategy requires routing between expensive frontier models (Claude, GPT) and cost-effective local models (Llama, Mistral) based on task complexity. You'll build and own this infrastructure.
Responsibilities
- check_circle CI/CD pipelines – Automated build, test, and deploy on every push
- check_circle Infrastructure as code – Terraform/Pulumi for reproducible environments
- check_circle Monitoring & alerting – Know when things break before customers do
- check_circle Incident response – Own uptime and reliability
- check_circle Daily deploys – Enable the team to ship to production every day safely
- check_circle Model serving infrastructure – Deploy and serve LLMs (local and API-based)
- check_circle Model router – Build the abstraction layer that routes requests to appropriate models
- check_circle GPU infrastructure – Manage inference servers for local models (Llama, Mistral)
- check_circle Cost optimization – Track and optimize model usage costs
- check_circle Model versioning – Safe rollouts and rollbacks for prompt/model changes
- check_circle Developer experience – Make the team faster through better tooling
- check_circle Scaling – Prepare infrastructure for growth
- check_circle Infrastructure security – Server hardening, network security, firewall configuration, VPC design
- check_circle Secrets management – Vault, AWS Secrets Manager, or similar; no secrets in code
- check_circle Access control – IAM policies, least-privilege principles, SSO integration
- check_circle Vulnerability scanning – Automated scanning in CI/CD, dependency audits, container scanning
- check_circle Intrusion detection – CloudTrail, GuardDuty, or similar; alert on suspicious activity
- check_circle Encryption – Data at rest and in transit; key management
- check_circle Incident response – Work with fractional CISO to implement detection, containment, and recovery procedures
- check_circle Compliance – Support audits and maintain security documentation
- check_circle CI/CD quality gates – Automated tests run on every push; bad code doesn't deploy
- check_circle Test environment management – Staging environments that mirror production
- check_circle LLM output monitoring – Track hallucinations, wrong tool calls, response quality in production
- check_circle Security scanning – Automated vulnerability scanning in CI pipeline
- check_circle Alerting & anomaly detection – Know when something breaks before customers do
- check_circle Cloud: AWS (EC2, RDS, S3, Lambda)
- check_circle Containers: Docker
- check_circle CI/CD: GitHub Actions
- check_circle Database: PostgreSQL (RDS)
- check_circle Caching: Redis
- check_circle Model serving: vLLM, Ollama, or similar for local inference
- check_circle GPU compute: AWS/GCP GPU instances or dedicated inference providers
- check_circle Model routing: Custom abstraction layer for model selection
- check_circle Observability: Datadog, Grafana, or similar for unified monitoring
- check_circle 3+ years DevOps/SRE/Platform engineering experience
- check_circle Strong AWS experience (EC2, RDS, Lambda, IAM, VPC)
- check_circle Infrastructure as code (Terraform, Pulumi, or CloudFormation)
- check_circle CI/CD pipeline design and maintenance
- check_circle Docker and container orchestration
- check_circle Monitoring and observability tools
- check_circle MLOps experience – Model deployment, serving, monitoring
- check_circle GPU infrastructure – Managing inference workloads
- check_circle Experience with LLM serving (vLLM, TGI, Ollama)
- check_circle Kubernetes experience
- check_circle Cost optimization mindset
- check_circle Experience serving both frontier APIs and local models
- check_circle LangChain/LangSmith or similar LLM observability
- check_circle Startup experience – comfort with ambiguity and speed
- check_circle Texas location
- check_circle AI-augmented development – We use AI tools extensively; you'll automate everything possible
- check_circle Daily deploys – Your pipelines enable the team to ship constantly
- check_circle Async communication – Written updates, minimal meetings
- check_circle On-call rotation – Shared responsibility for production (small team = everyone contributes)
Benefits
- check_circle 401(k)
- check_circle Dental insurance
- check_circle Health insurance
- check_circle Paid time off
- check_circle Vision insurance
Tags & Focus Areas
Fulltime Remote Ai Mlops Generative Ai