Loading…
Loading…
Loading…
Go from first principles to running a production ML platform: data versioning, feature stores, CI/CD/CT, Kubernetes serving, monitoring, drift detection, security, cost, LLMOps, and incident response.
Haithem
Formateur
Les articles payants se règlent par virement bancaire. Après la commande, les instructions complètes apparaissent dans Mes achats (connexion requise). Nous activons l'accès après vérification de votre reçu.
Secure checkout · Card or bank transfer
Ce cours inclut :
MLOps is what separates a model that works in a notebook from a system that keeps working in production while data, code, and the world around it keep changing. This course takes you from the ML lifecycle and team topologies all the way to running a governed, monitored, cost-aware ML platform on Kubernetes.
You'll build reproducible training pipelines on versioned data, wire experiment tracking and a model registry into CI/CD/CT pipelines with real promotion gates, containerize and serve models with canary and shadow rollouts, monitor for drift and quality regressions, threat-model and secure multi-tenant systems, operate LLM/RAG pipelines, and run a live incident drill with a blameless postmortem.
ML engineers, data scientists moving into production ownership, and platform/DevOps engineers picking up ML-specific operational practice. Comfortable with Python and the command line; deep ML theory is not required — this course is about operating ML systems, not deriving algorithms.
MLOps Engineering: The Complete Course
$49.99
The ML lifecycle and MLOps maturity levels
Team topologies: platform vs embedded ML engineering
Packaging and project structure for ML repositories
Testing ML code: unit, property, and data tests
Deterministic seeds and environment capture
Config-driven training with structured configuration
Batch vs streaming ingestion and ETL/ELT for ML
Data quality checks and data contracts
Data versioning fundamentals with DVC
Lineage tracking and reproducing training data
Feature store architecture: offline and online stores
Point-in-time correctness and training-serving skew
Experiment tracking with MLflow
Weights & Biases and experiment governance at scale
Portable model formats: ONNX and TorchScript
Container-based packaging and artifact signing
The MLflow Model Registry and stage transitions
Rollback-by-registry-pointer
CI pipelines for ML repositories
Testing data contracts and model behavior in CI
Automated retraining triggers
Promotion gates and approval workflows
Docker fundamentals for ML: GPU images and dependency layers
Multi-stage builds and image security scanning
Kubernetes primitives for training: Jobs and CronJobs
Autoscaling inference with HPA and KEDA
Airflow DAGs for training pipelines
Kubeflow Pipelines, Prefect, and Dagster
Terraform basics for ML infrastructure
Environment parity and IaC review gates
Managed ML platforms: SageMaker, Vertex AI, and Azure ML
Cost, vendor lock-in, and multi-cloud abstraction patterns
Batch scoring jobs versus online inference
Online serving protocols: REST and gRPC, and request batching
Triton Inference Server and TorchServe
BentoML and KServe: multi-model serving on Kubernetes
Kafka and streaming feature pipelines
Real-time scoring architectures and delivery tradeoffs
Canary and shadow deployments for models
Blue-green deployments and online A/B testing
Golden signals and SLOs for ML services
Dashboards for prediction distributions and slice metrics
Data drift versus concept drift
Statistical drift tests: PSI and KS, and automated alerting
Threat modeling ML systems
Access control, tenant isolation, and compliance for ML
GPU cost management and spot/preemptible training
Cost attribution and cost-aware architecture decisions
Prompt and version management, and RAG pipeline operations
LLM evaluation harnesses, latency/cost tradeoffs, and guardrails
On-call runbooks and incident drills for ML services
Postmortems and disaster recovery for models, data, and registries
Explaining predictions with SHAP and LIME
Fairness and bias auditing as a release gate
Quantization, pruning, and knowledge distillation
Industry case studies and MLOps interview preparation
Capstone I: End-to-End MLOps Platform Build
Capstone II: Production Incident Game Day & Technical Defense
Aucun avis pour le moment. Soyez le premier !
Inscrivez-vous au cours pour participer à la discussion.
Aucun commentaire. Lancez la discussion !